Hard to Collect Data?
Data not accurate?
Data Too Expensive?
Here's the fix.

A fusion motion-capture system for Physical AI data collection
Priced per seat like a tool, not a lab.

The capture glove — soft, ultra-thin, no rigid exoskeleton Data capture glove (part of mocap system)
±1°
Joint-angle accuracy
±3mm
Fingertip position
0
For each glove
15 bend channels
4 stretch channels
0
Battery · a full shift

Data is the bottleneck of Physical AI. Capturing it shouldn't cost like a lab.

01 · The problem

Every capture method fails somewhere.
Usually at grasping.

A serious data operation needs three things at once: accuracy you can trust as ground truth, continuity through hours of real manipulation, and a per-seat cost that scales to fleets. Grasping is occlusion — and every lost track is a retake.

Solution
Precision
Anti-interference
Comfort
Cost efficiency
Tactile
IMU (inertial)
EMG (muscle signal)
Electromagnetic
Vision only
Matrix Inno fusion
● More dots = stronger IMU drifts · EMG reads intent, not posture · EMF fails near metal · vision fails exactly at grasp
02 · Our system

A soft glove. A wearable camera.
Nothing else.

Vision tells you where the hand is. The glove tells you what it's doing — even when the camera can't see it. The two cover each other's blind spots.

System diagram: head-mounted depth camera gives absolute hand position, stretch-sensor glove tracks finger motion continuously, fusion engine outputs accurate continuous hand-motion data

Vision anchors · glove never blinks

Absolute when visible.
Continuous when not.

A head-mounted depth camera provides absolute posture: 21 keypoints are regressed, and depth back-projection recovers true 3D joint angles. Meanwhile the glove streams 19 channels of sensor data and calculates 20+ DOFs on each hand — continuously, whatever the camera sees.

Step 01

Time-synced capture

Vision computes true 3D joint angles while the glove captures raw values on every channel, time-synced.

Step 02

Paired sampling

Whenever vision locks a joint cleanly, the system logs a paired sample: visual angle + sensor value.

Step 03

Per-hand mapping

Samples build a per-channel fit from sensor value to joint angle — tuned automatically to each operator's hand.

Step 04

Occlusion-proof output

Vision keeps applying local corrections; when the hand is hidden, the glove carries the motion on its own.

Calibrate only when vision is trustworthy.

A visual angle becomes a calibration sample only when keypoint confidence is high. Our cutting edge opportunistic recalibration mechanism cancels drift for each single joint so the full hand never needs to be visible at once.

Confidence gate timeline: calibrate when the hand is clearly visible, skip when occluded or blurry — the glove runs solo

Confidence gate · calibrate only on clean frames

Live capture software: camera view, depth field, skeleton overlay, and 19 channels streaming

Live capture

Skeleton lock, depth field, and 19 channels streaming — per-joint angles in real time.

Skeleton lock on a gloved hand through YOLO detection and keypoint regression

Robust on gloved hands

A skin-toned glove plus high-contrast joint markers keeps keypoint accuracy at ≈ 0.95–0.98 after transfer learning.

03 · Validation

Tested, not promised.

Lab-grade accuracy at tool-grade cost — verified against professional optical mocap and cycled to failure on automated rigs.

Mean fingertip error (mm) · lower is better

Vision based solution on market 6–10
Matrix Inno Mocap ±3

Public depth-based benchmarks are dominated by fast motion, severe self-occlusion, and ambiguous frames. Our calibration samples are taken precisely when those conditions are absent.

500,000+ cycles on the rig.
No issue.

Sensor elements were cycled to 20%+ elongation half a million times on an automated endurance rig. The glove maintains stable, continuous and accurate output. Our vision system continuously calibrates the sensors in real time, ensuring every captured motion is as accurate as the very first.

Endurance rig with sensor element relaxed
Rig · 20% elongation
Endurance rig with sensor element stretched to 20% elongation
Rig · relaxed

Endurance & repeatability

Bending cycles (20% elongation)500,000+
Repeat-bend reproducibility (glove sensor level)±2%
The glove in hand — soft textile with a compact electronics module

Soft, ultra-thin, no exoskeleton

Operators move naturally and collect through a full shift without fatigue.

Fusion flow: depth camera ground truth and 19-channel glove signal feed one mapping — camera sees hand, calibrate; hand occluded, glove keeps going

Interference-free by design

Resistive channels drop no frames under occlusion, changing light, or electromagnetic noise.

04 · Why it scales

Scale to hundreds of operators,
no mocap studio required

Built to scale

Glove plus one commodity depth camera — a cost structure built for fleet-scale collection, not a calibrated capture volume.

Never blinks

Hands buried in a bin, wrapped around a tool, out of frame — the data keeps flowing. No retakes, no gaps in the dataset.

Tactile-ready

The platform natively hosts tactile arrays. Posture capture today extends to posture-plus-force tomorrow — same glove, same pipeline.

Your robots are waiting for this data.

Contact us for a trial

sales@matrixinno.com

Matrix Inno