Multi-Person 3D Action Capture with Anti-Projection Ray Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing action capture systems struggle to accurately reconstruct three-dimensional joint points for multiple persons in crowded scenes due to camera perspective obstructions and mis-recognition of 2D joint points, leading to incorrect matching and reconstruction.

Innovation Solution

A method involving multiple cameras to obtain synchronous video frames, recognize and position 2D joint points, calculate anti-projection rays, and cluster coordinates of shortest distances between these rays to achieve a 2D personnel matching scheme, followed by 3D reconstruction for each person.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If traditional optical action capture systems are used, then interaction capability is provided, but direct human-computer interaction is not achieved and system complexity increases

Engineering Contradiction:
Improvedirect human-computer interactionVSAvoidthird-party hardware and specialized clothing
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical and optical action capture systems with a computer vision-based system using deep learning algorithms. The system processes ordinary video frames to estimate human postures, eliminating the need for specialized wearables and complex hardware while enabling direct human-computer interaction through natural movement recognition

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If OpenPose is used for single person recognition, then 3D joint point reconstruction is achieved, but in crowded scenes with multiple persons, incorrect matching occurs leading to reconstruction errors

Engineering Contradiction:
Improve3D joint point reconstruction accuracyVSAvoidpersonnel matching accuracy in crowded scenes
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the action capture process into distinct stages: 2D joint point detection in video frames, 2D personnel matching across multiple frames, and 3D reconstruction. By separating these functions and adding intermediate matching steps, the system correctly associates 2D joint points with specific persons before 3D reconstruction, preventing misattribution in crowded scenes

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces 2D personnel matching as an intermediary step between 2D joint point detection and 3D reconstruction. This intermediate process correctly identifies and associates joint points with specific persons in crowded scenes, ensuring accurate person-joint point correspondence before performing 3D reconstruction

Inventive Principle:
Principle #24Intermediary (Mediator)

3Area of stationary object

If multiple cameras are used to capture crowded scenes, then coverage is improved, but camera perspective obstructions cause mis-recognition of 2D joint points

Engineering Contradiction:
Improvescene coverage areaVSAvoid2D joint point recognition accuracy
Core Design Contradiction:
Area of stationary objectVSMeasurement precision

Solution Approach 1:

The patent introduces 2D personnel matching as an intermediary verification step that cross-checks joint point detections across multiple camera views. This intermediate process resolves ambiguities caused by perspective obstructions by confirming which detected joint points belong to which persons based on spatial consistency across cameras

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12417654B2Method for capturing three dimensional actions of multiple persons, storage medium and electronic device
Publication Date: 2025.09.16 UNILUMIN GRP
  • US12417654B2 patent drawing
  • US12417654B2 patent drawing
  • US12417654B2 patent drawing

AI summary

Disclosed are a method for capturing three dimensional actions of multiple persons, a storage media and an electronic device. The method includes: obtaining synchronous video frames of multiple cameras, recognizing and positioning joint points in each synchronous video frame of each camera, to obtain two dimensional (2D) joint point of each person in each camera; calculating an anti-projection ray of each 2D joint point, and clustering coordinates of two endpoints of a shortest distance between two anti-projection rays to obtain a 2D personnel matching scheme, wherein the anti-projection ray is from a camera pointing to its corresponding 2D joint point; and performing a three dimensional (3D) reconstruction for each person according to the 2D personnel matching scheme, to generate 3D information of each person for capturing 3D actions.