Hybrid calibration-free multi-person motion capture system based on multiple LiDARs and multiple extensible mobile cameras

The A3Cap method uses multi-LiDAR and scalable mobile multi-camera system, using data fusion and noise reduction technology, to solve the problem of insufficient flexibility and robustness of traditional motion capture methods in dynamic environments, and achieves high-precision multi-person motion capture.

CN120371126APending Publication Date: 2025-07-25SHANGHAI TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510446923.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing motion capture methods are not flexible and robust in dynamic or complex environments, relying on sensor calibration matrix leads to low accuracy, and the noise of multiple sensors is severely affected.

Method used

Using the A3Cap method, through multi-LiDAR and scalable mobile multi-camera, a universal motion estimator, an anti-noise tracker and an online track optimizer is used to realize data fusion and noise reduction of multiple vision sensors, eliminate calibration matrix dependence, and improve robustness and real-timeness.

Benefits of technology

It realizes flexible, robust and efficient human motion capture in dynamic environments, supports any combination of sensors, improves accuracy and is suitable for online robot applications.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The technical scheme of the invention discloses a hybrid calibration-free multi-person motion capture system based on multiple LiDARs (Light Detection and Ranging) and multiple extensible mobile cameras. The invention discloses a novel method named A3Cap, which can capture multi-person movement in real time and with high quality in an unlimited environment and get rid of limitation on capture time, capture place or visual sensor. According to the method disclosed by the invention, through effective network design, the advantages of various visual sensors are fully utilized, the dependence on a calibration matrix is eliminated, and challenges such as sensor noise and online deployment requirements are solved. According to the A3Cap method, the advantages of various visual sensors are fully utilized, and flexible, robust and efficient human body motion capture in a dynamic environment is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a hybrid uncalibrated multi-person motion capture system based on multiple LiDARs and scalable mobile multi-cameras, which can perform human motion capture using any number of visual sensors at any time and any place. Background Art

[0002] Human motion capture plays an important role in fields such as animation, healthcare, sports science, and robotics. Traditional motion capture methods usually rely on wearable sensors such as markers, inertial measurement units, and ego-cameras, but these methods have problems such as sensor drift, discomfort for the wearer, and difficulty in capturing certain actions in dynamic or complex environments. In recent years, vision sensor-based methods have received extensive attention due to their convenience in dynamic environments. However, these methods usually require fixed sensor types, quantities, and perspectives, rely on accurate sensor calibration matrices, which limit the flexibility of the motion capture system. In addition, the additional sensor noise and calibration errors introduced by multi-sensor configurations may affect the accuracy and robustness of motion capture. Summary of the Invention

[0003] The object of the present invention is to provide a flexible, robust, and motion capture method applicable to various scenarios.

[0004] To achieve the above object, the technical solution of the present invention discloses

[0005] The present invention discloses a new method called A3Cap, which can capture multi-person motion in real time and with high quality in an unrestricted environment, getting rid of the limitations on the capture time, location, or visual sensors. The method disclosed by the present invention makes full use of the advantages of various visual sensors through an effective network design, eliminates the dependence on the calibration matrix, and solves challenges such as sensor noise and online deployment requirements.

[0006] The A3Cap method provided by the present invention realizes flexible, robust, and efficient human motion capture in a dynamic environment by making full use of the advantages of multiple visual sensors. Compared with existing methods, the present invention has the following beneficial effects:

[0007] 1) Flexibility: Supports any combination of visual sensors, without being restricted by sensor types, quantities, and perspectives;

[0008] 2) Robustness: Through a noise-resistant trajectory tracker and an online trajectory optimizer, the robustness against sensor noise and different perspectives is improved;

[0009] 3) High efficiency: Real-time motion capture is achieved, suitable for practical scenarios such as online robot applications. High precision: It performs excellently in multiple datasets and real scenarios, outperforming the current state-of-the-art methods. Detailed implementation manners

[0010] The following further elaborates the present invention in combination with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0011] A hybrid calibration-free multi-person motion capture system based on multi-LiDAR and scalable mobile multi-cameras disclosed in an embodiment of the present invention includes the following key components:

[0012] General motion estimator - This component realizes the precise fusion of data of different sensor types, quantities, and perspectives through the human body center feature space and the human body center feature integrator, thereby accurately estimating the global direction, posture, and shape of the human body. The specific implementation steps are as follows:

[0013] Step 101, Feature extraction: Extract 3D human joints from the point clouds of multiple LiDARs and convert them to the human body center space;

[0014] Step 102, Feature fusion: Use learnable tokens and attention mechanisms to fuse 3D and 2D features and extract a unified feature representation.

[0015] Step 103, Pose decoding: Decode the human body pose and shape through a bidirectional RNN layer and predict SMPL parameters.

[0016] Noise-resistant trajectory tracker - This component improves the translation accuracy by iteratively correcting the global human body position, gradually reducing the prediction space, and reducing the variance caused by noise. The specific implementation steps are as follows:

[0017] Step 201, Offset prediction: Predict the offset between the normalized center of the point cloud and the coordinate origin;

[0018] Step 202, Iterative correction: Gradually correct the global human body position through multiple iterations to finally determine the position of the human body in the LiDAR coordinate system.

[0019] Online real-time motion capture subsystem - This subsystem enhances the accuracy and smoothness of the global trajectory by using an online trajectory optimizer and ensures real-time performance. The specific implementation steps are as follows:

[0020] Step 301, Data preprocessing: Detect, track, and match the input image and point cloud data;

[0021] Step 302, Trajectory optimization: Optimize the current prediction based on past and current prediction results to ensure the natural smoothness of the trajectory.

[0022] Through the above invention content, the present invention provides a high-quality human motion capture solution applicable to various scenarios, with broad application prospects.

Claims

1. A hybrid calibration-free multi-person motion capture system based on multi-LiDAR and scalable mobile multi-cameras, characterized in that, Including: A general motion estimator, which is used to integrate data of different sensor types, quantities, and perspectives through a human body center feature space and a human body center feature integrator, so as to accurately estimate the global direction, posture, and shape of the human body; A noise-resistant trajectory tracker, which iteratively corrects the global human body position, gradually narrows the prediction space, reduces the variance caused by noise, and improves the translation accuracy; An online real-time motion capture subsystem, which uses an online trajectory optimizer to enhance the accuracy and smoothness of the global trajectory by using temporal consistency to ensure real-time performance.

2. A hybrid calibration-free multi-person motion capture system based on multi-LiDAR and scalable mobile multi-cameras as claimed in claim 1, characterized in that, The specific implementation steps of the general motion estimator are as follows: Step 101, Feature extraction: Extract 3D human body joints from the point clouds of multiple LiDARs and transform them into the human body center space; Step 102, Feature fusion: Use learnable tokens and attention mechanisms to fuse 3D and 2D features and extract a unified feature representation; Step 103, Pose decoding: Decode the human body pose and shape through a bidirectional RNN layer and predict SMPL parameters.

3. A hybrid calibration-free multi-person motion capture system based on multi-LiDAR and scalable mobile multi-cameras as claimed in claim 1, characterized in that, The specific implementation steps of the noise-resistant trajectory tracker are as follows: Step 201, Offset prediction: Predict the offset between the normalized center of the point cloud and the coordinate origin; Step 202, Iterative correction: Gradually correct the global human body position through multiple iterations and finally determine the position of the human body in the LiDAR coordinate system.

4. A hybrid calibration-free multi-person motion capture system based on multi-LiDAR and scalable mobile multi-cameras as claimed in claim 1, wherein, The specific implementation steps of the online real-time motion capture subsystem are as follows: Step 301, Data preprocessing: Detect, track, and match the input image and point cloud data; Step 302, Trajectory optimization: Optimize the current prediction based on past prediction results and current prediction results to ensure the natural smoothness of the trajectory.