Calibration system for limb motion capture

By using cameras, multi-view data fusion technology and nonlinear optimization algorithms in the motion capture system, the problem of inaccurate motion capture in the existing technology in natural scenes and complex environments is solved, and the motion capture and calibration effects with high precision, portability and adaptability are achieved.

CN120047620AInactive Publication Date: 2025-05-27SHANDONG SPORT UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510177753.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing motion capture system is difficult to achieve high-precision body motion capture and calibration in natural scenes, outdoor scenes and complex environments, and has defects such as high equipment costs, dependence on laboratory environments, occlusion problems and drift errors.

Method used

The camera, multi-view data fusion technology and nonlinear optimization algorithm are used to achieve high-precision motion capture and calibration. The hardware module includes a high-performance motion camera, a fixed structure and a time synchronization unit, the data processing module includes a calibration unit, a structured light recovery unit, a skeleton reconstruction unit and an optimization unit, and the output module provides motion capture data, three-dimensional scene structure and visual display.

Benefits of technology

Overcome the shortcomings of traditional optical and inertial capture systems, achieve high-precision motion capture and calibration in complex environments, and improve the portability, adaptability and accuracy of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047620A_ABST
    Figure CN120047620A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of motion capture, and particularly discloses a limb motion capture calibration system which comprises a hardware module, a data processing module and an output module. The hardware module comprises a camera, a fixing structure and a time synchronization unit; the data processing module comprises a calibration unit, a structured light recovery unit, a skeleton reconstruction unit and an optimization unit; and the output module comprises motion capture data, a three-dimensional scene structure and visual display. According to the invention, through the camera, the multi-view data fusion technology and the nonlinear optimization algorithm, the capturing and calibration of the limb movement are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of motion capture, and more specifically, to a calibration system for limb motion capture. Background Art

[0002] Traditional motion capture systems mainly include optical capture systems and inertial measurement systems, and these technologies are widely used in fields such as film and television production, sports analysis, and virtual reality. Existing public document 1 (Research on Skiing Motion Capture System Based on Optical Sensors [J]. Electrical Industry, 2024, (12): 44 - 48.) uses optical sensors to capture the motion trajectories and postures of athletes, and analyzes and evaluates the actions of athletes in real time through the Unity3D platform. However, this optical capture system relies on marker points to track actions, and its high equipment cost and dependence on laboratory environments significantly limit the usage scenarios. Although such systems can provide high-precision data in a controlled environment, they perform poorly in natural light, outdoor scenes, and large-scale captures, and may also lose data when there are occlusions, resulting in incomplete motion capture. Inertial measurement systems measure the angles and movements of body parts through accelerometers and gyroscopes. Existing public document 2 (Research on Human Motion Capture Method Based on IMU [D]. Nanjing University of Posts and Telecommunications, 2022. DOI: 10.27251 / d.cnki.gnjdc.2022.001171.) captures human motions based on inertial sensors, but this technology can only capture relative motions, is difficult to provide global position data, and is prone to drift errors during long-term use. The common limitation of both technologies is the difficulty in balancing portability, accuracy, and scene adaptability, especially in complex and changing environments. These deficiencies make it difficult for current technologies to meet the requirements of motion capture in natural scenes. For example, it is impossible to efficiently capture and calibrate limb motions in outdoor sports, dynamic actions, or complex scenes. Summary of the Invention

[0003] To overcome the above-mentioned defects of the prior art, the present invention provides a calibration system for limb motion capture, which realizes high-precision motion capture and calibration through cameras, multi-view data fusion technology, and non-linear optimization algorithms, and overcomes the defects of traditional optical and inertial capture systems such as dependence on laboratory environments, occlusion problems, and drift errors, so as to solve the problems raised in the above background art.

[0004] To achieve the above object, the present invention provides the following technical solutions: A calibration system for limb motion capture, comprising a hardware module, a data processing module and an output module; the hardware module includes a camera, a fixing structure and a time synchronization unit, and the data processing module includes a calibration unit, a structured light recovery unit, a skeleton reconstruction unit and an optimization unit; the output module includes motion capture data, three-dimensional scene structure and visual display; the optimization unit comprehensively considers the image matching error and the skeleton constraint factors, and realizes the calibration of the motion data through a non-linear optimization algorithm. The non-linear optimization algorithm includes the following steps: Step M1, initial joint positions and rotation angles are generated by the skeleton reconstruction unit as the initial values for optimization. The projection matrix is calculated using the initial pose and calibration parameters of the camera to provide preliminary constraints for optimization; Step M2, in each iteration, calculate the error function and its gradient at the current parameters, and update the joint positions and rotation angles according to the error change: , where is the error function and is the Jacobian matrix of , is the damping factor for adjusting the step size, and the updated parameters are: , where is the three-dimensional position of the updated joint point, is the three-dimensional position of the joint point before update, is the update amount of the three-dimensional position of the joint point, is the new rotation angle of the updated joint, is the rotation angle of the joint before update, is the update amount of the joint angle; Step M3, stop the optimization process when any of the following conditions is met: 1. The parameter update amounts and are lower than the preset threshold; 2. The descent rate of the error function

[0005] is lower than the preset threshold; 3. The preset maximum number of iterations is reached.As a further aspect of the present invention, the calibration system uses a high-performance motion camera with a frame rate of 60fps and a high-definition resolution of 720p, which can provide sufficient image capture capabilities in fast-action scenarios. The wide-angle lens of the camera extends the capture range of a single camera, reducing the problem of information loss due to large movement amplitudes or perspective changes. In addition, the camera is small in size and light in weight (e.g., 94 grams), making it easy to be directly fixed to the joint parts of the user's body without affecting movement flexibility. Each camera is fixed in areas such as the user's waist, chest, thighs, and calves. The cameras installed on the legs face outward to reduce the influence of limb occlusion during movement. In addition, in areas where occlusion or perspective changes are likely to occur during the user's movement (such as the waist or chest), multiple cameras are additionally installed to provide redundant data, and multi-view information is integrated through subsequent algorithms.

[0006] As a further aspect of the present invention, the fixing structure ensures the stability of the cameras installed on the user's body during dynamic capture while not hindering the normal activities of the user. The fixing structure uses Velcro, elastic straps, or other lightweight and adjustable fixing devices to firmly install the cameras on the user's body.

[0007] As a further aspect of the present invention, the time synchronization unit ensures data consistency among multiple cameras during limb movement capture. A time synchronization unit based on audio signals is used as the main synchronization method. The clapboard signal generates a clear and high-intensity sound for marking the starting time point of multi-camera recording. The clapboard sound is recorded by the built-in microphones of each camera. By analyzing the peak positions in these audio signals, the time offset of each camera can be accurately determined, and time alignment processing is performed on the video frames. For the audio and video data recorded by each camera, the time synchronization unit combines the audio analysis results to timestamp the captured video frames, thereby mapping each frame of data to a unified timeline.

[0008] As a further aspect of the present invention, the calibration unit is responsible for calibrating the internal and external parameters of all cameras. The internal parameters include focal length, principal point, and distortion coefficient, and the external parameters include position and attitude. Before use, all cameras need to perform internal parameter calibration based on the fisheye lens distortion model. A fixed and wide-angle fisheye lens is used. By pre-shooting a set of checkerboard pattern images with known geometric features, the focal length and principal point of the camera can be estimated using these checkerboard pattern images. Assuming the internal parameter matrix of the camera is , its form is: , where is the focal length of the camera in the axis direction, and is the camera in the The focal length in the axis direction, is the coordinate of the camera principal point in the axis direction, is the coordinate of the camera principal point in the axis direction. Using a polynomial distortion model combined with the internal parameter matrix , the distortion coefficients are corrected.

[0009] As a further aspect of the present invention, the structured light recovery unit is responsible for extracting key points from video data and reconstructing the three-dimensional trajectory of the camera and the sparse scene point cloud. Before extracting the key points, a set of reference images are taken, and the reference images are panoramic photos of the captured environment. The reference images are used to initialize the global camera position, thereby reducing drift errors. Feature points are extracted from the reference images using the Scale-Invariant Feature Transform (SIFT) algorithm. The calculation formula of the SIFT algorithm is as follows: Assume there are two reference images and , , where is the position of the th feature point in image , is the Scale-Invariant Feature Transform; and the fundamental matrix is estimated by the RANSAC method to obtain geometrically consistent matching point pairs; then, triangulation is performed by the Direct Linear Transformation (DLT) algorithm to convert the matching point pairs into points in 3D space, and a sparse point cloud of the scene is constructed.

[0010] An iterative absolute and relative camera registration method is adopted to achieve high-density reconstruction of the camera position, and the accuracy can be guaranteed even in the case of large view angle changes. The iterative absolute and relative camera registration method includes the following steps: Step Y1, select a pair of images with a sufficient number of matching points and that cannot be explained by a homography matrix as the initial input. Extract the feature points of this pair of images using the Scale-Invariant Feature Transform, find the matching points, estimate the fundamental matrix using the RANSAC method, calculate the essential matrix through the camera internal parameter matrix , decompose the rotation matrix and translation vector from the essential matrix, and use two-view bundle adjustment to optimize the result and minimize the reprojection error; Step Y2, as more images are added, gradually increase the number of registered cameras. For each new image , find the matching points between it and the existing registered images, and estimate the pose of the new camera through the PnP algorithm: , where is the rotation matrix decomposed from the essential matrix, is the translation vector decomposed from the essential matrix, is the image the position of the th feature point, which is a 3D point obtained by triangulation. The poses of all registered cameras and the positions of the 3D points are optimized again using bundle adjustment; Step Y3, when the perspective changes greatly or the background texture is insufficient, the camera cannot successfully reconstruct through the absolute registration method. A relative camera registration method is introduced. For the cameras that have been successfully registered , find the 2D-2D matching points between it and the unregistered camera , and estimate the pose of the unregistered camera through the PnP algorithm: , optimize the poses of all registered cameras and the positions of the 3D points again using bundle adjustment to ensure overall consistency; continuously iterate the above steps until all cameras are densely reconstructed; after each iteration, evaluate the pose estimation accuracy of the unregistered cameras. If the preset threshold is met, include it in the set of registered cameras.

[0011] As a further solution of the present invention, the bone reconstruction unit generates a personalized bone model through the user's range of motion test, associates the bone model with the camera data, and provides an accurate benchmark for motion capture. The workflow of the bone reconstruction unit starts from the user's range of motion test. The user needs to perform a series of predefined full-body actions to cover the complete range of motion of each joint. The actions include basic postures such as raising the hand, squatting, and turning around. These motions are captured by cameras installed on various parts of the user's body, motion feature points are extracted, and the same points under multiple camera views are found through a matching algorithm. Using the internal and external parameters of the camera, the feature points are located in three-dimensional space through triangulation, the three-dimensional positions of the joints are calculated, and then through the hierarchical relationship of the bone model, the direction vectors between adjacent bones are calculated, and the rotation angles of the joints are determined using the included angles of the direction vectors. The bone reconstruction unit also combines multi-view data fusion and three-dimensional reconstruction techniques to extract consistent bone parameters from the capture results of multiple cameras.

[0012] The generation of the bone model depends on the kinematic parameterization method. Using a predefined joint hierarchy, the joint positions of the user are associated with the root node and child nodes of the bone. The root node is the waist, and the child nodes are the knees and elbows. The degrees of freedom (DOF) of each joint are also defined during the reconstruction process. For example, the single-degree-of-freedom rotation of the knee joint or the three-degree-of-freedom spherical motion of the shoulder joint. Through the parameterization method, the bone model can accurately describe the user's anatomical structure and motion characteristics.

[0013] As a further aspect of the present invention, the optimization unit globally optimizes the initially reconstructed skeletal and motion data to improve the accuracy and stability of limb motion capture. The optimization unit comprehensively considers factors such as image matching errors and skeletal constraints, and achieves high-precision calibration of motion data through a non-linear optimization algorithm. First, the joint positions and angles are calibrated by minimizing the image reprojection error to ensure that the reconstructed skeletal model is consistent with the camera image data. The reprojection error refers to the deviation between the two-dimensional point obtained by projecting a three-dimensional joint point onto the image plane in the camera coordinate system and the actual image observation point. The error is expressed by the formula: , where is the error, is the number of cameras, is the number of joint points, is the th three-dimensional position of the joint point, is the th two-dimensional position of the th joint point actually observed on the image plane by the th camera, is the projection function of the th camera, which is used to project the three-dimensional point

[0014] Secondly, constraints on the invariance of bone length and the range of joint rotation angles are introduced into the human kinematic model to eliminate abnormal data that does not conform to the natural laws of human motion. The invariance of bone length means that the distance between adjacent joints remains constant during motion. The constraint on the invariance of bone length is used to correct the deviation of joint positions caused by camera errors or incomplete data. The calculation formula for optimizing the constraint is: , and are the three-dimensional coordinates of two adjacent joints, is the bone length between two adjacent joints, is the Euclidean distance between two joints; the range of rotation angles of each joint is restricted by the physiological structure. The constraint on the range of joint rotation angles is: , where is the rotation angle of joint j, is the minimum allowable rotation angle of joint j, is the maximum allowable rotation angle of joint j.

[0015] As a further aspect of the present invention, the output module processes, formats the optimized limb motion capture data, and provides outputs in various forms to meet the application requirements of different scenarios. The output module includes motion capture data output, three-dimensional scene structure output, and visual display. The motion capture data output outputs the generated joint position and joint angle data in a unified format; the three-dimensional scene structure output provides the user with environment information related to the action, and builds a basic three-dimensional model of the motion capture scene through the sparse point cloud and camera trajectory data generated by multi-view cameras; the visual display integrates the rendering engine to integrate the skeleton model, joint motion trajectory, and sparse scene point cloud into an interactive three-dimensional view.

[0016] Technical effects and advantages of a calibration system for limb motion capture according to the present invention: Through cameras, multi-view data fusion technology, and non-linear optimization algorithms, the present invention achieves high-precision motion capture and calibration, overcomes the defects of traditional optical and inertial capture systems such as dependence on laboratory environments, occlusion problems, and drift errors. Its hardware module design is lightweight, supports portable installation, and provides a stable time synchronization function to ensure data consistency in dynamic scenarios; the data processing module uses structured light recovery, skeleton reconstruction, and global optimization technologies to accurately extract joint positions and joint angles, and at the same time introduces human kinematic constraints to exclude abnormal data; the output module supports formatted output of motion data, three-dimensional scene modeling, and visual display. Compared with the prior art, the present invention improves the portability, adaptability, and accuracy of the motion capture system. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic structural diagram of a calibration system for limb motion capture according to the present invention.

[0018] Figure 2 It is a distribution diagram of reflective balls on the human body based on an optical capture system in the prior art.

[0019] Figure 3 It is a schematic diagram of sensor wearing based on an inertial measurement system in the prior art.

[0020] Figure 4 It is a schematic diagram of the camera of a calibration system for limb motion capture according to the present invention fixed on the user's body area.

[0021] Figure 5 It is a three-dimensional structure diagram after the absolute camera registration method of the present invention is iterated.

[0022] Figure 6 It is a three-dimensional structure diagram after the relative camera registration method of the present invention is iterated. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] Embodiment 1 Referring to Figure 1 the structural schematic diagram shown, the embodiment of the present invention provides a calibration system for limb motion capture, which includes a hardware module, a data processing module, and an output module. The hardware module includes a camera, a fixing structure, and a time synchronization unit. The data processing module includes a calibration unit, a structured light recovery unit, a bone reconstruction unit, and an optimization unit; the output module includes motion capture data, a three-dimensional scene structure, and a visual display.

[0025] In this embodiment, referring to Figure 2 the distribution diagram of reflective spheres on the human body shown, it is a limb motion capture system based on an optical capture system in the prior art. An optical sensor is used to capture the motion trajectory and posture of the user, and the user's actions are analyzed and evaluated in real time through the Unity3D platform. However, it performs poorly in natural light, outdoor scenes, and large-scale capture; referring to Figure 3 the schematic diagram of sensor wearing shown, it is a limb motion capture system based on an inertial measurement system in the prior art, but there is still a certain gap in the measurement accuracy of the designed motion capture device.

[0026] Furthermore, the calibration system uses a high-performance motion camera, which has a frame rate of 60fps and a high-definition resolution of 720p, and can provide sufficient image capture capabilities in fast-action scenes. The wide-angle lens of the camera expands the capture range of a single camera, reducing the problem of information loss caused by large motion amplitudes or perspective changes. In addition, the camera is small in size and light in weight (such as 94 grams), which is convenient to be directly fixed on the joint parts of the user's body without affecting the motion flexibility. Referring to Figure 4 the schematic diagram shown, each camera is fixed in areas such as the waist, chest, thighs, and calves of the user's body. The cameras installed on the legs face outward to reduce the influence of limb occlusion during movement. In addition, in areas where occlusion or perspective changes are likely to occur during the user's actions (such as the waist or chest), multiple additional cameras are installed to provide redundant data, and multi-view information is integrated through subsequent algorithms.

[0027] Furthermore, the fixing structure ensures the stability of the camera mounted on the user's body during dynamic capture, while not interfering with the user's normal activities. The fixing structure uses Velcro, elastic straps or other lightweight and adjustable fixing devices to firmly mount the camera on the user's body.

[0028] Furthermore, the time synchronization unit ensures data consistency among multiple cameras during limb motion capture. It uses an audio signal-based time synchronization unit as the main synchronization method. The clap signal generates a clear and high-intensity sound for marking the starting time point of multi-camera recording. The clap sound is recorded by the built-in microphone of each camera. By analyzing the peak positions in these audio signals, the time offset of each camera can be accurately determined, and the video frames can be time-aligned. For the audio and video data recorded by each camera, the time synchronization unit combines the audio analysis results to timestamp the captured video frames, thereby mapping each frame of data to a unified timeline.

[0029] Furthermore, the calibration unit is responsible for calibrating the internal and external parameters of all cameras. The internal parameters include focal length, principal point, and distortion coefficient, and the external parameters include position and attitude. Before use, all cameras need to be internally calibrated based on the fisheye lens distortion model. Using a fixed and wide-angle fisheye lens, by pre-shooting a set of checkerboard pattern images with known geometric features, the focal length and principal point of the camera can be estimated using these checkerboard pattern images. Assume the internal parameter matrix of the camera is in the form of: , where is the focal length of the camera in the axis direction, is the focal length of the camera in the axis direction, is the coordinate of the camera's principal point in the axis direction, is the coordinate of the camera's principal point in the axis direction. The polynomial distortion model is combined with the internal parameter matrix to correct the distortion coefficient.

[0030] Furthermore, the structured light recovery unit is responsible for extracting key points from the video data and reconstructing the 3D trajectory of the camera and the sparse scene point cloud. Before extracting the key points, a set of reference images are taken. The reference images are panoramic photos of the capture environment. The reference images are used to initialize the global camera position, thereby reducing the drift error. The scale-invariant feature transform algorithm is used to extract feature points from the reference images. The calculation formula of the scale-invariant feature transform algorithm is: Assume there are two reference images and , wherein, is the position of the th feature point in the image ; is Scale - Invariant Feature Transform (SIFT); and the fundamental matrix is estimated by the RANSAC method to obtain geometrically consistent matching point pairs; then, triangulation is performed by the Direct Linear Transformation (DLT) algorithm to convert the matching point pairs into points in 3D space, constructing a sparse point cloud of the scene.

[0031] An iterative absolute and relative camera registration method is adopted to achieve high - density reconstruction of the camera position, ensuring accuracy even in the case of large view changes. The iterative absolute and relative camera registration method includes the following steps: Step Y1, select a pair of images with a sufficient number of matching points that cannot be explained by a homography matrix as the initial input. Extract the feature points of this pair of images by Scale - Invariant Feature Transform (SIFT), find the matching points, estimate the fundamental matrix using the RANSAC method, calculate the essential matrix through the camera intrinsic parameter matrix , decompose the rotation matrix and translation vector from the essential matrix, and use two - view bundle adjustment to optimize the result and minimize the reprojection error; Step Y2, as more images are added, gradually increase the number of registered cameras. For each new image , find the matching points between it and the existing registered images, and estimate the pose of the new camera through the PnP algorithm: , wherein, is the rotation matrix decomposed from the essential matrix, is the translation vector decomposed from the essential matrix, is the position of the th feature point in the image ; is the 3D point obtained by triangulation. Again, use bundle adjustment to optimize the poses of all registered cameras and the positions of the 3D points; Step Y3, when the view changes greatly or the background texture is insufficient, the camera cannot be successfully reconstructed by the absolute registration method. Introduce the relative camera registration method. For the cameras that have been successfully registered , find the 2D - 2D matching points between it and the unregistered camera , and estimate the pose of the unregistered camera through the PnP algorithm: , use bundle adjustment again to optimize the poses of all registered cameras and the positions of 3D points to ensure overall consistency; continuously iterate the above steps until high-density reconstructions are obtained for all cameras; after each iteration, evaluate the pose estimation accuracy of unregistered cameras, and if the preset threshold is met, include them in the set of registered cameras.

[0032] Furthermore, the bone reconstruction unit generates a personalized bone model through the user's range of motion test and associates the bone model with the camera data to provide an accurate benchmark for motion capture. The workflow of the bone reconstruction unit starts with the user's range of motion test, where the user needs to perform a series of predefined full-body movements to cover the complete range of motion of each joint. The movements include basic poses such as raising the hand, squatting, and turning. These movements are captured by cameras installed on various parts of the user's body, motion feature points are extracted, and the same points are found from multiple camera views through a matching algorithm. Using the internal and external parameters of the cameras, the feature points are located in three-dimensional space through triangulation, the three-dimensional positions of the joints are calculated, and then through the hierarchical relationship of the bone model, the direction vectors between adjacent bones are calculated, and the rotation angles of the joints are determined using the angles between the direction vectors. The bone reconstruction unit also combines multi-view data fusion and three-dimensional reconstruction techniques to extract consistent bone parameters from the capture results of multiple cameras.

[0033] In this embodiment, the bone reconstruction unit mainly includes the following functions: 1. 3D point cloud reconstruction based on multi-view camera data, obtaining human key points from multiple camera views and calculating 3D joint coordinates using the triangulation algorithm; 2. Generation of a personalized bone model, determining parameters such as individual joint lengths and rotation angle limits through the user's range of motion test, thereby generating a bone model that conforms to the user's characteristics. 3. Bone pose optimization, using a non-linear optimization method to optimize the initially generated bone model to make it conform to human kinematic constraints and reduce errors. 4. Fusion with action data, matching the bone model with the captured action data to achieve high-precision limb motion capture. The implementation process of bone reconstruction includes: 1. 3D point cloud reconstruction, extracting human key points from images captured by multi-view cameras through scale-invariant feature transform, and using the random sample consensus method to remove mis-matched points; 2. Suppose the positions of a certain joint point captured by two cameras are and , and their corresponding camera projection matrices are and , then the 3D coordinates of the joint point can be solved by the following equations: , , solved using the direct linear transformation method: , where , the optimal solution is obtained through singular value decomposition.

[0034] The specific steps for the bone reconstruction unit to generate a personalized bone model are as follows: Step K1, collect the user's motion data, including when the user stands with their arms relaxed, collect the 3D coordinates of their standard standing position, dynamic motion range data: upper limb movement, such as the arm being lifted forward, swung backward, and extended laterally; lower limb movement, such as knee bending, single-leg standing, and squatting; spinal movement, such as left and right rotation, and forward and backward bending; special actions, such as finger bending and head rotation.

[0035] Step K2, use the SIFT algorithm to detect the key points of the user's body, such as: the tip of the nose, ears, chin, shoulders, elbows, wrists, fingers, hips, knees, ankles, and toes.

[0036] Step K3, three-dimensional bone reconstruction, set the projection matrix of the first camera: , the projection matrix of the second camera: , where is the camera internal parameter matrix, is the rotation matrix of the first camera, is the rotation matrix of the second camera, is the translation vector of the first camera, is the translation vector of the second camera, and the 3D position of the key point is obtained by solving: , , is the pixel coordinate of the body key point in the 2D image of the first camera, is the pixel coordinate of the body key point in the 2D image of the second camera. The optimal solution is calculated using singular value decomposition .

[0037] Furthermore, the optimization unit globally optimizes the preliminarily reconstructed bones and motion data to improve the accuracy and stability of limb motion capture. The optimization unit comprehensively considers factors such as image matching error and bone constraints, and realizes high-precision calibration of motion data through a non-linear optimization algorithm. First, calibrate the joint positions and angles by minimizing the image reprojection error to ensure that the reconstructed bone model is consistent with the camera image data. The reprojection error refers to the deviation between the 2D point obtained by projecting the 3D joint point onto the image plane in the camera coordinate system and the actual image observation point. The error is expressed by the formula: , where is the error, is the number of cameras, is the number of joint points, is the th three-dimensional position of the joint point, is the The two-dimensional position of the th joint point actually observed by the camera on the image plane, is the projection function of the th camera, which is used to project the three-dimensional point onto the two-dimensional image plane.

[0038] Secondly, the kinematic model of the human body is introduced with the constraints of bone length invariance and joint rotation angle range to eliminate abnormal data that does not conform to the natural movement law of the human body. The bone length invariance means that the distance between adjacent joints remains constant during movement. The bone length invariance constraint is used to correct the joint position deviation caused by camera error or incomplete data. The calculation formula for optimizing the constraint is: , and are the three-dimensional coordinates of two adjacent joints, is the bone length between two adjacent joints, is the Euclidean distance between two joints; the rotation angle range of each joint is restricted by the physiological structure. The joint rotation angle range constraint is: , where is the rotation angle of joint j, is the minimum allowable rotation angle of joint j, is the maximum allowable rotation angle of joint j.

[0039] The non-linear optimization algorithm is adopted to iteratively adjust the bone parameters, which include joint positions and rotation angles, and gradually reduce the error function . The non-linear optimization algorithm includes the following steps: Step M1, the initial joint positions and rotation angles are generated by the bone reconstruction unit as the initial values for optimization. The projection matrix is calculated using the initial pose and calibration parameters of the camera to provide preliminary constraints for optimization; Step M2, in each iteration, calculate the error function and its gradient at the current parameters, and update the joint positions and rotation angles according to the error change: , where is the Jacobian matrix of the error function and , is the damping factor for adjusting the step size. The updated parameters are: , where is the three-dimensional position of the updated joint point, is the three-dimensional position of the joint point before update, is the update amount of the three-dimensional position of the joint point, is the new rotation angle of the updated joint, is the rotation angle of the joint before update, is the update amount of the joint angle; Step M3, the optimization process stops when any of the following conditions is met: 1. The parameter update amount and is lower than the preset threshold; 2. The descent rate of the error function is lower than the preset threshold; 3. The preset maximum number of iterations is reached.

[0040] Furthermore, the output module processes, formats the optimized limb motion capture data, and provides outputs in various forms to meet the application requirements of different scenarios. The output module includes motion capture data output, three-dimensional scene structure output, and visualization display. The motion capture data output outputs the generated joint position and joint angle data in a unified format; the three-dimensional scene structure output provides the user with environment information related to the action, and builds a basic three-dimensional model of the motion capture scene through the sparse point cloud and camera trajectory data generated by multi-view cameras; the visualization display integrates the rendering engine to integrate the bone model, joint point motion trajectory, and sparse scene point cloud into an interactive three-dimensional view.

[0041] The present invention realizes high-precision motion capture and calibration through cameras, multi-view data fusion technology, and non-linear optimization algorithms, overcomes the defects of traditional optical and inertial capture systems such as dependence on laboratory environment, occlusion problems, and drift errors. Its hardware module design is lightweight, supports portable installation and provides a stable time synchronization function to ensure data consistency in dynamic scenarios; the data processing module adopts structured light recovery, bone reconstruction, and global optimization technology, which can accurately extract joint positions and joint angles, and at the same time introduces human kinematic constraints to exclude abnormal data; the output module supports formatted output of motion data, three-dimensional scene modeling, and visualization display. Compared with the prior art, the present invention improves the portability, adaptability, and accuracy of the motion capture system.

[0042] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0043] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A calibration system for body motion capture, characterized in that: It includes a hardware module, a data processing module and an output module; the hardware module includes a camera, a fixed structure and a time synchronization unit, the data processing module includes a calibration unit, a structured light recovery unit, a bone reconstruction unit and an optimization unit; the output module includes motion capture data, a three-dimensional scene structure and a visual display; the optimization unit comprehensively considers image matching errors and bone constraint factors, and realizes the calibration of motion data through a nonlinear optimization algorithm, and the nonlinear optimization algorithm includes the following steps: Step M1, initial joint position and rotation angle Generated by the skeleton reconstruction unit as the initial value of optimization, the projection matrix is ​​calculated using the initial pose of the camera and the calibration parameters to provide preliminary constraints for the optimization; Step M2: In each iteration, calculate the error function under the current parameters And its gradient, update the joint position and rotation angle according to the error change: ,in, is the error function and The Jacobian matrix of To adjust the damping factor of the step size, the updated parameters are: ,in, is the updated 3D position of the joint point, is the three-dimensional position of the joint point before updating, is the update amount of the three-dimensional position of the joint point, is the new rotation angle of the joint after update, To update the rotation angle of the front joint, is the update amount of the joint angle; Step M3, the optimization process stops when any of the following conditions are met:

1. Parameter update amount and Below the preset threshold; 2. Error function The rate of decrease is lower than the preset threshold; 3. The preset maximum number of iterations is reached.

2. A body motion capture calibration system according to claim 1, characterized in that The structured light recovery unit is responsible for extracting key points from the video data and reconstructing the three-dimensional trajectory of the camera and the sparse scene point cloud. Before extracting the key points, a set of reference images is taken, and the reference images are panoramic photos of the captured environment; the reference images are used to initialize the global camera position, thereby reducing the drift error; the feature points are extracted from the reference images using the scale-invariant feature transformation algorithm, and the calculation formula of the scale-invariant feature transformation algorithm is: Assuming there are two reference images and , ,in, For images Middle The location of the feature points, It is a scale-invariant feature transformation; and the basic matrix is ​​estimated through the RANSAC method to obtain geometrically consistent matching point pairs; then, the matching point pairs are triangulated through the direct linear transformation algorithm to convert them into points in 3D space and construct a sparse point cloud of the scene.

3. A body motion capture calibration system according to claim 1, characterized in that The structured light recovery unit adopts an iterative absolute and relative camera registration method to achieve high-density reconstruction of the camera position, and can ensure accuracy even when the viewing angle changes greatly. The iterative absolute and relative camera registration method includes the following steps: Step Y1, select a pair of images with matching points that cannot be explained by the homography matrix as the initial input, extract the feature points of the pair of images through scale-invariant feature transformation, find the matching points, estimate the basic matrix using the RANSAC method, and use the camera intrinsic parameter matrix Calculate the essential matrix, decompose the rotation matrix and translation vector from the essential matrix, and use two-view bundle adjustment to optimize the result and minimize the reprojection error; Step Y2: As more images are added, the number of registered cameras is gradually increased. For each new image , find the matching points between it and the existing registered image, and estimate the pose of the new camera through the PnP algorithm: ,in, is the rotation matrix decomposed from the essential matrix, is the translation vector decomposed from the essential matrix, For images Middle The location of the feature points, For the 3D points obtained by triangulation, bundle adjustment is used again to optimize the poses of all registered cameras and the positions of 3D points; Step Y3: When the viewing angle changes greatly or the background texture is insufficient, the camera cannot be successfully reconstructed through the absolute registration method. The relative camera registration method is introduced to reconstruct the camera that has been successfully registered. , find the camera that is not registered 2D-2D matching points between , estimate the pose of the unregistered camera using the PnP algorithm: , use bundle adjustment again to optimize the poses and positions of 3D points of all registered cameras to ensure overall consistency; continue to iterate the above steps until all cameras are reconstructed with high density; after each iteration, evaluate the pose estimation accuracy of the unregistered camera, and if it meets the preset threshold, it will be included in the set of registered cameras.

4. A body motion capture calibration system according to claim 1, characterized in that: The optimization unit globally optimizes the preliminarily reconstructed skeleton and motion data to improve the accuracy and stability of limb motion capture; the optimization unit comprehensively considers image matching errors and skeleton constraint factors, and realizes high-precision calibration of motion data through a nonlinear optimization algorithm; first, the joint position and angle are calibrated by minimizing the image reprojection error to ensure that the reconstructed skeleton model is consistent with the camera image data. The reprojection error refers to the deviation between the two-dimensional point after the three-dimensional joint point is projected onto the image plane in the camera coordinate system and the actual image observation point. The error is expressed by the formula: ,in, is the error, is the number of cameras, is the number of joint points, For the The three-dimensional position of the joint points, For the The first camera actually observes the The two-dimensional position of the joint points, For the The projection function of the camera is used to transform the three-dimensional point Projection onto the 2D image plane.

5. A body motion capture calibration system according to claim 1, characterized in that: The optimization unit also introduces bone length invariance constraints and joint rotation angle range constraints to the human kinematics model to eliminate abnormal data that does not conform to the natural movement law of the human body; bone length invariance means that the distance between adjacent joints remains constant during the movement process, and the bone length invariance constraint is used to correct the joint position deviation caused by camera error or incomplete data. The calculation formula of the optimization constraint is: , and are the three-dimensional coordinates of two adjacent joints, is the bone length between two adjacent joints, is the Euclidean distance between two joints; the rotation angle range of each joint is limited by the physiological structure, and the joint rotation angle range is constrained as follows: ,in, is the rotation angle of joint j, is the minimum allowed rotation angle of joint j, is the maximum allowed rotation angle of joint j.

6. A body motion capture calibration system according to claim 1, characterized in that: The calibration unit is responsible for calibrating the intrinsic parameters and extrinsic parameters of all cameras. The intrinsic parameters include focal length, principal point and distortion coefficient, and the extrinsic parameters include position and posture. Before use, the cameras need to be calibrated based on the fisheye lens distortion model. A fixed and wide-angle fisheye lens is used. A set of checkerboard pattern images with known geometric features are pre-shot, and the focal length and principal point of the camera are estimated using the checkerboard pattern images. Assume that the intrinsic parameter matrix of the camera is , which is of the form: ,in, The camera is The focal length in the axial direction, The camera is The focal length in the axial direction, is the camera principal point at The coordinates in the axis direction, is the camera principal point at Coordinates in the axial direction, using a polynomial distortion model combined with an internal parameter matrix , and correct the distortion coefficient.

7. A body motion capture calibration system according to claim 1, characterized in that: The calibration system uses a high-performance sports camera with a frame rate of 60fps and a resolution of 720p. The camera's wide-angle lens expands the capture range. The camera is small in size and light in weight, making it easy to fix on the user's body joints. Each camera is fixed on the user's waist, chest, thigh and calf areas, and the camera installed on the leg faces outward.

8. The body motion capture calibration system according to claim 1, characterized in that: The fixing structure uses Velcro, elastic straps or other lightweight, adjustable fixing devices to firmly mount the camera on the user's body.

9. A body motion capture calibration system according to claim 1, characterized in that: The time synchronization unit adopts a time synchronization method based on audio signals, marks the starting time point through a clapper signal, analyzes the peak value of the audio signal to determine the time offset, performs time alignment processing on the video frames and marks the timestamps, and maps each frame of data to a unified time axis.

10. A body motion capture calibration system according to claim 1, characterized in that: The motion capture data output of the output module outputs the joint point positions and joint angle data in a unified format.