Motion capture system, method and storage medium
Through the multimodal sensor system, the accuracy and robustness of optical motion capture when applied to humans or animals is solved, and high-precision and robust motion capture effect are achieved.
Patent Information
- Application Number
- CN202510033678.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-13
AI Technical Summary
When optical motion capture technology is applied to humans or animals, it causes inaccurate and unruly capture due to movement or temperature changes of skin, hair, etc.
A multimodal sensor system is adopted, including inertial sensors, non-visible image sensors and visible image sensors, and the accuracy and robustness of motion capture are improved through the fusion of perceptual data of different modes.
High-precision motion capture of the measured target is achieved, the robustness of the system is enhanced, and it can work effectively under different lighting conditions.
Smart Images

Figure CN119996794A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of motion capture technology, and in particular, relates to a motion capture system, method and storage medium. Background Art
[0002] Optical motion capture mainly relies on optical devices such as near-infrared motion capture cameras to monitor and track specific markers installed on the target. When a marker is captured by two or more cameras at the same time, the three-dimensional spatial position of the point can be determined. Through continuous shooting, the movement trajectory of the marker over time can be tracked. In traditional optical motion capture systems, multiple near-infrared cameras are arranged around the target, and their overlapping fields of view constitute the activity range of the target. In order to facilitate identification and processing, special luminous or reflective markers are usually affixed to the surface of the target. However, when optical motion capture technology is applied to deformable bodies such as humans and animals, the movement of the skin, hair, etc. on the surface of the object or the change of temperature may cause the optical motion capture to be inaccurate and non-robust. Summary of the invention
[0003] The present application aims to solve at least one of the technical problems existing in the related art. To this end, the present application proposes a motion capture system, method and storage medium, which obtains motion data of a measured target based on perception data of different modalities, thereby improving the accuracy and robustness of the motion capture system for capturing motion of the measured target.
[0004] In a first aspect, the present application provides a motion capture system, the system comprising: a motion capture module, comprising sensors of two or more modalities, for sensing a target to be sensed and obtaining sensing data of different modalities;
[0005] A processing unit, connected to the motion capture module, for processing the perception data of the different modes to obtain the motion data of the measured target;
[0006] The motion capture module includes at least one inertial sensor, which is arranged on the measured object and is used to measure inertial data of the measured object during movement, and output the inertial data to the processing unit.
[0007] In the above technical solution, the motion capture system includes a motion capture module and a processing unit. The motion capture module includes sensors of two or more modes, which are used to sense the target to be measured and obtain perception data of different modes. The processing unit is connected to the motion capture module and is used to process the perception data of different modes to obtain motion data of the target to be measured, thereby realizing motion capture of the target to be measured. The motion capture module includes at least one inertial sensor, which is arranged on the target to be measured to measure the inertial data of the target to be measured during movement. The motion data of the target to be measured is obtained based on the perception data of different modes, thereby improving the accuracy and robustness of the motion capture system for motion capture of the target to be measured.
[0008] According to one embodiment of the present application, the motion capture module further includes:
[0009] at least one non-visible light image sensor, the non-visible light image sensor being used to sense the target in a near-infrared or infrared band and output an image in a near-infrared or infrared band; and / or,
[0010] At least one visible light image sensor, wherein the visible light image sensor is used to sense the target in the visible light band and output a visible light image.
[0011] In the above technical solution, the motion capture module also includes at least one non-visible light image sensor and at least one visible light image sensor, wherein the non-visible light image sensor is used to sense the target to be measured in the near-infrared or infrared band and output images in the near-infrared or infrared band, and the visible light image sensor is used to sense the target to be measured in the visible light band and output visible light images, thereby obtaining motion data of the target to be measured based on perception data of different modalities, thereby improving the accuracy and robustness of the motion capture system in capturing motion of the target to be measured.
[0012] According to an embodiment of the present application, when the motion capture system includes at least one non-visible light image sensor, the motion capture system further includes at least one near-infrared or infrared fill light source for illuminating the target;
[0013] At least one marking point is arranged on the measured target, and under the illumination of the supplementary light source, the marking point has a contrast with the background in the imaging of the non-visible light image sensor.
[0014] In the above technical solution, when the motion capture system includes at least one non-visible light image sensor, the motion capture system includes at least one near-infrared or infrared fill light source for illuminating the target to be measured, and at least one marker point is arranged on the target to be measured. Under the illumination of the fill light source, the marker point has a contrast with the background in the imaging of the non-visible light image sensor, making the marker point more prominent, thereby improving the accuracy of identifying and locating the marker point in the imaging of the non-visible light image sensor, thereby improving the accuracy of the motion capture system in capturing the motion of the target to be measured.
[0015] According to one embodiment of the present application, the motion capture system also includes: a first image processing unit, connected to the non-visible light image sensor, for tracking and spatially positioning the marker points in the near-infrared or infrared band image, and outputting the coordinates of the marker points to the processing unit.
[0016] In the above technical solution, the motion capture system also includes a first image processing unit, which is connected to the non-visible light image sensor and is used to track and spatially locate marker points in images in the near-infrared or infrared bands, and output the coordinates of the marker points to the processing unit, thereby improving the flexibility of the motion capture system.
[0017] According to an embodiment of the present application, when the motion capture system includes at least one visible light image sensor, the motion capture system further includes: a second image processing unit connected to the visible light image sensor, configured to:
[0018] Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points;
[0019] The motion posture of the measured target in the visible light image is estimated to obtain a second posture estimation result of the measured target.
[0020] In the above technical scheme, the second image processing unit is used to perform key point detection on the visible light image, obtain the coordinates and IDs of the key points, estimate the motion posture of the target under measurement in the visible light image, and obtain a second pose estimation result. The second pose estimation result can be used to fuse with the pose estimation result obtained based on the perception data of other modalities, and obtain the motion data of the target under measurement, thereby improving the accuracy and robustness of the motion capture system for motion capture of the target under measurement.
[0021] According to one embodiment of the present application, the processing unit is configured to perform at least one of the following:
[0022] Performing three-dimensional reconstruction of the marker points according to the marker point coordinates;
[0023] Performing ID tracking on the marker point to obtain a marker point ID tracking result;
[0024] Calculate the first pose estimation result of the measured object according to the coordinates of at least three non-collinear landmark points on the measured object;
[0025] Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points;
[0026] Performing motion posture estimation on the target to be measured in the visible light image to obtain a second posture estimation result of the target to be measured;
[0027] According to the inertial data, a motion posture of the measured target is estimated to obtain a third posture estimation result of the measured target;
[0028] Reprojecting the marker point onto the visible light image and comparing it with the ID of the key point to verify the correctness of the marker point ID tracking result or to perform ID identification on the marker point that has lost tracking;
[0029] The first pose estimation result, the second pose estimation result and the third pose estimation result are fused and calculated to obtain motion data of the measured target.
[0030] In the above technical scheme, the processing unit performs three-dimensional reconstruction of the marker points according to the coordinates of the marker points, performs ID tracking on the marker points to obtain the marker point ID tracking results, calculates the first pose estimation result of the measured target according to the coordinates of at least three non-collinear marker points on the measured target, performs key point detection on the visible light image to obtain the coordinates and IDs of the key points, estimates the motion posture of the measured target in the visible light image to obtain the second pose estimation result, estimates the motion posture of the measured target according to the inertial data to obtain the third pose estimation result of the measured target, reprojects the marker points onto the visible light image, and compares them with the IDs of the key points to verify the correctness of the marker point ID tracking results or to perform ID identification on the marker points that have lost tracking, thereby realizing the pose estimation result of the measured target based on the perception data of different modalities, and fusing the pose estimation results obtained based on the perception data of each modality to obtain the motion data of the measured target, thereby improving the accuracy and robustness of the motion capture system for the motion capture of the measured target.
[0031] According to one embodiment of the present application, the processing unit is configured to perform at least one of the following:
[0032] Performing three-dimensional reconstruction of the marker points according to the marker point coordinates;
[0033] Performing ID tracking on the marker point to obtain a marker point ID tracking result;
[0034] Calculate the first pose estimation result of the measured object according to the coordinates of at least three non-collinear landmark points on the measured object;
[0035] According to the inertial data, a motion posture of the measured target is estimated to obtain a third posture estimation result of the measured target;
[0036] Reprojecting the marker point onto the visible light image and comparing it with the ID of the key point to verify the correctness of the marker point ID tracking result or to perform ID identification on the marker point that has lost tracking;
[0037] The first pose estimation result, the second pose estimation result and the third pose estimation result are fused and calculated to obtain motion data of the measured target.
[0038] In the above technical scheme, the processing unit performs three-dimensional reconstruction of the marker points according to the coordinates of the marker points, performs ID tracking on the marker points to obtain the marker point ID tracking results, calculates the first pose estimation result of the measured target according to the coordinates of at least three non-collinear marker points on the measured target, estimates the motion posture of the measured target according to the inertial data to obtain the third pose estimation result of the measured target, reprojects the marker points onto the visible light image, and compares them with the IDs of the key points to verify the correctness of the marker point ID tracking results or to perform ID identification on the marker points that have lost tracking, thereby achieving the pose estimation result of the measured target based on the perception data of different modalities, and fusing the pose estimation results obtained based on the perception data of each modality to obtain the motion data of the measured target, thereby improving the accuracy and robustness of the motion capture system for the motion capture of the measured target.
[0039] In a second aspect, the present application provides a motion capture method, the method comprising:
[0040] Perceive the target to be measured and obtain perception data of different modes;
[0041] Processing the sensing data of the different modes to obtain motion data of the measured target;
[0042] The sensing of the target to be measured to obtain sensing data of different modes includes:
[0043] The inertial data of the measured target during its motion is measured.
[0044] In the above technical scheme, the target to be measured is sensed to obtain perception data of different modalities, including sensing the target to be measured in the near-infrared or infrared band, outputting an image of the near-infrared or infrared band, wherein the image of the near-infrared or infrared band includes at least one marker point, sensing the target to be measured in the visible light band, outputting a visible light image, and measuring inertial data of the target to be measured during movement, processing the perception data of different modalities to obtain motion data of the target to be measured, realizing the fusion of perception data of different modalities, and improving the accuracy and robustness of the motion capture method.
[0045] According to an embodiment of the present application, the sensing of the target to obtain sensing data of different modalities further includes:
[0046] sensing the target in a near-infrared or infrared band, and outputting an image in the near-infrared or infrared band, wherein the image in the near-infrared or infrared band includes at least one marker point; and / or,
[0047] The target is sensed in the visible light band and a visible light image is output.
[0048] In the above technical scheme, the target to be measured is sensed to obtain perception data of different modalities, including sensing the target to be measured in the near-infrared or infrared band, outputting an image in the near-infrared or infrared band, wherein the image in the near-infrared or infrared band includes at least one marker point, sensing the target to be measured in the visible light band, outputting a visible light image, and then processing the perception data of different modalities to obtain motion data of the target to be measured, thereby realizing the fusion of perception data of different modalities.
[0049] According to an embodiment of the present application, the processing of the perception data of different modalities to obtain the motion data of the measured target includes:
[0050] Tracking and spatially locating the marker points in the image of the near-infrared or infrared band to obtain the coordinates of the marker points;
[0051] Performing three-dimensional reconstruction of the marker points according to the marker point coordinates;
[0052] Performing ID tracking on the marker point to obtain a marker point ID tracking result;
[0053] Calculate the first pose estimation result of the measured object according to the coordinates of at least three non-collinear landmark points on the measured object;
[0054] Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points;
[0055] Performing motion posture estimation on the target to be measured in the visible light image to obtain a second posture estimation result of the target to be measured;
[0056] According to the inertial data, a motion posture of the measured target is estimated to obtain a third posture estimation result of the measured target;
[0057] Reprojecting the marker point onto the visible light image and comparing it with the ID of the key point to verify the correctness of the marker point ID tracking result or to perform ID identification on the marker point that has lost tracking;
[0058] The first pose estimation result, the second pose estimation result and the third pose estimation result are fused and calculated to obtain motion data of the measured target.
[0059] In the above technical scheme, the marker points in the image of near-infrared or infrared bands are tracked and spatially positioned to obtain the coordinates of the marker points; the marker points are three-dimensionally reconstructed, the marker points are ID tracked to obtain the marker point ID tracking results, the first pose estimation result of the measured target is calculated according to the coordinates of at least three non-collinear marker points on the measured target, the key point detection is performed on the visible light image to obtain the coordinates and ID of the key points, the motion posture of the measured target in the visible light image is estimated to obtain the second pose estimation result, the motion posture of the measured target is estimated according to the inertial data to obtain the third pose estimation result of the measured target, the marker points are reprojected onto the visible light image, and compared with the ID of the key points to verify the correctness of the marker point ID tracking result or to perform ID identification on the marker points that have lost tracking, so as to realize the pose estimation result of the measured target based on the perception data of different modalities, and to fuse the pose estimation results obtained based on the perception data of each modality to obtain the motion data of the measured target, thereby improving the accuracy and robustness of the motion capture system for the motion capture of the measured target.
[0060] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the motion capture method as described in the second aspect above when executing the computer program.
[0061] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the motion capture method as described in the second aspect above.
[0062] In a fifth aspect, the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the motion capture method as described in the second aspect.
[0063] In a sixth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the motion capture method as described in the second aspect above.
[0064] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0066] Figure 1 is one of the structural schematic diagrams of the motion capture system provided in some embodiments of the present application;
[0067] Figure 2 This is the second structural schematic diagram of the motion capture system provided by some embodiments of the present application;
[0068] Figure 3 This is the third structural schematic diagram of the motion capture system provided by some embodiments of the present application;
[0069] Figure 4 This is a fourth structural diagram of a motion capture system provided by some embodiments of the present application;
[0070] Figure 5 This is the fifth structural diagram of the motion capture system provided by some embodiments of the present application;
[0071] Figure 6 is a flowchart of a motion capture method provided by some embodiments of the present application;
[0072] Figure 7 It is a schematic diagram of the structure of an electronic device provided in some embodiments of the present application.
[0073] Description of reference numerals:
[0074] 101: motion capture module; 102: processing unit; 103: inertial sensor;
[0075] 104: non-visible light image sensor; 105: visible light image sensor; 106: fill light source;
[0076] 107: first image processing unit; 108: second image processing unit; 700: electronic device;
[0077] 701: processor; 702: memory. DETAILED DESCRIPTION
[0078] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.
[0079] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0080] The motion capture system, method and storage medium provided in the embodiments of the present application are described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.
[0081] The motion capture method provided in the embodiment of the present application can be executed by a motion capture system. In the embodiment of the present application, the motion capture system provided in the embodiment of the present application is described by taking the motion capture method executed by the motion capture system as an example.
[0082] Figure 1 FIG. 1 is one of the structural diagrams of the motion capture system provided in some embodiments of the present application. Figure 1 As shown, the motion capture system includes:
[0083] The motion capture module 101 includes sensors of two or more modes, and is used to sense the target to be measured and obtain sensing data of different modes;
[0084] A processing unit 102, connected to the motion capture module 101, is used to process the perception data of different modes to obtain motion data of the measured target;
[0085] The motion capture module 101 includes at least one inertial sensor 103;
[0086] The inertial sensor 103 is disposed on the measured object, and is used to measure inertial data of the measured object during movement, and output the inertial data to the processing unit.
[0087] It should be noted that the target to be measured in the embodiments of the present application may be one or more targets, and the present application does not impose any limitation on this.
[0088] It is understandable that the motion capture module 101 includes sensors of two or more modalities, and these sensors can be of different types, for example: non-visible light image sensors, visible light image sensors or inertial sensors, etc. Each modality of sensor can sense the target to be measured to obtain perception data of a specific modality, such as a near-infrared band image sensor senses the target to be measured to obtain a near-infrared band image, a visible light image sensor senses the target to be measured to obtain a visible light image, and an inertial sensor measures the inertial data of the target to be measured during movement, etc. Through sensors of two or more modalities, the motion capture module 101 can sense the target to be measured from multiple modalities, thereby improving the comprehensiveness and accuracy of the perception of the target to be measured. Optionally, the Zhang Zhengyou calibration method is used to calibrate the sensors of each modality.
[0089] Figure 2 This is the second structural diagram of the motion capture system provided by some embodiments of the present application. Figure 2 As shown, in some embodiments, the motion capture module 101 further includes:
[0090] At least one non-visible light image sensor 104, the non-visible light image sensor 104 is used to sense the target in the near infrared or infrared band and output an image in the near infrared or infrared band; and / or,
[0091] At least one visible light image sensor 105, wherein the visible light image sensor 105 is used to sense the target in the visible light band and output a visible light image.
[0092] It can be understood that the motion capture module 101 includes at least one inertial sensor 103 and at least one non-visible light image sensor 104;
[0093] Alternatively, it includes at least one inertial sensor 103 and at least one visible light image sensor 105;
[0094] Alternatively, it includes at least one inertial sensor 103 , at least one non-visible light image sensor 104 and at least one visible light image sensor 105 .
[0095] It can be understood that the non-visible light image sensor 104 is used to sense the target to be measured in the near-infrared or infrared band and output an image in the near-infrared or infrared band. Since the light in the near-infrared and infrared bands is invisible to the human eye, the non-visible light image sensor is not limited by visible light conditions and is more suitable for environments with undesirable visible light conditions. The motion capture module includes at least one non-visible light image sensor 104, which improves the applicability of the motion capture system under different lighting conditions; the visible light image sensor 105 is used to sense the target to be measured in the visible light band and output a visible light image. Under good lighting conditions, the visible light image sensor 105 can output a high-resolution color image of the target to be measured; the inertial sensor 103 is set on the target to be measured to measure the inertial data of the target to be measured during movement, including data reflecting the motion state and direction change, such as three-axis angular velocity, three-axis acceleration, three-axis magnetometer data, etc. By fusing the inertial data, the anti-occlusion robustness of the motion solution of the target to be measured can be improved.
[0096] Optionally, the inertial sensor disposed on the target to be measured is bound to the target to be measured through device identification information such as a serial number (SN).
[0097] In some embodiments, the motion capture module 101 includes at least one non-visible light image sensor 104, at least one visible light image sensor 105 and at least one inertial sensor 103, so that the motion capture module 101 can perceive the target to be measured from optical, visual and inertial modalities. The optical modality corresponds to the near-infrared or infrared band image output by the non-visible light image sensor 104, the visual modality corresponds to the visible light image output by the visible light image sensor 105, and the inertial modality corresponds to the inertial data output by the inertial sensor, which improves the comprehensiveness and accuracy of the perception of the target to be measured, obtains motion data based on the perception data of different modalities, and also improves the accuracy of the motion capture system in capturing the motion of the target to be measured.
[0098] The processing unit 102 is used to process the perception data of different modes, such as data matching, fusion and analysis, to obtain the motion data of the measured target. The motion data of the measured target describes the motion state of the measured target in three-dimensional space, including the motion posture data of the measured object, so as to realize the motion capture of the measured target. It can be understood that the motion data is obtained based on the perception data of different modes, and the fusion of the perception data of different modes is realized, which improves the accuracy of the motion capture system in capturing the motion of the measured target. The sensor of a single mode may be affected by occlusion or other factors, resulting in the output perception data being inaccurate. Fusion of the perception data of sensors of multiple modes can improve the robustness of the motion capture system.
[0099] In the above technical solution, the motion capture system includes a motion capture module and a processing unit. The motion capture module includes sensors of two or more modes, which are used to sense the target to be measured and obtain perception data of different modes. The processing unit is connected to the motion capture module and is used to process the perception data of different modes to obtain motion data of the target to be measured, thereby realizing motion capture of the target to be measured. The motion capture module includes at least one inertial sensor, which is arranged on the target to be measured to measure the inertial data of the target to be measured during movement. The motion data of the target to be measured is obtained based on the perception data of different modes, thereby improving the accuracy and robustness of the motion capture system for motion capture of the target to be measured.
[0100] Figure 3 This is the third structural diagram of the motion capture system provided by some embodiments of the present application. Figure 3 As shown, in the case where the motion capture system includes at least one non-visible light image sensor 104, the motion capture system also includes at least one near-infrared or infrared band fill light source 106 for illuminating the target;
[0101] At least one marking point is arranged on the measured object. Under the illumination of the supplementary light source 106 , the marking point has a contrast with the background in the imaging of the non-visible light image sensor 104 .
[0102] The near-infrared or infrared band fill light source 106 is used to provide near-infrared or infrared band light to illuminate the target to be measured, so that the non-visible light image sensor 104 can perceive the target to be measured more clearly, thereby improving the quality of the near-infrared or infrared band image output by the non-visible light image sensor 104, and also enabling the non-visible light image sensor 104 to still perceive the target to be measured when the near-infrared or infrared band lighting conditions are not ideal.
[0103] At least one mark point is arranged on the target to be measured. The mark point is a special mark or light-emitting point, also called a marker. Under the illumination of the fill light source 106, the mark point will form a contrast with the background in the imaging of the non-visible light image sensor, making the mark point more prominent in the image of the near-infrared or infrared band output by the non-visible light image sensor 104, and forming a contrast with the surrounding background, making the mark point easier to identify and track in subsequent image processing and analysis, thereby improving the accuracy of identifying and locating the mark point.
[0104] In the above technical solution, when the motion capture system includes at least one non-visible light image sensor, the motion capture system includes at least one near-infrared or infrared fill light source for illuminating the target to be measured, and at least one marker point is arranged on the target to be measured. Under the illumination of the fill light source, the marker point has a contrast with the background in the imaging of the non-visible light image sensor, making the marker point more prominent, thereby improving the accuracy of identifying and locating the marker point in the imaging of the non-visible light image sensor, thereby improving the accuracy of the motion capture system in capturing the motion of the target to be measured.
[0105] Figure 4 FIG. 4 is a schematic diagram of the structure of the motion capture system provided in some embodiments of the present application. Figure 4 As shown, the motion capture system also includes: a first image processing unit 107, which is connected to the non-visible light image sensor 104, and is used to track and spatially locate the marker points in the near-infrared or infrared band image, and output the coordinates of the marker points to the processing unit 102.
[0106] In some embodiments, the motion capture module 101 in the motion capture system includes at least one non-visible light image sensor 104, at least one visible light image sensor 105, and at least one inertial sensor 103, which are used to sense the target to be measured and obtain near-infrared or infrared band images, visible light images, and inertial data. In some embodiments, the processing unit 102 processes the near-infrared or infrared band images, visible light images, and inertial data to obtain motion data of the target to be measured.
[0107] In some embodiments, the motion capture system also includes a first image processing unit 107, which uses image processing technology to track marker points in images in the near-infrared or infrared bands, and obtains the coordinates of the marker points through spatial positioning calculations. The processing unit 102 receives the coordinates of the marker points output by the first image processing unit 107, thereby improving the flexibility of the motion capture system.
[0108] Optionally, the first image processing unit 107 is an embedded processing unit, such as a Field Programmable Gate Array (FPGA), etc., which is embedded in the motion capture module 101 and connected to the non-visible light image sensor 104 .
[0109] In the above technical solution, the motion capture system also includes a first image processing unit, which is connected to the non-visible light image sensor and is used to track and spatially locate marker points in images in the near-infrared or infrared bands, and output the coordinates of the marker points to the processing unit, thereby improving the flexibility of the motion capture system.
[0110] like Figure 4 As shown, in some embodiments, when the motion capture module 101 in the motion capture system includes at least one non-visible light image sensor 104, at least one visible light image sensor 105 and at least one inertial sensor 103, and the motion capture system further includes a first image processing unit 107, the processing unit 102 is used to perform at least one of the following:
[0111] Performing three-dimensional reconstruction of the marker points according to the marker point coordinates;
[0112] Performing ID tracking on the marker point to obtain a marker point ID tracking result;
[0113] Calculate the first pose estimation result of the measured object according to the coordinates of at least three non-collinear landmark points on the measured object;
[0114] Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points;
[0115] Performing motion posture estimation on the target to be measured in the visible light image to obtain a second posture estimation result of the target to be measured;
[0116] According to the inertial data, a motion posture of the measured target is estimated to obtain a third posture estimation result of the measured target;
[0117] Reprojecting the marker point onto the visible light image and comparing it with the ID of the key point to verify the correctness of the marker point ID tracking result or to perform ID identification on the marker point that has lost tracking;
[0118] The first pose estimation result, the second pose estimation result and the third pose estimation result are fused and calculated to obtain motion data of the measured target.
[0119] The motion capture module 101 includes at least one non-visible light image sensor 104, which can sense the object under test from at least one perspective to obtain images in the near-infrared or infrared bands, and the first image processing unit 107 uses image processing technology to track the marker points in these near-infrared or infrared band images, obtain the coordinates of the marker points through spatial positioning calculation and output them to the processing unit 102, that is, the processing unit 102 receives the coordinates of the marker points.
[0120] The processing unit 102 performs multi-view 3D reconstruction on the object to be measured by using a triangulation method according to the coordinates of the marker points to obtain the 3D coordinates of the marker points.
[0121] The processing unit 102 tracks each marker point to determine its position change in continuous image frames or time series, and assigns a unique ID to each marker point and tracks the ID of each marker point to achieve the ability to identify and associate the same marker point during the dynamic capture of the measured target. Optionally, the processing unit 102 uses a method such as Kalman filtering to track the marker point ID to obtain a marker point ID tracking result, which includes the three-dimensional coordinates of each marker point and its corresponding image frame sequence or time series.
[0122] Further, the processing unit 102 calculates a first pose estimation result PO based on the three-dimensional coordinates of at least three non-collinear marker points. The pose estimation result is an estimation of motion pose data, and the motion pose data of the measured target includes position data and pose data of the measured target in three-dimensional space.
[0123] Optionally, the processing unit 102 uses a deep learning algorithm, such as YOLOv8, to perform key point detection on the visible light image to obtain the coordinates and ID of the key point, where the ID is a unique identifier of the key point.
[0124] Optionally, the processing unit 102 uses algorithms such as MediaPipe to estimate the motion posture of the target in the visible light image to obtain a second posture estimation result PA.
[0125] Optionally, the processing unit 102 uses a complementary filtering algorithm to estimate the motion posture of the target according to the inertial data to obtain a third pose estimation result of the target, wherein the spatial positioning is replaced by the displacement vector T obtained in the first pose estimation result PO.
[0126] Furthermore, the processing unit 102 reprojects the marker point onto the visible light image according to the three-dimensional coordinates of the marker point, converts the three-dimensional spatial coordinates of the marker point into two-dimensional reprojection coordinates, and compares the projection of the marker point in the visible light image with the key points detected in the visible light image, including comparing their coordinate positions and IDs to verify the accuracy of the marker point ID tracking results, thereby verifying the correctness of marker ID tracking or performing ID identification on marker points that have lost tracking.
[0127] Optionally, the processing unit 102 uses methods such as Kalman filtering to perform fusion calculations on the first pose estimation result PO, the second pose estimation result PA and the third pose estimation result PB to obtain motion data P of the target under test. The motion data of the target under test describes the motion posture of the target under test in three-dimensional space, thereby realizing motion capture of the target under test.
[0128] Through the above method, the problem of marker ID tracking error or inability to correctly identify the ID can be solved, and the motion posture can still be reliably estimated through visible light images and inertial sensors after the marker is lost, avoiding the complete loss of motion data.
[0129] In the above technical scheme, the processing unit performs three-dimensional reconstruction of the marker points according to the coordinates of the marker points, performs ID tracking on the marker points to obtain the marker point ID tracking results, calculates the first pose estimation result of the measured target according to the coordinates of at least three non-collinear marker points on the measured target, performs key point detection on the visible light image to obtain the coordinates and IDs of the key points, estimates the motion posture of the measured target in the visible light image to obtain the second pose estimation result, estimates the motion posture of the measured target according to the inertial data to obtain the third pose estimation result of the measured target, reprojects the marker points onto the visible light image, and compares them with the IDs of the key points to verify the correctness of the marker point ID tracking results or to perform ID identification on the marker points that have lost tracking, thereby realizing the pose estimation result of the measured target based on the perception data of different modalities, and fusing the pose estimation results obtained based on the perception data of each modality to obtain the motion data of the measured target, thereby improving the accuracy and robustness of the motion capture system for the motion capture of the measured target.
[0130] Figure 5 FIG. 5 is a schematic diagram of the structure of the motion capture system provided in some embodiments of the present application. Figure 5 As shown, in the case where the motion capture system includes at least one visible light image sensor, the motion capture system further includes: a second image processing unit 108, connected to the visible light image sensor 105, for:
[0131] Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points;
[0132] The motion posture of the measured target in the visible light image is estimated to obtain a second posture estimation result of the measured target.
[0133] The second image processing unit 108 is connected to the visible light image sensor 105 and is used to perform one or more of the following processing on the visible light image: feature extraction, feature recognition, feature space positioning calculation, motion data resolution, and send the processing results to the processing unit 102.
[0134] Among them, feature extraction refers to extracting features from visible light images that can be used for subsequent motion capture of the target under test. Features may include edges, corners, textures, etc. in the image; feature recognition refers to the process of further determining the specific identity or category of these features based on feature extraction; feature space positioning calculation refers to estimating the spatial position and posture of the target under test by identifying and using features in the image; motion data solution refers to the process of determining the motion posture data of the target under test. The motion posture data of the target under test includes the position data and posture data of the target under test in three-dimensional space.
[0135] The second image processing unit 108 processes the visible light image, and the processing unit 102 receives the processing result sent by the second image processing unit 108 .
[0136] Optionally, the second image processing unit 108 is an embedded processing unit, such as a field programmable gate array (FPGA), etc., which is embedded in the motion capture module 101, and can also exist in the motion capture system in the form of a host computer.
[0137] In the above technical solution, the second image processing unit is used to perform feature extraction, feature recognition, feature space positioning calculation or motion data solution on the visible light image, and send the processing results to the processing unit. The second image processing unit can be an embedded processing unit or it can exist in the motion capture system in the form of a host computer, thereby improving the flexibility of the motion capture system.
[0138] Optionally, the second image processing unit 108 uses a deep learning algorithm, such as YOLOv8, to detect key points in the visible light image to obtain the coordinates and IDs of the key points, where the ID is a unique identifier of the key points. Optionally, the second image processing unit 108 can use algorithms such as MediaPipe to estimate the motion posture of the target under test in the visible light image to obtain a second pose estimation result.
[0139] The second image processing unit is used to detect key points of the visible light image, obtain the coordinates and IDs of the key points, estimate the motion posture of the target under test in the visible light image, and obtain a second pose estimation result. The second pose estimation result can be used to fuse with the pose estimation results obtained based on perception data of other modalities to obtain motion data of the target under test, thereby improving the accuracy and robustness of the motion capture system for motion capture of the target under test.
[0140] like Figure 5As shown, in the case where the motion capture module 101 in the motion capture system includes at least one non-visible light image sensor 104, at least one visible light image sensor 105 and at least one inertial sensor 103, and the motion capture system further includes a first image processing unit 107 and a second image processing unit 108, the processing unit 102 is used to perform at least one of the following:
[0141] Performing three-dimensional reconstruction of the marker points according to the marker point coordinates;
[0142] Performing ID tracking on the marker point to obtain a marker point ID tracking result;
[0143] Calculate the first pose estimation result of the measured object according to the coordinates of at least three non-collinear landmark points on the measured object;
[0144] According to the inertial data, a motion posture of the measured target is estimated to obtain a third posture estimation result of the measured target;
[0145] Reprojecting the marker point onto the visible light image and comparing it with the ID of the key point to verify the correctness of the marker point ID tracking result or to perform ID identification on the marker point that has lost tracking;
[0146] The first pose estimation result, the second pose estimation result and the third pose estimation result are fused and calculated to obtain motion data of the measured target.
[0147] It can be understood that the motion capture module 101 includes at least one non-visible light image sensor 104, which can sense the object under test from at least one perspective to obtain images in the near-infrared or infrared bands, and the first image processing unit 107 uses image processing technology to track the marker points in these near-infrared or infrared band images, obtain the marker point coordinates through spatial positioning calculation and output them to the processing unit 102, that is, the processing unit 102 receives the marker point coordinates.
[0148] The processing unit 102 performs multi-view 3D reconstruction on the object to be measured by using a triangulation method according to the coordinates of the marker points to obtain the 3D coordinates of the marker points.
[0149] The processing unit 102 tracks each marker point to determine its position change in continuous image frames or time series, and assigns a unique ID to each marker point and tracks the ID of each marker point to achieve the ability to identify and associate the same marker point during the dynamic capture of the measured target. Optionally, the processing unit 102 uses a method such as Kalman filtering to track the marker point ID to obtain a marker point ID tracking result, which includes the three-dimensional coordinates of each marker point and its corresponding image frame sequence or time series.
[0150] Furthermore, the processing unit 102 calculates a first pose estimation result PO based on the three-dimensional coordinates of at least three non-collinear landmark points.
[0151] Optionally, the processing unit 102 uses a complementary filtering algorithm to estimate the motion posture of the target according to the inertial data to obtain a third pose estimation result of the target, wherein the spatial positioning is replaced by the displacement vector T obtained in the first pose estimation result PO.
[0152] Furthermore, the processing unit 102 reprojects the marker point onto the visible light image according to the three-dimensional coordinates of the marker point, converts the three-dimensional spatial coordinates of the marker point into two-dimensional reprojection coordinates, and compares the projection of the marker point in the visible light image with the key points detected in the visible light image, including comparing their coordinate positions and IDs to verify the accuracy of the marker point ID tracking results, thereby verifying the correctness of marker ID tracking or performing ID identification on marker points that have lost tracking.
[0153] Optionally, the processing unit 102 uses methods such as Kalman filtering to perform fusion calculations on the first pose estimation result PO, the second pose estimation result PA and the third pose estimation result PB to obtain motion data P of the target under test. The motion data of the target under test describes the motion posture of the target under test in three-dimensional space, thereby realizing motion capture of the target under test.
[0154] In the above technical scheme, the processing unit performs three-dimensional reconstruction of the marker points according to the coordinates of the marker points, performs ID tracking on the marker points to obtain the marker point ID tracking results, calculates the first pose estimation result of the measured target according to the coordinates of at least three non-collinear marker points on the measured target, estimates the motion posture of the measured target according to the inertial data to obtain the third pose estimation result of the measured target, reprojects the marker points onto the visible light image, and compares them with the IDs of the key points to verify the correctness of the marker point ID tracking results or to perform ID identification on the marker points that have lost tracking, thereby achieving the pose estimation result of the measured target based on the perception data of different modalities, and fusing the pose estimation results obtained based on the perception data of each modality to obtain the motion data of the measured target, thereby improving the accuracy and robustness of the motion capture system for the motion capture of the measured target.
[0155] The motion capture system provided in the above-mentioned embodiments of the present application realizes optical-visual-inertial multi-modal fusion motion capture, and can be used for high-precision, high-robustness motion tracking, positioning and kinematic solution of multiple targets.
[0156] Figure 6 is a flow chart of a motion capture method provided by some embodiments of the present application. Figure 6As shown, the motion capture method includes: step 610 and step 620.
[0157] Step 610: sensing the target to obtain sensing data of different modes;
[0158] The sensing of the target to be measured to obtain sensing data of different modes includes:
[0159] The inertial data of the measured target during its motion is measured.
[0160] In some embodiments, sensing the target to obtain sensing data of different modalities further includes:
[0161] sensing the target in a near-infrared or infrared band, and outputting an image in the near-infrared or infrared band, wherein the image in the near-infrared or infrared band includes at least one marker point; and / or,
[0162] Tracking and spatially locating the marker points in the image of the near-infrared or infrared band to obtain the coordinates of the marker points;
[0163] The target is sensed in the visible light band and a visible light image is output.
[0164] Step 620: Process the perception data of different modes to obtain motion data of the measured target.
[0165] Optionally, the processing the perception data of different modalities to obtain the motion data of the measured target includes:
[0166] Tracking and spatially locating the marker points in the image of the near-infrared or infrared band to obtain the coordinates of the marker points;
[0167] Performing three-dimensional reconstruction of the marker points according to the marker point coordinates;
[0168] Performing ID tracking on the marker point to obtain a marker point ID tracking result;
[0169] Calculate the first pose estimation result of the measured object according to the coordinates of at least three non-collinear landmark points on the measured object;
[0170] Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points;
[0171] Performing motion posture estimation on the target to be measured in the visible light image to obtain a second posture estimation result of the target to be measured;
[0172] According to the inertial data, a motion posture of the measured target is estimated to obtain a third posture estimation result of the measured target;
[0173] Reprojecting the marker point onto the visible light image and comparing it with the ID of the key point to verify the correctness of the marker point ID tracking result or to perform ID identification on the marker point that has lost tracking;
[0174] The first pose estimation result, the second pose estimation result and the third pose estimation result are fused and calculated to obtain motion data of the measured target.
[0175] The marker points are reconstructed in three dimensions, and the marker points are tracked by ID to obtain the marker point ID tracking results. The first pose estimation result of the measured target is calculated according to the coordinates of at least three non-collinear marker points on the measured target. The key point detection is performed on the visible light image to obtain the coordinates and IDs of the key points. The motion posture of the measured target in the visible light image is estimated to obtain the second pose estimation result. The motion posture of the measured target is estimated according to the inertial data to obtain the third pose estimation result of the measured target. The marker points are reprojected onto the visible light image and compared with the IDs of the key points to verify the correctness of the marker point ID tracking results or to perform ID identification on the marker points that have lost tracking. The pose estimation result of the measured target is obtained based on the perception data of different modalities, and the pose estimation results obtained based on the perception data of each modality are fused to obtain the motion data of the measured target, thereby improving the accuracy and robustness of the motion capture system for the motion capture of the measured target.
[0176] In the above technical scheme, the target to be measured is sensed to obtain perception data of different modalities, including sensing the target to be measured in the near-infrared or infrared band, outputting an image of the near-infrared or infrared band, wherein the image of the near-infrared or infrared band includes at least one marker point, tracking and spatially locating the marker point in the image of the near-infrared or infrared band to obtain the coordinates of the marker point, sensing the target to be measured in the visible light band, outputting a visible light image, and measuring the inertial data of the target to be measured during movement, processing the perception data of different modalities to obtain the motion data of the target to be measured, realizing the fusion of perception data of different modalities, and improving the accuracy and robustness of the motion capture method.
[0177] For understanding of the motion capture method provided in the embodiment of the present application, one can refer to the aforementioned description of the motion capture system, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0178] In some embodiments, Figure 7As shown, an embodiment of the present application also provides an electronic device 700, including a processor 701, a memory 702, and a computer program stored in the memory 702 and executable on the processor 701. When the program is executed by the processor 701, each process of the above-mentioned motion capture method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0179] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.
[0180] An embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned motion capture method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0181] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0182] An embodiment of the present application also provides a computer program product, including a computer program, which implements the above-mentioned motion capture method when executed by a processor.
[0183] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0184] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned motion capture method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0185] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0186] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0187] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0188] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
[0189] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0190] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present application, and that the scope of the present application is defined by the claims and their equivalents.
Claims
1. A motion capture system, characterized in that: include: The motion capture module includes two or more modal sensors, which are used to sense the target and obtain sensing data of different modalities; A processing unit, connected to the motion capture module, for processing the perception data of the different modes to obtain the motion data of the measured target; The motion capture module includes at least one inertial sensor, which is arranged on the measured object and is used to measure inertial data of the measured object during movement, and output the inertial data to the processing unit.
2. The motion capture system according to claim 1, characterized in that: The motion capture module also includes: at least one non-visible light image sensor, the non-visible light image sensor being used to sense the target in a near-infrared or infrared band and output an image in a near-infrared or infrared band; and / or, At least one visible light image sensor, wherein the visible light image sensor is used to sense the target in the visible light band and output a visible light image.
3. The motion capture system according to claim 2, characterized in that: In the case where the motion capture system includes at least one non-visible light image sensor, the motion capture system further includes at least one near-infrared or infrared fill light source for illuminating the target; At least one marking point is arranged on the measured target, and under the illumination of the supplementary light source, the marking point has a contrast with the background in the imaging of the non-visible light image sensor.
4. The motion capture system according to claim 3, characterized in that: The motion capture system also includes: a first image processing unit, connected to the non-visible light image sensor, for tracking and spatially locating the marker points in the image of the near-infrared or infrared band, and outputting the coordinates of the marker points to the processing unit.
5. The motion capture system according to claim 2 or 4, characterized in that: In the case where the motion capture system includes at least one visible light image sensor, the motion capture system further includes: a second image processing unit connected to the visible light image sensor, configured to: Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points; The motion posture of the measured target in the visible light image is estimated to obtain a second posture estimation result of the measured target.
6. The motion capture system according to claim 5, characterized in that: The processing unit is configured to perform at least one of the following: Performing three-dimensional reconstruction of the marker points according to the marker point coordinates; Performing ID tracking on the marker point to obtain a marker point ID tracking result; Calculate the first pose estimation result of the measured object according to the coordinates of at least three non-collinear landmark points on the measured object; According to the inertial data, a motion posture of the measured target is estimated to obtain a third posture estimation result of the measured target; Reprojecting the marker point onto the visible light image and comparing it with the ID of the key point to verify the correctness of the marker point ID tracking result or to perform ID identification on the marker point that has lost tracking; The first pose estimation result, the second pose estimation result and the third pose estimation result are fused and calculated to obtain motion data of the measured target.
7. A motion capture method, based on the motion capture system according to any one of claims 1 to 6, characterized in that: include: Perceive the target to be measured and obtain perception data of different modes; Processing the sensing data of the different modes to obtain motion data of the measured target; The sensing of the target to be measured to obtain sensing data of different modes includes: The inertial data of the measured target during its motion is measured.
8. The motion capture method according to claim 7, characterized in that: The sensing of the measured target to obtain sensing data of different modes also includes: Sensing the target in a near-infrared or infrared band, and outputting an image in the near-infrared or infrared band, wherein the image in the near-infrared or infrared band includes at least one marker point; and / or, The target is sensed in the visible light band and a visible light image is output.
9. The motion capture method according to claim 8, characterized in that: The processing of the sensing data of the different modes to obtain the motion data of the measured target includes: Tracking and spatially locating the marker points in the image of the near-infrared or infrared band to obtain the coordinates of the marker points; Performing three-dimensional reconstruction of the marker points according to the marker point coordinates; Performing ID tracking on the marker point to obtain a marker point ID tracking result; Calculate the first pose estimation result of the measured object according to the coordinates of at least three non-collinear landmark points on the measured object; Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points; Performing motion posture estimation on the target to be measured in the visible light image to obtain a second posture estimation result of the target to be measured; According to the inertial data, a motion posture of the measured target is estimated to obtain a third posture estimation result of the measured target; Reprojecting the marker point onto the visible light image and comparing it with the ID of the key point to verify the correctness of the marker point ID tracking result or to perform ID identification on the marker point that has lost tracking; The first pose estimation result, the second pose estimation result and the third pose estimation result are fused and calculated to obtain motion data of the measured target.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the motion capture method as described in any one of claims 7 to 9 is implemented.
Citation Information
Cited By
Hand motion capture system and method and readable storage medium
CN121095318A
Motion capture system and motion capture method
WO2026149260A1