Motion capture system, motion capture method, and storage medium
By using multimodal sensors and upper computer processing units in the motion capture system to process perceived data of different modes, the problem of insufficient accuracy and robustness in complex environments in the prior art is solved, and higher motion capture accuracy and applicability are achieved.
Patent Information
- Application Number
- CN202510033674.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-13
AI Technical Summary
The existing optical motion capture technology has problems of insufficient accuracy and robustness in complex environments.
The motion capture system using a multimodal sensor, including a non-visible image sensor and a visible image sensor, perceives the measured target through an optical lens, acquires perceived data of different modes, and processes these data using a computer processing unit to obtain the motion position data of the measured target.
The motion capture accuracy and robustness of the motion capture system to the measured target is improved, and its applicability is enhanced in complex environments.
Smart Images

Figure CN119996792A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of motion capture technology, and in particular, relates to a motion capture system, a motion capture method and a storage medium. Background Art
[0002] Optical motion capture technology is widely used in many fields due to its high precision and fast response. This technology mainly relies on near-infrared motion capture cameras to monitor and track specific markers installed on the target object. Multiple near-infrared cameras are arranged around the target object. When the marker is captured by two or more cameras at the same time, the three-dimensional spatial position of the point can be determined. Through continuous shooting, the motion trajectory of the marker over time can be tracked. However, in complex environments, there are certain limitations in perceiving the object only through near-infrared cameras. Therefore, it is necessary to propose a system or method that can improve the above limitations to improve the accuracy and robustness of motion capture. Summary of the invention
[0003] The present application aims to solve at least one of the technical problems existing in the related art. To this end, the present application proposes a motion capture system, a motion capture method and a storage medium, which realizes the motion capture of the measured target, obtains motion posture data based on the perception data of different modalities, and improves the accuracy and robustness of the motion capture system for the motion capture of the measured target.
[0004] In a first aspect, the present application provides a motion capture system, the system comprising:
[0005] At least one motion capture camera, each of which includes sensors of multiple modes and an optical lens arranged in front of each of the sensors, wherein the sensor is used to sense the target through the optical lens to obtain sensing data of different modes;
[0006] The host computer processing unit is connected to the motion capture camera and is used to process the perception data of different modes to obtain the motion posture data of the measured target.
[0007] In the above technical solution, the motion capture system includes at least one motion capture camera, each motion capture camera includes sensors of multiple modalities, the sensors are used to sense the target to be measured through an optical lens arranged in the front, and obtain perception data of different modalities. The host computer processing unit is connected to the motion capture camera, and is used to process the perception data of different modalities to obtain motion posture data of the target to be measured, thereby realizing motion capture of the target to be measured, and obtaining motion posture data based on perception data of different modalities, thereby improving the accuracy and robustness of the motion capture system for motion capture of the target to be measured.
[0008] According to one embodiment of the present application, the motion capture camera includes at least one non-visible light image sensor and at least one visible light image sensor.
[0009] The non-visible light image sensor is used to sense the target in the near infrared or infrared band and output an image in the near infrared or infrared band;
[0010] The visible light image sensor is used to sense the target in the visible light band and output a visible light image.
[0011] In the above technical solution, the motion capture camera includes at least one non-visible light image sensor and at least one visible light image sensor. The non-visible light image sensor is used to sense the target to be measured in the near-infrared or infrared band and output images in the near-infrared or infrared band. The visible light image sensor is used to sense the target to be measured in the visible light band and output visible light images, thereby realizing optical and visual dual-modal perception of the target to be measured, and improving the accuracy and robustness of the motion capture system in capturing motion of the target to be measured.
[0012] According to one embodiment of the present application, the motion capture camera further includes at least one fill light source in the near infrared or infrared band, for illuminating the target to be measured;
[0013] At least one marking point is arranged on the measured target, and under the illumination of the supplementary light source, the marking point has a contrast with the background in the imaging of the non-visible light image sensor.
[0014] In the above technical solution, the motion capture camera includes at least one fill light source in the near-infrared or infrared band, which is used to illuminate the target to be measured, and at least one marker point is arranged on the target to be measured. Under the illumination of the fill light source, the marker point has a contrast with the background in the imaging of the non-visible light image sensor, making the marker point more prominent, thereby improving the accuracy of identifying and locating the marker point in the imaging of the non-visible light image sensor, and thereby improving the accuracy of the motion capture camera in capturing the motion of the target to be measured.
[0015] According to one embodiment of the present application, the motion capture camera further includes: an embedded processing unit connected to the non-visible light image sensor and the visible light image sensor,
[0016] The embedded processing unit is used for:
[0017] Tracking and spatially locating the marker points in the image of the near-infrared or infrared band, and sending the coordinates of the marker points to the host computer processing unit;
[0018] The visible light image is processed by one or more of the following: feature extraction, feature recognition, feature space positioning calculation, motion data solution, and the processing result is sent to the host computer processing unit.
[0019] In the above technical solution, the motion capture camera also includes an embedded processing unit, which is connected to the non-visible light image sensor and the visible light image sensor, and is used to track and spatially locate the marker points in the near-infrared or infrared band images, and output the coordinates of the marker points to the host processing unit, perform feature extraction, feature recognition, feature spatial positioning calculation or motion data solution on the visible light image, and send the processing results to the host processing unit, thereby improving the flexibility of the motion capture system.
[0020] According to one embodiment of the present application, the embedded processing unit is specifically used for:
[0021] Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points;
[0022] Estimating the motion posture of the target in the visible light image to obtain first-position posture data;
[0023] Matching the coordinates of the key point with the coordinates of the marker point to determine an ID for the marker point;
[0024] The coordinates of the marker point, the ID of the marker point and the first pose data are sent to the host computer processing unit.
[0025] In the above technical scheme, the embedded processing unit is used to detect key points of the visible light image, obtain the coordinates and IDs of the key points, estimate the motion posture of the target in the visible light image, obtain the first pose data, match the coordinates of the key points with the coordinates of the marker points, determine the ID for the marker points, and send the coordinates of the marker points, the ID of the marker points and the first pose data to the host computer processing unit, thereby realizing data processing of the visible light image in the embedded processing unit and improving the flexibility of the motion capture system.
[0026] According to one embodiment of the present application, the host computer processing unit is used to:
[0027] According to the coordinates and ID of the marker point, the marker point is three-dimensionally reconstructed to obtain the three-dimensional coordinates of the marker point;
[0028] According to the three-dimensional coordinates of at least three non-collinear marker points, the position and posture of the measured object are solved to obtain second position and posture data;
[0029] The first posture data and the second posture data are fused and calculated to obtain motion posture data of the measured target.
[0030] In the above technical scheme, the host computer processing unit is used to perform three-dimensional reconstruction of the marker points according to the coordinates and ID of the marker points to obtain the three-dimensional coordinates of the marker points, and to solve the pose of the measured target according to the three-dimensional coordinates of at least three non-collinear marker points to obtain second pose data. The first pose data and the second pose data are fused and calculated to obtain the motion pose data of the measured target, thereby realizing the fusion of motion pose data obtained from perception data based on different modalities, and improving the accuracy and robustness of the motion capture system for motion capture of the measured target.
[0031] In a second aspect, the present application provides a motion capture method, the method comprising:
[0032] Perceive the target to be measured and obtain perception data of different modes;
[0033] The perception data of the different modes are processed to obtain the motion posture data of the measured target.
[0034] In the above technical scheme, the target to be measured is sensed to obtain perception data of different modalities, and the perception data of different modalities are processed to obtain motion posture data of the target to be measured, thereby realizing the fusion of perception data of different modalities and improving the accuracy and robustness of the motion capture method for motion capture of the target to be measured.
[0035] According to an embodiment of the present application, sensing the target to obtain sensing data of different modalities includes:
[0036] Sensing the measured target in a near-infrared or infrared band, and outputting an image of the near-infrared or infrared band, wherein the image of the near-infrared or infrared band includes at least one marking point;
[0037] Tracking and spatially locating the marker points in the image of the near-infrared or infrared band to obtain the coordinates of the marker points;
[0038] Sensing the measured target in the visible light band and outputting a visible light image;
[0039] Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points;
[0040] Estimating the motion posture of the target in the visible light image to obtain first-position posture data;
[0041] The coordinates of the key point and the coordinates of the marker point are matched, an ID is determined for the marker point, and the ID of the marker point is obtained.
[0042] In the above technical scheme, the target to be measured is sensed in the near-infrared or infrared band, and an image of the near-infrared or infrared band is output, the image of the near-infrared or infrared band includes at least one marker point, and the marker point in the image of the near-infrared or infrared band is tracked and spatially positioned to obtain the coordinates of the marker point; the target to be measured is sensed in the visible light band, and a visible light image is output, and key point detection is performed on the visible light image to obtain the coordinates and ID of the key point; the motion posture of the target to be measured in the visible light image is estimated to obtain the first position posture data; the coordinates of the key point are matched with the coordinates of the marker point, and the ID of the marker point is determined to obtain the ID of the marker point, which helps to obtain the motion data of the target to be measured based on the perception data of different modalities, thereby improving the accuracy and robustness of the motion capture system for the motion capture of the target to be measured.
[0043] According to an embodiment of the present application, the processing of the perception data of different modalities to obtain the motion posture data of the measured target includes:
[0044] According to the coordinates and ID of the marker point, the marker point is three-dimensionally reconstructed to obtain the three-dimensional coordinates of the marker point;
[0045] According to the three-dimensional coordinates of at least three non-collinear marker points, the position and posture of the measured object are solved to obtain second position and posture data;
[0046] The first posture data and the second posture data are fused and calculated to obtain motion posture data of the measured target.
[0047] In the above technical scheme, the marker points are three-dimensionally reconstructed according to the coordinates and ID of the marker points to obtain the three-dimensional coordinates of the marker points, and the pose of the measured target is solved according to the three-dimensional coordinates of at least three non-collinear marker points to obtain second pose data. The first pose data and the second pose data are fused and calculated to obtain the motion pose data of the measured target, thereby realizing the fusion of motion pose data obtained from perception data of different modalities, and improving the accuracy and robustness of the motion capture method for motion capture of the measured target.
[0048] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the motion capture method as described in the second aspect above when executing the computer program.
[0049] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the motion capture method as described in the second aspect above.
[0050] In a fifth aspect, the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the motion capture method as described in the second aspect.
[0051] In a sixth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the motion capture method as described in the second aspect above.
[0052] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0054] Figure 1 is one of the structural schematic diagrams of the motion capture system provided in some embodiments of the present application;
[0055] Figure 2 This is the second structural schematic diagram of the motion capture system provided by some embodiments of the present application;
[0056] Figure 3 This is the third structural schematic diagram of the motion capture system provided by some embodiments of the present application;
[0057] Figure 4 This is a fourth structural diagram of a motion capture system provided by some embodiments of the present application;
[0058] Figure 5 is a flowchart of a motion capture method provided by some embodiments of the present application;
[0059] Figure 6 It is a schematic diagram of the structure of an electronic device provided in some embodiments of the present application.
[0060] Description of reference numerals:
[0061] 101: motion capture camera; 102: sensor; 103: optical lens;
[0062] 104: host computer processing unit; 105: fill light source; 106: embedded processing unit;
[0063] 1021: non-visible light image sensor; 1022: visible light image sensor;
[0064] 600: electronic device; 601: processor; 602: memory. DETAILED DESCRIPTION
[0065] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.
[0066] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0067] The motion capture system, motion capture method and storage medium provided in the embodiments of the present application are described in detail below with reference to the accompanying drawings through specific embodiments and their application scenarios.
[0068] Figure 1 FIG. 1 is one of the structural diagrams of the motion capture system provided in some embodiments of the present application. Figure 1 As shown, the motion capture system includes:
[0069] At least one motion capture camera 101, each of the motion capture cameras 101 includes sensors 102 of multiple modes, and an optical lens 103 arranged in front of each of the sensors 102, the sensor 102 is used to sense the target through the optical lens 103 to obtain sensing data of different modes;
[0070] The host computer processing unit 104 is connected to the motion capture camera 101 and is used to process the perception data of different modes to obtain the motion posture data of the measured target.
[0071] It should be noted that the target to be measured in the embodiments of the present application may be one or more targets, and the present application does not impose any limitation on this.
[0072] It can be understood that the motion capture camera 101 includes sensors 102 of multiple modes and an optical lens 103 arranged in front of each sensor 102, which can be used to perceive the target to be measured and obtain perception data of different modes of the target to be measured. The optical lens 103 can be used to adjust and focus the light captured by the sensor 102 so that each sensor 102 can perceive the target to be measured more clearly and accurately. The sensor 102 and the optical lens 103 work in coordination to obtain perception data of different modes of the target to be measured.
[0073] Each motion capture camera 101 includes sensors 102 of multiple modes, and these sensors 102 can be of different types, for example: non-visible light image sensors, visible light image sensors, etc. Each modality of sensor 102 can sense the target to be measured and obtain perception data of a specific modality, such as a near-infrared band image sensor senses the target to be measured to obtain a near-infrared band image, and a visible light image sensor senses the target to be measured to obtain a visible light image, etc. Through sensors 102 of multiple modes, each motion capture camera 101 can sense the target to be measured from multiple modalities, thereby improving the comprehensiveness and accuracy of the perception of the target to be measured. Optionally, the Zhang Zhengyou calibration method is used to calibrate the sensors of each modality.
[0074] It can be understood that the motion capture system includes at least one motion capture camera 101, that is, the motion capture system can include multiple motion capture cameras 101. When the line of sight of a motion capture camera 101 is blocked or fails, other motion capture cameras 101 can continue to perceive the target under measurement, thereby improving the robustness of the motion capture system. Multiple motion capture cameras can perceive the target under measurement from different angles (multiple viewpoints), thereby improving the comprehensiveness and accuracy of the perception of the target under measurement. Multiple motion capture cameras work in coordination to improve the accuracy and robustness of motion capture.
[0075] In some embodiments, the motion capture camera 101 includes at least one non-visible light image sensor and at least one visible light image sensor, which are used to perceive the target to be measured, and obtain a non-visible light image and a visible light image of the target to be measured, respectively, to achieve optical and visual dual-modal perception of the target to be measured.
[0076] The host computer processing unit 104 is used to process the perception data of different modes, such as data matching, fusion and analysis, to obtain the motion posture data of the measured target, and the motion posture data of the measured target includes the position data and posture data of the measured target in three-dimensional space, thereby realizing the motion capture of the measured target. It can be understood that the motion posture data is obtained based on the perception data of different modes, and the fusion of the perception data of different modes is realized, which improves the accuracy of the motion capture system for the motion capture of the measured target. The sensor of a single mode may be affected by occlusion or other factors, resulting in the output perception data being inaccurate. The fusion of the perception data of sensors of multiple modes can improve the robustness of the motion capture system.
[0077] In the above technical solution, the motion capture system includes at least one motion capture camera, each motion capture camera includes sensors of multiple modalities, the sensors are used to sense the target to be measured through an optical lens arranged in the front, and obtain perception data of different modalities. The host computer processing unit is connected to the motion capture camera, and is used to process the perception data of different modalities to obtain motion posture data of the target to be measured, thereby realizing motion capture of the target to be measured, and obtaining motion posture data based on perception data of different modalities, thereby improving the accuracy and robustness of the motion capture system for motion capture of the target to be measured.
[0078] Figure 2 This is the second structural diagram of the motion capture system provided by some embodiments of the present application. Figure 2 As shown, in some embodiments, the motion capture camera 101 includes at least one non-visible light image sensor 1021 and at least one visible light image sensor 1022.
[0079] The non-visible light image sensor 1021 is used to sense the target in the near infrared or infrared band and output the image in the near infrared or infrared band;
[0080] The visible light image sensor 1022 is used to sense the target in the visible light band and output a visible light image.
[0081] It can be understood that the non-visible light image sensor 1021 is used to sense the target to be measured in the near-infrared or infrared band and output images in the near-infrared or infrared band. Since the light in the near-infrared and infrared bands is invisible to the human eye, the non-visible light image sensor is not limited by visible light conditions and is more suitable for environments with undesirable visible light conditions. The motion capture camera includes at least one non-visible light image sensor 1021, which improves the applicability of the motion capture system under different lighting conditions; the visible light image sensor 1022 is used to sense the target to be measured in the visible light band and output visible light images. Under good lighting conditions, the visible light image sensor 1022 can output high-resolution color images.
[0082] The motion capture camera 101 includes at least one non-visible light image sensor 1021 and at least one visible light image sensor 1022, so that the motion capture camera can perceive the target from both optical and visual modes. The optical mode corresponds to the image of the near-infrared or infrared band output by the non-visible light image sensor 1021, and the visual mode corresponds to the visible light image output by the visible light image sensor 1022, thereby improving the comprehensiveness and accuracy of the perception of the target, obtaining motion posture data based on perception data of different modes, and also improving the accuracy of the motion capture system in capturing the motion of the target.
[0083] In the above technical solution, the motion capture camera includes at least one non-visible light image sensor and at least one visible light image sensor. The non-visible light image sensor is used to sense the target to be measured in the near-infrared or infrared band and output images in the near-infrared or infrared band. The visible light image sensor is used to sense the target to be measured in the visible light band and output visible light images, thereby realizing optical and visual dual-modal perception of the target to be measured, and improving the accuracy and robustness of the motion capture system in capturing motion of the target to be measured.
[0084] Figure 3 This is the third structural diagram of the motion capture system provided by some embodiments of the present application. Figure 3 As shown, in some embodiments, the motion capture camera 101 further includes at least one fill light source 105 in the near infrared or infrared band, for illuminating the target to be measured;
[0085] At least one marking point is arranged on the measured target, and under the illumination of the supplementary light source, the marking point has a contrast with the background in the imaging of the non-visible light image sensor.
[0086] The near-infrared or infrared band fill light source 105 is used to provide near-infrared or infrared band light to illuminate the target to be measured, so that the non-visible light image sensor 1021 can perceive the target to be measured more clearly, thereby improving the quality of the near-infrared or infrared band image output by the non-visible light image sensor 1021, and also enabling the non-visible light image sensor 1021 to still perceive the target to be measured when the near-infrared or infrared band lighting conditions are not ideal.
[0087] At least one mark point is arranged on the target to be measured. The mark point is a special mark or light-emitting point, also called a marker. Under the illumination of the fill light source 105, the mark point will form a contrast with the background in the imaging of the non-visible light image sensor, making the mark point more prominent in the near-infrared or infrared band image output by the non-visible light image sensor, and forming a contrast with the surrounding background, making the mark point easier to identify and track in subsequent image processing and analysis, thereby improving the accuracy of identifying and locating the mark point.
[0088] In the above technical solution, the motion capture camera includes at least one fill light source in the near-infrared or infrared band, which is used to illuminate the target to be measured, and at least one marker point is arranged on the target to be measured. Under the illumination of the fill light source, the marker point has a contrast with the background in the imaging of the non-visible light image sensor, making the marker point more prominent, thereby improving the accuracy of identifying and locating the marker point in the imaging of the non-visible light image sensor, and thereby improving the accuracy of the motion capture camera in capturing the motion of the target to be measured.
[0089] like Figure 3 As shown, in some embodiments, the host computer processing unit 104 is specifically used to:
[0090] Tracking and spatially locating the marker points in the image of the near-infrared or infrared band to obtain the coordinates of the marker points;
[0091] Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points;
[0092] Estimating the motion posture of the target in the visible light image to obtain first-position posture data;
[0093] Matching the coordinates of the key point with the coordinates of the marker point, determining an ID for the marker point, and obtaining the ID of the marker point;
[0094] According to the coordinates and ID of the marker point, the marker point is three-dimensionally reconstructed to obtain the three-dimensional coordinates of the marker point;
[0095] According to the three-dimensional coordinates of at least three non-collinear marker points, the position and posture of the measured object are solved to obtain second position and posture data;
[0096] The first posture data and the second posture data are fused and calculated to obtain motion posture data of the measured target.
[0097] It can be understood that at least one marker is arranged on the target to be measured. Under the illumination of the fill light source, these marker points contrast with the background in the near-infrared or infrared band image output by the non-visible light image sensor, and appear as bright spots. Through image processing technology, the upper computer processing unit 104 can track the marker points in the near-infrared or infrared band image, and obtain the coordinates of the marker points through spatial positioning calculation.
[0098] Optionally, the host computer processing unit 104 uses a deep learning algorithm, such as YOLOv8, to detect key points in the visible light image to obtain the coordinates and ID of the key points, where the ID is a unique identifier of the key point. Optionally, the host computer processing unit 104 uses an algorithm such as MediaPipe to estimate the motion posture of the target under test in the visible light image to obtain the first pose data. The motion pose data of the target under test includes the position data and pose data of the target under test in three-dimensional space.
[0099] The host computer processing unit 104 matches the coordinates of the marker points with the coordinates of the key points, establishes a correspondence between the marker points and the key points, and determines an ID for each marker point. Optionally, determining an ID for a marker point includes: comparing the key point coordinates with the marker point coordinates, determining the correspondence between the marker point and the key point by finding the nearest match, assigning a consistent ID value to the marker points that match the key points at the same spatial position, so that the marker points have unique and time-domain consistent IDs.
[0100] The host computer processing unit 104 reconstructs the marker points in three dimensions using triangulation according to the coordinates and ID of the marker points to obtain the three-dimensional coordinates of the marker points; based on the three-dimensional coordinates of at least three non-collinear marker points, the host computer processing unit 104 can perform posture calculation on the target to obtain second posture data.
[0101] Optionally, a fusion algorithm such as Kalman filtering is used to fuse the first pose data and the second pose data to obtain the motion pose data of the target under test, thereby realizing the fusion of motion pose data obtained based on different modal perception data. The fusion calculation can effectively improve the anti-occlusion ability of the target motion solution while ensuring that the accuracy loss is controllable, thereby improving the accuracy and robustness of the motion capture system for motion capture of the target under test.
[0102] In some embodiments, the motion capture system includes at least one motion capture camera, which can perceive the object to be measured from at least one viewing angle to obtain near-infrared or infrared band images and visible light images, and use triangulation to perform three-dimensional reconstruction of the object to be measured to obtain the three-dimensional coordinates of the marker points, and solve the three-dimensional coordinate pose based on at least three non-collinear marker points, thereby improving the accuracy and robustness of the motion capture system in capturing motion of the target to be measured.
[0103] In the above technical scheme, the host computer processing unit is used to track and spatially locate the marker points in the near-infrared or infrared band image to obtain the coordinates of the marker points, perform key point detection on the visible light image to obtain the coordinates and ID of the key points, estimate the motion posture of the target under test in the visible light image to obtain the first pose data, match the coordinates of the key points with the coordinates of the marker points, determine the ID for the marker points, perform three-dimensional reconstruction of the marker points based on the coordinates and ID of the marker points to obtain the three-dimensional coordinates of the marker points, perform pose solution on the target under test based on the three-dimensional coordinates of at least three non-collinear marker points to obtain the second pose data, fuse the first pose data and the second pose data to obtain the motion pose data of the target under test, thereby realizing the fusion of motion pose data obtained based on different modal perception data, and improving the accuracy and robustness of the motion capture system for the motion capture of the target under test.
[0104] Figure 4 FIG. 4 is a schematic diagram of the structure of the motion capture system provided in some embodiments of the present application. Figure 4 As shown, in some embodiments, the motion capture camera further includes: an embedded processing unit 106 connected to the non-visible light image sensor 1021 and the visible light image sensor 1022,
[0105] The embedded processing unit 106 is used for:
[0106] Tracking and spatial positioning calculation are performed on the marker points in the near-infrared or infrared band image, and the coordinates of the marker points are sent to the host computer processing unit.
[0107] It can be understood that the motion capture camera 101 includes a non-visible light image sensor 1021 and a visible light image sensor 1022, which are used to sense the target to be measured and obtain images in the near-infrared or infrared bands and visible light images. The host computer processing unit 104 processes the images in the near-infrared or infrared bands and the visible light images to obtain the motion posture data of the target to be measured, thereby realizing motion capture of the target to be measured.
[0108] In some embodiments, the motion capture camera 101 may also include an embedded processing unit 106, which uses image processing technology to track marker points in images in the near-infrared or infrared bands, and obtains the coordinates of the marker points through spatial positioning calculations. The host computer processing unit 104 receives the coordinates of the marker points output by the embedded processing unit 106, thereby improving the flexibility of the motion capture system.
[0109] In the above technical solution, the motion capture camera also includes an embedded processing unit, which is connected to the non-visible light image sensor and the visible light image sensor, and is used to track and spatially locate the marker points in the near-infrared or infrared band images, and output the coordinates of the marker points to the host computer processing unit, thereby improving the flexibility of the motion capture system.
[0110] like Figure 4 As shown, in some embodiments, the embedded processing unit 106 is further used for:
[0111] The visible light image is subjected to one or more of the following processing: feature extraction, feature recognition, feature space positioning calculation, motion data solution, and the processing result is sent to the host computer processing unit 104 .
[0112] It can be understood that feature extraction refers to extracting features from visible light images that can be used for subsequent motion capture of the target under test. Features may include edges, corners, textures, etc. in the image; feature recognition refers to the process of further determining the specific identity or category of these features based on feature extraction; feature space positioning calculation refers to estimating the spatial position and posture of the target under test by identifying and using features in the image; motion data solution refers to the process of determining the motion posture data of the target under test.
[0113] In the above technical solution, the embedded unit is also used to perform feature extraction, feature recognition, feature space positioning calculation or motion data solution on the visible image, and send the processing results to the host computer processing unit, thereby improving the flexibility of the motion capture system.
[0114] like Figure 4 As shown, in some embodiments, the embedded processing unit 106 is specifically used for:
[0115] Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points;
[0116] Estimating the motion posture of the target in the visible light image to obtain first-position posture data;
[0117] Matching the coordinates of the key point with the coordinates of the marker point to determine an ID for the marker point;
[0118] The coordinates of the marker point, the ID of the marker point and the first pose data are sent to the host computer processing unit.
[0119] Optionally, the embedded processing unit 106 uses a deep learning algorithm, such as YOLOv8, to detect key points in the visible light image to obtain the coordinates and ID of the key points, where the ID is a unique identifier of the key point. Optionally, the embedded processing unit 106 uses an algorithm such as MediaPipe to estimate the motion posture of the target under test in the visible light image to obtain the first pose data.
[0120] The embedded processing unit 106 is also used to match the coordinates of the marker point with the coordinates of the key point, establish a correspondence between the marker point and the key point, and determine a unique ID for each marker point. Optionally, determining the ID for the marker point includes: comparing the key point coordinates with the marker point coordinates, determining the correspondence between the marker point and the key point by finding the nearest match, and assigning a consistent ID value to the marker point that matches the key point at the same spatial position, so that the marker point has a unique and time-domain consistent ID.
[0121] In the above technical scheme, the embedded processing unit is used to detect key points of the visible light image, obtain the coordinates and IDs of the key points, estimate the motion posture of the target in the visible light image, obtain the first pose data, match the coordinates of the key points with the coordinates of the marker points, determine the ID for the marker points, and send the coordinates of the marker points, the ID of the marker points and the first pose data to the host computer processing unit, thereby realizing data processing of the visible light image in the embedded processing unit and improving the flexibility of the motion capture system.
[0122] like Figure 4 As shown, when the motion capture system further includes an embedded processing unit 106, the host computer processing unit 104 is used to:
[0123] According to the coordinates and ID of the marker point, the marker point is three-dimensionally reconstructed to obtain the three-dimensional coordinates of the marker point;
[0124] According to the three-dimensional coordinates of at least three non-collinear marker points, the position and posture of the measured object are solved to obtain second position and posture data;
[0125] The first posture data and the second posture data are fused and calculated to obtain motion posture data of the measured target.
[0126] After the host computer processing unit 104 receives the coordinates of the marker point, the ID of the marker point and the first pose data sent by the embedded processing unit 106, it uses triangulation to perform three-dimensional reconstruction of the marker point according to the coordinates and ID of the marker point to obtain the three-dimensional coordinates of the marker point; based on the three-dimensional coordinates of at least three non-collinear marker points, the host computer processing unit 104 can perform pose solution on the target to obtain second pose data.
[0127] Optionally, the upper computer processing unit 104 adopts a fusion algorithm such as Kalman filtering to fuse the first pose data and the second pose data to obtain the motion pose data of the target under test, thereby realizing the fusion of motion pose data obtained based on perception data of different modalities. The fusion calculation can effectively improve the anti-occlusion ability of the target motion solution while ensuring that the accuracy loss is controllable, thereby improving the accuracy and robustness of the motion capture system for motion capture of the target under test.
[0128] In the above technical scheme, the host computer processing unit is used to perform three-dimensional reconstruction of the marker points according to the coordinates and ID of the marker points to obtain the three-dimensional coordinates of the marker points, and to solve the pose of the measured target according to the three-dimensional coordinates of at least three non-collinear marker points to obtain second pose data. The first pose data and the second pose data are fused and calculated to obtain the motion pose data of the measured target, thereby realizing the fusion of motion pose data obtained from perception data based on different modalities, and improving the accuracy and robustness of the motion capture system for motion capture of the measured target.
[0129] The motion capture method provided in the embodiment of the present application may be executed by a motion capture system or a functional module or functional entity in the motion capture system that can implement the motion capture method. The motion capture method provided in the embodiment of the present application is described below using the motion capture system as an example of the execution body.
[0130] Figure 5 is a flow chart of a motion capture method provided by some embodiments of the present application. Figure 5 As shown, the motion capture method, based on the above-mentioned motion capture system, includes: step 510 and step 520.
[0131] Step 510: sense the target to obtain sensing data of different modes.
[0132] Optionally, the target to be measured is sensed by a motion capture camera, which includes sensors of multiple modalities, and these sensors can be of different types, such as non-visible light image sensors, visible light image sensors, etc. Each modality of sensor can sense the target to be measured to obtain sensing data of a specific modality, such as a near-infrared band image sensor senses the target to be measured to obtain a near-infrared band image, and a visible light image sensor senses the target to be measured to obtain a visible light image, etc. Through sensors of multiple modalities, the target to be measured can be sensed from multiple modalities, thereby improving the comprehensiveness and accuracy of sensing the target to be measured.
[0133] Step 520: Process the perception data of different modes to obtain motion posture data of the measured target.
[0134] It is understandable that the perception data of different modes are processed, such as data matching, fusion and analysis, and finally the motion posture data of the measured target is obtained. The motion posture data of the measured target includes the position data and posture data of the measured target in three-dimensional space, thereby realizing the motion capture of the measured target. It is understandable that the motion posture data is obtained based on the perception data of different modes, and the fusion of perception data of different modes is realized, which improves the accuracy of the motion capture system for motion capture of the measured target. Sensors of a single mode may be affected by occlusion or other factors, resulting in inaccurate output perception data. Fusion of perception data of sensors of multiple modes can improve the robustness of the motion capture system.
[0135] In the above technical scheme, the target to be measured is sensed to obtain perception data of different modalities, and the perception data of different modalities are processed to obtain motion posture data of the target to be measured, thereby realizing the fusion of perception data of different modalities and improving the accuracy and robustness of the motion capture method for motion capture of the target to be measured.
[0136] Optionally, sensing the target to be measured to obtain sensing data of different modes includes:
[0137] Sensing the measured target in a near-infrared or infrared band, and outputting an image of the near-infrared or infrared band, wherein the image of the near-infrared or infrared band includes at least one marking point;
[0138] Tracking and spatially locating the marker points in the image of the near-infrared or infrared band to obtain the coordinates of the marker points;
[0139] Sensing the measured target in the visible light band and outputting a visible light image;
[0140] Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points;
[0141] Estimating the motion posture of the target in the visible light image to obtain first-position posture data;
[0142] The coordinates of the key point and the coordinates of the marker point are matched, an ID is determined for the marker point, and the ID of the marker point is obtained.
[0143] It can be understood that the target to be measured is sensed in the near-infrared or infrared band and an image of the near-infrared or infrared band is output. Since the light in the near-infrared and infrared bands is invisible to the human eye, the perception of the target to be measured in the near-infrared or infrared band is not restricted by visible light conditions and can be carried out in environments where visible light conditions are not ideal, thereby improving the applicability of the motion capture method under different lighting conditions.
[0144] It can be understood that the target to be measured is sensed in the visible light band and a visible light image is output. Under good lighting conditions, the target to be measured is sensed in the visible light band and a high-resolution color visible light image can be output.
[0145] It can be understood that at least one marker is arranged on the target to be measured. These marker points contrast with the background in the near-infrared or infrared band image and appear as bright spots. Through image processing technology, the marker points in the near-infrared or infrared band image are followed and the coordinates of the marker points are obtained through spatial positioning calculation.
[0146] Optionally, a deep learning algorithm, such as YOLOv8, is used to detect key points in the visible light image to obtain the coordinates and ID of the key points, where the ID is a unique identifier of the key point. Optionally, an algorithm such as MediaPipe can be used to estimate the motion posture of the target in the visible light image to obtain the first pose data.
[0147] The coordinates of the marker points are matched with the coordinates of the key points, a corresponding relationship between the marker points and the key points is established, and a unique ID is determined for each marker point. Optionally, determining the ID for the marker point includes: comparing the key point coordinates with the marker point coordinates, determining the corresponding relationship between the marker point and the key point by finding the nearest match, and assigning a consistent ID value to the marker points that match the key points at the same spatial position, so that the marker points have unique and time-domain consistent IDs.
[0148] In the above technical scheme, the target to be measured is sensed in the near-infrared or infrared band, and an image of the near-infrared or infrared band is output, the image of the near-infrared or infrared band includes at least one marker point, and the marker point in the image of the near-infrared or infrared band is tracked and spatially positioned to obtain the coordinates of the marker point; the target to be measured is sensed in the visible light band, and a visible light image is output, and key point detection is performed on the visible light image to obtain the coordinates and ID of the key point; the motion posture of the target to be measured in the visible light image is estimated to obtain the first position posture data; the coordinates of the key point are matched with the coordinates of the marker point, and the ID of the marker point is determined to obtain the ID of the marker point, which helps to obtain the motion data of the target to be measured based on the perception data of different modalities, thereby improving the accuracy and robustness of the motion capture system for the motion capture of the target to be measured.
[0149] Optionally, the processing of the perception data of different modalities to obtain motion posture data of the measured target includes:
[0150] According to the coordinates and ID of the marker point, the marker point is three-dimensionally reconstructed to obtain the three-dimensional coordinates of the marker point;
[0151] According to the three-dimensional coordinates of at least three non-collinear marker points, the position and posture of the measured object are solved to obtain second position and posture data;
[0152] The first posture data and the second posture data are fused and calculated to obtain motion posture data of the measured target.
[0153] It can be understood that, according to the coordinates and ID of the marker points, the triangulation method is used to reconstruct the marker points in three dimensions to obtain the three-dimensional coordinates of the marker points; based on the three-dimensional coordinates of at least three non-collinear marker points, the pose of the target is solved to obtain the second pose data.
[0154] Optionally, a fusion algorithm such as Kalman filtering is used to fuse the first pose data and the second pose data to obtain the motion pose data of the target under test, thereby realizing the fusion of motion pose data obtained from perception data based on different modalities. The fusion calculation can effectively improve the anti-occlusion ability of the target motion solution while ensuring that the accuracy loss is controllable, thereby improving the accuracy and robustness of the motion capture system for motion capture of the target under test.
[0155] In the above technical scheme, the marker points are three-dimensionally reconstructed according to the coordinates and ID of the marker points to obtain the three-dimensional coordinates of the marker points, and the pose of the measured target is solved according to the three-dimensional coordinates of at least three non-collinear marker points to obtain second pose data. The first pose data and the second pose data are fused and calculated to obtain the motion pose data of the measured target, thereby realizing the fusion of motion pose data obtained from perception data of different modalities, and improving the accuracy and robustness of the motion capture method for motion capture of the measured target.
[0156] For understanding of the motion capture method provided in the embodiment of the present application, one can refer to the aforementioned description of the motion capture system, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0157] In some embodiments, Figure 6 As shown, an embodiment of the present application also provides an electronic device 600, including a processor 602, a memory 602, and a computer program stored in the memory 602 and executable on the processor 601. When the program is executed by the processor 602, each process of the above-mentioned motion capture method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0158] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.
[0159] An embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned motion capture method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0160] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0161] An embodiment of the present application also provides a computer program product, including a computer program, which implements the above-mentioned motion capture method when executed by a processor.
[0162] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0163] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned motion capture method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0164] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0165] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0166] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0167] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
[0168] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0169] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present application, and that the scope of the present application is defined by the claims and their equivalents.
Claims
1. A motion capture system, characterized in that: include: At least one motion capture camera, each of which includes sensors of multiple modes and an optical lens arranged in front of each of the sensors, wherein the sensor is used to sense the target through the optical lens to obtain sensing data of different modes; The host computer processing unit is connected to the motion capture camera and is used to process the perception data of different modes to obtain the motion posture data of the measured target.
2. The motion capture system according to claim 1, characterized in that: The motion capture camera includes at least one non-visible light image sensor and at least one visible light image sensor, The non-visible light image sensor is used to sense the target in the near infrared or infrared band and output an image in the near infrared or infrared band; The visible light image sensor is used to sense the target in the visible light band and output a visible light image.
3. The motion capture system according to claim 2, characterized in that: The motion capture camera further comprises at least one fill light source in the near infrared or infrared band, for illuminating the measured target; At least one marking point is arranged on the measured target, and under the illumination of the supplementary light source, the marking point has a contrast with the background in the imaging of the non-visible light image sensor.
4. The motion capture system according to claim 3, characterized in that: The motion capture camera further includes: an embedded processing unit connected to the non-visible light image sensor and the visible light image sensor, The embedded processing unit is used for: Tracking and spatially locating the marker points in the image of the near-infrared or infrared band, and sending the coordinates of the marker points to the host computer processing unit; The visible light image is processed by one or more of the following: feature extraction, feature recognition, feature space positioning calculation, motion data solution, and the processing result is sent to the host computer processing unit.
5. The motion capture system according to claim 4, characterized in that: The embedded processing unit is specifically used for: Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points; Estimating the motion posture of the target in the visible light image to obtain first-position posture data; Matching the coordinates of the key point with the coordinates of the marker point to determine an ID for the marker point; The coordinates of the marker point, the ID of the marker point and the first pose data are sent to the host computer processing unit.
6. The motion capture system according to claim 5, characterized in that: The host computer processing unit is used for: According to the coordinates and ID of the marker point, the marker point is three-dimensionally reconstructed to obtain the three-dimensional coordinates of the marker point; According to the three-dimensional coordinates of at least three non-collinear marker points, the position and posture of the measured object are solved to obtain second position and posture data; The first posture data and the second posture data are fused and calculated to obtain motion posture data of the measured target.
7. A motion capture method, based on the motion capture system according to any one of claims 1 to 6, characterized in that: include: Perceive the target to be measured and obtain perception data of different modes; The perception data of the different modes are processed to obtain the motion posture data of the measured target.
8. The motion capture method according to claim 7, characterized in that: The sensing of the target to be measured to obtain sensing data of different modes includes: Sensing the measured target in a near-infrared or infrared band, and outputting an image of the near-infrared or infrared band, wherein the image of the near-infrared or infrared band includes at least one marking point; Tracking and spatially locating the marker points in the image of the near-infrared or infrared band to obtain the coordinates of the marker points; Sensing the measured target in the visible light band and outputting a visible light image; Performing key point detection on the visible light image to obtain key point information, wherein the key point information includes coordinates and IDs of the key points; Estimating the motion posture of the target in the visible light image to obtain first-position posture data; The coordinates of the key point and the coordinates of the marker point are matched, an ID is determined for the marker point, and the ID of the marker point is obtained.
9. The motion capture method according to claim 8, characterized in that: The processing of the sensing data of the different modes to obtain the motion posture data of the measured target includes: According to the coordinates and ID of the marker point, the marker point is three-dimensionally reconstructed to obtain the three-dimensional coordinates of the marker point; According to the three-dimensional coordinates of at least three non-collinear marker points, the position and posture of the measured object are solved to obtain second position and posture data; The first posture data and the second posture data are fused and calculated to obtain motion posture data of the measured target.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the motion capture method as described in any one of claims 7 to 9 is implemented.
Citation Information
Cited By
Hand motion capture system and method and readable storage medium
CN121095318A
Motion capture system and motion capture method
WO2026149260A1