Artificial intelligence driven video motion capture and attitude estimation system and method
Through the video motion capture and pose estimation system driven by artificial intelligence, combined with deep learning and optimization algorithms, the problems of high cost, insufficient robustness and poor real-time performance of motion capture and pose estimation in the existing technology are solved, and high-precision and robust motion capture and pose estimation are achieved to meet the real-time processing needs in complex scenarios.
Patent Information
- Application Number
- CN202510159967.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems such as high cost, insufficient robustness in complex scenarios, poor real-time performance, limited pose estimation accuracy and limited multi-objective tracking capabilities in motion capture and pose estimation.
The video motion capture and pose estimation system driven by artificial intelligence is adopted, including video input module, preprocessing module, object detection module, key point detection module, pose estimation module, motion capture module and output module. Through the combination of deep learning and optimization algorithms, the accuracy and robustness of pose estimation are improved, real-time processing needs are met, and multi-objective tracking is achieved.
It significantly improves the accuracy and robustness of pose estimation, meets the real-time processing needs in complex scenarios, and can handle motion capture of multiple people or objects at the same time, enhancing the practicality and efficiency of the system.
Smart Images

Figure CN120126211A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and particularly to an artificial intelligence-driven video action capture and pose estimation system and method. Background Art
[0002] Action capture and pose estimation technology is a technology that analyzes video or sensor data, extracts the motion information of a human body or an object, and reconstructs its pose. Traditional methods rely on expensive hardware devices (such as optical markers, inertial sensors, etc.) and complex calibration processes, with high costs and limited usage scenarios.
[0003] In recent years, methods based on computer vision and deep learning have gradually become mainstream, but they are insufficient in robustness to complex scenarios (such as occlusion, illumination changes), have poor real-time performance, are difficult to meet the high frame rate requirements, and the pose estimation accuracy is limited by the diversity and quality of training data, with limited tracking and recognition capabilities in multi-object scenarios.
[0004] Therefore, we propose an artificial intelligence-driven video action capture and pose estimation system and method. Summary of the Invention
[0005] The present invention mainly solves the technical problems existing in the above-mentioned prior art, and provides an artificial intelligence-driven video action capture and pose estimation system and method.
[0006] To achieve the above object, the present invention adopts the following technical solutions. An artificial intelligence-driven video action capture and pose estimation system includes a video input module, a preprocessing module, an object detection module, a key point detection module, a pose estimation module, an action capture module, and an output module. The video input module includes a high-definition camera interface and a network video stream receiving module, which can be compatible with a variety of video formats and resolutions, ensuring the diversity and flexibility of input data. The video input module is used to receive video streams or image sequences.
[0007] Preferably, the preprocessing module includes a noise filter, an image super-resolution algorithm, and a frame rate interpolation algorithm, effectively improving the video quality and providing high-quality input for subsequent processing.
[0008] Preferably, the preprocessing module is responsible for denoising, resolution enhancement, and frame rate optimization of the input video.
[0009] Preferably, the object detection module includes the YOLOv5 object detection algorithm, which can accurately and quickly locate the objects in the video and maintain a high detection rate even in complex scenarios.
[0010] Preferably, the object detection module detects human or object targets in the video based on a deep learning model.
[0011] Preferably, the key point detection module includes an adaptive key point detection algorithm, which can dynamically adjust the positions of key points, effectively cope with occlusion and illumination changes, and improve the accuracy and robustness of key point detection.
[0012] Preferably, the key point detection module uses a convolutional neural network or a graph convolutional network to extract the key point information of the target.
[0013] Preferably, the pose estimation module includes a Transformer model, which can accurately capture the dynamic pose changes of the target and maintain high accuracy even under fast movement or complex poses.
[0014] Preferably, the pose estimation module optimizes the key point sequence through a spatio-temporal modeling algorithm to estimate the pose of the target.
[0015] Preferably, the motion capture module includes an inverse dynamics algorithm and a Kalman filtering algorithm, which can smooth the motion sequence, reduce noise, and improve the coherence and authenticity of motion capture.
[0016] Preferably, the motion capture module combines a kinematic model with an optimization algorithm to generate the motion sequence of the target.
[0017] Preferably, the output module includes a 3D model, motion data, and a visualization interface.
[0018] Preferably, the output module displays the motion capture results in the form of a 3D model, motion data, or a visualization interface, facilitating intuitive understanding and analysis by the user.
[0019] A method for video motion capture and pose estimation driven by artificial intelligence, including the above-mentioned system for video motion capture and pose estimation driven by artificial intelligence, specifically comprising the following steps:
[0020] The first step: Video preprocessing: Input video data and perform preprocessing;
[0021] The second step: Object detection: Use an object detection model to locate the target in the video;
[0022] The third step: Key point extraction: Extract the key point information of the target to generate a preliminary pose estimation;
[0023] The fourth step: Pose estimation: Optimize the pose estimation result through spatio-temporal modeling;
[0024] The fifth step: Motion capture: Combine a kinematic model to generate a motion sequence;
[0025] The sixth step: Result output: Output the motion capture results and visualize them.
[0026] The present invention provides an artificial intelligence-driven video action capture and pose estimation system and method, which has the following beneficial effects:
[0027] 1. The artificial intelligence-driven video action capture and pose estimation system and method integrate a video input module, a preprocessing module, an object detection module, a key point detection module, a pose estimation module, an action capture module, and an output module. By combining deep learning and optimization algorithms, the accuracy of pose estimation is improved. With lightweight models and hardware acceleration technologies, the real-time processing requirements are met, and it can adapt to complex scenarios such as occlusion and lighting changes, and can simultaneously process the action capture of multiple people or objects.
[0028] 2. The artificial intelligence-driven video action capture and pose estimation system and method significantly improve the pose estimation accuracy and robustness in complex scenarios through the application of an adaptive key point detection algorithm and a Transformer model.
[0029] 3. The artificial intelligence-driven video action capture and pose estimation system and method further optimize the smoothness and coherence of the action sequence, reduce noise interference, and enhance the realism and credibility of action capture by introducing inverse dynamics algorithms and Kalman filtering algorithms.
[0030] 4. The artificial intelligence-driven video action capture and pose estimation system and method achieve full-process automated processing from video input to action capture result output through modular system design and innovative technical solutions. It not only improves the accuracy and real-time performance of action capture, but also significantly enhances the robustness in complex scenarios, providing strong technical support for fields such as sports training, medical rehabilitation, film and television production, and virtual reality.
[0031] 5. The artificial intelligence-driven video action capture and pose estimation system and method can accurately identify and track multiple targets in the video through the built-in multi-target tracking mechanism in the system. Even when there is occlusion or crossing between targets, it can maintain a stable tracking effect, and can simultaneously process the action capture of multiple people or objects, greatly improving the practicality and efficiency of the system. Brief Description of the Drawings
[0032] Figure 1 It is the system module diagram of the present invention;
[0033] Figure 2 It is the method flow chart of the present invention. Detailed Embodiments
[0034] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only exemplary. For those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained by extending the provided drawings.
[0035] The structures, ratios, sizes, etc. illustrated in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope that can be covered by the technical content disclosed in the present invention.
[0036] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0037] In the description of the embodiments of the present invention, it should be noted that the orientation or positional relationships indicated by the terms "center", "upper", "lower", "inner", "outer", "side", etc. are based on the orientation or positional relationships shown in the drawings, or the orientation or positional relationships in which the products of the invention are usually placed during use. This is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0038] In the description of the embodiments of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "set", "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present invention can be understood according to specific circumstances.
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present invention.
[0040] Example 1: An artificial intelligence-driven video action capture and pose estimation system, as Figure 1 shown, includes a video input module, a preprocessing module, an object detection module, a key point detection module, a pose estimation module, an action capture module, and an output module. The video input module is used to receive video streams or image sequences. The video input module includes a high-definition camera interface and a network video stream receiving module, which can be compatible with various video formats and resolutions, ensuring the diversity and flexibility of input data. The preprocessing module is responsible for denoising, resolution enhancement, and frame rate optimization of the input video. The preprocessing module includes a noise filter, an image super-resolution algorithm, and a frame rate interpolation algorithm, effectively improving the video quality and providing high-quality input for subsequent processing. The object detection module detects human or object targets in the video based on a deep learning model. The object detection module includes the YOLOv5 object detection algorithm, which can accurately and quickly locate the targets in the video and maintain a high detection rate even in complex scenes. The key point detection module uses a convolutional neural network or a graph convolutional network to extract the key point information of the target. The key point detection module includes an adaptive key point detection algorithm, which can dynamically adjust the key point positions, effectively coping with occlusion and illumination changes, and improving the accuracy and robustness of key point detection. The pose estimation module optimizes the key point sequence through a spatio-temporal modeling algorithm to estimate the pose of the target. The pose estimation module includes a Transformer model, which can accurately capture the dynamic pose changes of the target and maintain high accuracy even in fast motion or complex poses. The action capture module combines a kinematic model with an optimization algorithm to generate the action sequence of the target. The action capture module includes an inverse dynamics algorithm and a Kalman filtering algorithm, which can smooth the action sequence, reduce noise, and improve the coherence and authenticity of action capture. The output module includes a 3D model, motion data, and a visualization interface. The output module displays the action capture results in the form of a 3D model, motion data, or a visualization interface, facilitating intuitive understanding and analysis by users. By integrating the video input module, the preprocessing module, the object detection module, the key point detection module, the pose estimation module, the action capture module, and the output module, and combining deep learning with optimization algorithms, the accuracy of pose estimation is improved. By using lightweight models and hardware acceleration technologies, the real-time processing requirements are met, and it can adapt to complex scenes such as occlusion and illumination changes, and can simultaneously process the action capture of multiple people or objects.
[0041] Example 2: On the basis of Example 1, as Figure 1As shown, the preprocessing module is responsible for denoising, resolution enhancement, and frame rate optimization of the input video. The preprocessing module includes a noise filter, an image super-resolution algorithm, and a frame rate interpolation algorithm, which effectively improve the video quality and provide high-quality input for subsequent processing. The object detection module detects human or object targets in the video based on a deep learning model. The object detection module includes the YOLOv5 object detection algorithm, which can accurately and quickly locate the targets in the video and maintain a high detection rate even in complex scenarios. Through the application of the adaptive key point detection algorithm and the Transformer model, the accuracy and robustness of pose estimation in complex scenarios are significantly improved.
[0042] Embodiment 3: On the basis of Embodiment 1 and Embodiment 2, as Figure 1 shown, the key point detection module uses a convolutional neural network or a graph convolutional network to extract the key point information of the target. The key point detection module includes an adaptive key point detection algorithm, which can dynamically adjust the key point positions, effectively handle the problems of occlusion and illumination changes, and improve the accuracy and robustness of key point detection. The pose estimation module optimizes the key point sequence through a spatio-temporal modeling algorithm to estimate the pose of the target. The pose estimation module includes a Transformer model, which can accurately capture the dynamic pose changes of the target and maintain high accuracy even in fast motion or complex poses. By introducing the inverse dynamics algorithm and the Kalman filtering algorithm, the smoothness and coherence of the action sequence are further optimized, the noise interference is reduced, and the realism and credibility of action capture are improved.
[0043] Embodiment 4: On the basis of Embodiment 1, Embodiment 2, and Embodiment 3, as Figure 1 shown, the action capture module combines a kinematic model with an optimization algorithm to generate the action sequence of the target. The action capture module includes an inverse dynamics algorithm and a Kalman filtering algorithm, which can smooth the action sequence, reduce noise, and improve the coherence and realism of action capture. The output module includes a 3D model, motion data, and a visualization interface. The output module displays the action capture results in the form of a 3D model, motion data, or visualization interface, which is convenient for users to intuitively understand and analyze. Through the modular system design and innovative technical solutions, the full-process automatic processing from video input to action capture result output is realized. It not only improves the accuracy and real-time performance of action capture, but also significantly enhances the robustness in complex scenarios, providing strong technical support for fields such as sports training, medical rehabilitation, film and television production, and virtual reality. Through the built-in multi-target tracking mechanism of the system, multiple targets in the video can be accurately identified and tracked, and stable tracking effects can be maintained even when the targets are occluded or crossed, enabling the simultaneous processing of the action capture of multiple people or objects, which greatly improves the practicality and efficiency of the system.
[0044] Example 5: Based on Example 1, Example 2, Example 3, and Example 4, as Figure 2 shown, the method for video action capture and pose estimation driven by artificial intelligence specifically includes the following steps:
[0045] The first step: Video preprocessing: Input video data and perform preprocessing;
[0046] The second step: Object detection: Use an object detection model to locate the objects in the video;
[0047] The third step: Key point extraction: Extract the key point information of the objects and generate a preliminary pose estimation;
[0048] The fourth step: Pose estimation: Optimize the pose estimation result through spatio-temporal modeling;
[0049] The fifth step: Action capture: Combine a kinematic model to generate an action sequence;
[0050] The sixth step: Result output: Output the action capture result and visualize it.
[0051] Working principle of the present invention: Input data is obtained through a camera or video file, and image enhancement techniques (such as super-resolution reconstruction) are used to improve the video quality. The YOLOv5 model is adopted to detect targets, and the HRNet is used to extract key point information. The Transformer model is used to perform spatio-temporal modeling on the key point sequence to optimize the pose estimation result. The kinematic model is combined to generate an action sequence, which is output in the form of a 3D model or motion data. The artificial intelligence-driven video action capture and pose estimation system and method mainly rely on the integrated application of deep learning and optimization algorithms. After the video input module receives the video stream or image sequence, the preprocessing module intervenes first. By means of denoising, resolution enhancement, and frame rate optimization, etc., the video quality is effectively improved, laying a solid foundation for subsequent processing. Subsequently, the target detection module uses advanced deep learning models, such as the YOLOv5 target detection algorithm, to quickly and accurately locate the targets in the video. No matter how complex the environment is, a high detection rate can be maintained. The key point detection module adopts a convolutional neural network or a graph convolutional network to dynamically adjust the key point positions, effectively coping with challenges such as occlusion and illumination changes, and improving the accuracy and robustness of key point detection. Immediately afterwards, the pose estimation module accurately captures the dynamic pose changes of the targets through the Transformer model and spatio-temporal modeling algorithms, optimizing the key point sequence, and maintaining high accuracy even under fast motion or complex poses. The action capture module combines the inverse dynamics algorithm and the Kalman filtering algorithm to further optimize the smoothness and coherence of the action sequence, reduce noise interference, and enhance the realism and credibility of action capture. Finally, the output module displays the action capture results in the form of a 3D model, motion data, or a visualization interface, facilitating the user's intuitive understanding and analysis. Through the modular design and innovative technical solutions, the entire system realizes the full-process automatic processing from video input to action capture result output, not only improving the accuracy and real-time performance of action capture, but also significantly enhancing the robustness in complex scenarios. Such an intelligent video action capture and pose estimation system and method provide strong technical support for fields such as sports training, medical rehabilitation, film and television production, and virtual reality, promoting the rapid development and application of related technologies.
[0052] The above has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. Artificial intelligence driven video motion capture and posture estimation system, characterized by: It includes a video input module, a preprocessing module, a target detection module, a key point detection module, a posture estimation module, a motion capture module and an output module. The video input module includes a high-definition camera interface and a network video stream receiving module.
2. The artificial intelligence driven video motion capture and posture estimation system according to claim 1, characterized in that: The preprocessing module includes a noise filter, an image super-resolution algorithm and a frame rate interpolation algorithm.
3. The artificial intelligence driven video motion capture and posture estimation system according to claim 1, characterized in that: The target detection module includes the YOLOv5 target detection algorithm.
4. The artificial intelligence driven video motion capture and posture estimation system according to claim 1, characterized in that: The key point detection module includes an adaptive key point detection algorithm.
5. The artificial intelligence driven video motion capture and posture estimation system according to claim 1, characterized in that: The posture estimation module includes a Transformer model.
6. The artificial intelligence driven video motion capture and posture estimation system according to claim 1, characterized in that: The motion capture module includes an inverse dynamics algorithm and a Kalman filter algorithm.
7. The artificial intelligence driven video motion capture and posture estimation system according to claim 1, characterized in that: The output module includes a 3D model, motion data and a visualization interface.
8. A method for video motion capture and posture estimation driven by artificial intelligence, characterized in that: The artificial intelligence-driven video motion capture and posture estimation system comprising any one of claims 1 to 7 specifically comprises the following steps: Step 1: Video preprocessing: input video data and preprocess it; Step 2: Target detection: Use the target detection model to locate the target in the video; Step 3: Key point extraction: Extract the key point information of the target and generate a preliminary pose estimate; Step 4: Posture estimation: Optimize the posture estimation results through spatiotemporal modeling; Step 5: Motion capture: Generate motion sequences in combination with kinematic models; Step 6: Result output: Output the motion capture results and visualize them.