Mark-free vaulting horse image processing and intelligent motion capturing system based on computer vision technology

By using multiple cameras and image processing devices, combined with machine vision and deep learning technologies, the problem of accuracy in capturing vault postures in complex backgrounds was solved, achieving label-free high-precision motion capture and improving the scientific nature and safety of vault training.

CN121963053APending Publication Date: 2026-05-01CHINA INST OF SPORT SCI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA INST OF SPORT SCI
Filing Date
2026-01-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately capture the vaulter's posture in complex contexts, especially with complex movements and high computational loads, and sensors may compromise athlete safety.

Method used

Employing multiple cameras and image processing and intelligent motion capture devices, the system utilizes machine vision and deep learning technologies to identify human joints and perform 3D reconstruction. By optimizing video data through a calibration module, it achieves label-free motion capture.

Benefits of technology

High-precision video acquisition and key data extraction were achieved in complex environments, improving the scientific level of vault training and reducing the risk of athlete injury.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963053A_ABST
    Figure CN121963053A_ABST
Patent Text Reader

Abstract

The invention relates to a vaulting horse mark-free image processing and intelligent motion capturing system based on a computer vision technology, and belongs to the field of computer image processing, and the system comprises a plurality of cameras which are used for collecting vaulting horse motion video images of a moving target at a fixed point; the image processing and intelligent motion capture device respectively extracts feature points from different camera video images, adopts machine vision and deep learning technology to automatically identify human body articulation points in motion, tracks motions in motion, analyzes the motions in real time, reconstructs and obtains three-dimensional space coordinate information of a moving target, and carries out real-time tracking on the three-dimensional space coordinate information of the moving target. And performing optimization processing on the three-dimensional space coordinate information to obtain motion capture data. The system provided by the invention can realize integrated operation of video acquisition, motion capture and key data extraction under a large-view-field complex background.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a markerless image processing and intelligent motion capture system for vaulting horses based on computer vision technology. Background Technology

[0002] With the continuous development of competitive sports worldwide and the constant improvement of training levels, human physical capabilities are approaching their limits, and the focus of competition in sports is constantly shifting. The factors determining the outcome of matches are gradually shifting from athletes' innate talents to the scientific nature of training. In the fierce competition of contemporary high-level competitive sports, the power and role of science and technology are becoming increasingly prominent, and improving athletes' competitive level through scientific training has become a consensus in the sports world.

[0003] For the vault event in gymnastics, the key challenges in training and assistance lie mainly in the video measurement environment and the vault technique itself. The impact of the training or competition environment on video measurement is reflected in low signal-to-noise ratios due to factors such as lighting, and the presence of many objects in the background that interfere with the athlete's recognition. The vault technique itself consists of multiple directions and various tumbling and turning movements, and the entire movement is short and fast, making visual capture difficult.

[0004] The existing methods for solving human pose in motion scenarios have the following drawbacks: 1) Pose estimation in complex backgrounds In complex contexts, such as complex competition environments or the presence of distracting figures in the background, athletes' body postures are more difficult to discern than in ideal conditions, making it easier for joint point analysis results to be incorrect.

[0005] 2) Pose estimation for complex movements For certain specific movements, such as somersaults or twists in a tucked or bent-over position, general algorithms may not provide accurate posture estimation results. This is especially common in vault training.

[0006] 3) The computation time of neural network models.

[0007] Since neural network models are computationally very expensive, they are the main performance bottleneck when deploying algorithms. Furthermore, pose resolution requires processing data from multiple cameras simultaneously, which further increases latency. Therefore, the model's latency must be optimized during the final deployment of the algorithm.

[0008] 4) Damage risk prevention and control For sports like gymnastics, which involve complex techniques, fast movements, and high risks, the inability to use motion capture devices such as sensors attached to the athlete's body during training would not only affect the athlete's natural and realistic training but also increase the risk of injury.

[0009] Therefore, developing a device that can collect and reproduce the real data of vault technique movements to the greatest extent possible, and improve the scientific level of vault training, is one of the technical problems that needs to be solved. Summary of the Invention

[0010] The purpose of this invention is to at least address one of the aforementioned technical deficiencies.

[0011] Therefore, the purpose of this invention is to propose a markerless image processing and intelligent motion capture system for vaulting horses based on computer vision technology, which can realize integrated operation of video acquisition, motion capture and key data extraction under complex backgrounds with a large field of view.

[0012] To achieve the above objectives, embodiments of the present invention provide a markerless image processing and intelligent motion capture system for vaulting horses based on computer vision technology, comprising: multiple cameras and an image processing and intelligent motion capture device, wherein... The multiple cameras are used to capture video images of the vaulting motion of the moving target at fixed points; The image processing and intelligent motion capture device is used to extract feature points from video images from different cameras, automatically identify human joints in motion using machine vision and deep learning technologies, track and analyze movements in real time, reconstruct and obtain the three-dimensional spatial coordinate information of the moving target, and optimize the three-dimensional spatial coordinate information to obtain motion capture data. The image processing and intelligent motion capture device includes: a calibration module, an acquisition module, an analysis module, a two-dimensional posture analysis unit, and a human posture filtering unit. The calibration module is used to collect calibration videos of the calibration rod swinging at each camera position and verify whether the calibration results meet expectations; if not, the calibration video is corrected. The acquisition module is used to acquire video data captured by the camera, and to perform data cropping and calculation on the video data to obtain the kinematic data from the most recently acquired video data; The analysis module is used to analyze the dynamic parameters, key parameters, and motion trajectories of video data; The two-dimensional posture analysis unit identifies the human body outline from the two-dimensional image of the moving target captured by the camera; detects the key points of the human skeleton, analyzes the coordinates of the key points of the human skeleton; and matches the posture data of the target athlete in the 2D posture of multiple perspectives at the same time, further analyzes and calculates the joint position of the human body with the three-dimensional fusion algorithm, generates the three-dimensional spatial coordinates of the key points of the human skeleton, and obtains motion capture data. The human posture filtering unit is used to smooth the calculated three-dimensional spatial coordinates of key points of the human skeleton, reduce high-frequency components in the data, remove jitter and abrupt changes in the data, and obtain filtered motion capture data.

[0013] Furthermore, the multiple cameras are calibrated to obtain camera structural parameters, internal parameters, and distortion coefficients for three-dimensional spatial positioning; The calibration of the structural parameters among the multiple cameras needs to ensure that the multiple cameras take pictures of the same calibration rod at the same time; by synchronously acquiring images of the moving target, the images of the moving target in different cameras at the same time are obtained.

[0014] Furthermore, the multiple cameras include 8 cameras, located at 8 video information acquisition positions; the core acquisition area is the area from the take-off board to the landing area; two cameras are placed in a group, respectively at the middle and rear section of the track, the vaulting apparatus, the middle of the landing area on both sides, and at appropriate positions along the longitudinal extension line at the end of the field; wherein, at least 3 other cameras can be identified in the frame of each camera position.

[0015] Furthermore, the calibration module performs calibration correction by: decomposing the calibration video into images, dragging and saving the calibration box frame by frame, recalculating the calibration after each frame is modified, verifying the calibration result again, and confirming the calibration is complete if the calibration result is ideal.

[0016] Furthermore, the acquisition module performs data cropping on the video data, including: selecting the scene to be cropped in the interface, waiting for the video format conversion to complete, and then cropping the video segment calculated according to the required precise data. The acquisition module performs data calculations on the video data, including: after completing the data cropping operation, it will automatically calculate the kinematic data of the most recently acquired data or select the data file to be calculated according to the user's calculation needs.

[0017] Furthermore, the two-dimensional pose parsing unit uses deep learning technology to train two neural network models, including a human target detection model and a human pose estimation model. The human target detection model is used to detect the human body contour of the athlete from the image; the human pose estimation model is used to analyze the detected athlete coordinates and calculate the image coordinates of 25 joint points of the human body.

[0018] Furthermore, the analysis module performs dynamic parameter analysis on the video data, including setting the truncation frequency, the start frame and the end frame of the analysis segment, and analyzing the real-time changes in the kinematic data of the acquired object. The analysis module performs key parameter analysis on the video data, including: analyzing the kinematic analysis indicators of the vault, adjusting the number of frames in which the key parameters occur based on the calculation results and the actual situation of the video, confirming the time when the key kinematic parameters occur, inputting the corresponding frame number into the corresponding wireframe, and the key parameters will change to the value at the corresponding time, and finally outputting a key parameter analysis report. The analysis module performs motion trajectory analysis on video data, including: selecting key human points to be analyzed, projecting their motion paths onto various viewpoints, supporting switching between different viewpoints to view the motion paths of key points, and outputting a motion trajectory coordinate report.

[0019] Furthermore, it also includes: a basic settings module and a region drawing unit, wherein, The basic settings module is used to set the preset parameters of the camera, multiple calibration parameters of the calibration module, multiple acquisition parameters of the acquisition module, and multiple analysis parameters of the analysis module. The region drawing unit is used to customize the data acquisition region under the user's operation, including: custom drawing the acquisition and calculation region and setting the automatic acquisition region.

[0020] Furthermore, it also includes a comparison unit, which is used to compare and analyze the kinematic data and video footage of the same or different acquisition objects in different videos.

[0021] Furthermore, the image processing and intelligent motion capture device is also used to present the motion capture data in a visual manner and support the export of vaulting technical analysis reports, including: displaying the three-dimensional reconstructed virtual skeleton model and the recorded video synchronously or separately, supporting forward and backward slow motion playback or frame-by-frame playback, fusion and comparison of two-dimensional video images and three-dimensional models, comparison of three-dimensional models on the same screen, comparison of data curves on the same screen, and providing evaluation and suggestions through sports biomechanics datasets and cluster analysis.

[0022] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows: 1. This invention realizes the integrated operation of video acquisition, motion capture and key data extraction under complex backgrounds with a large field of view.

[0023] 2. This invention does not require Mark point analysis. It uses an optical camera to automatically identify and track key points and limb segments of the human body, employs a high-frequency synchronous multi-view camera to collect human motion data, and then uses a deep learning algorithm to perform two-dimensional human posture detection on the synchronous video. Finally, it uses computer geometry principles to estimate three-dimensional human posture and reconstructs the three-dimensional posture representation of the moving target.

[0024] 3. This invention corrects human posture errors using temporal and spatial information, eliminating outliers and inserting estimates based on motion information. It then utilizes inverse kinematics to re-filter and correct human joints and bones, thereby establishing more realistic human motion.

[0025] 4. The markerless testing method employed in this invention does not cause physical or psychological stress to athletes during the testing process. It maximizes the collection and reproduction of authentic data, providing quantifiable key parameters for vault technique analysis, thus improving the scientific level of vault training and enhancing its quality and efficiency. With continuous in-depth analysis and mining of competition and training data, this system will provide gymnasts with more scientific and accurate data for specialized technical training, promoting the continuous improvement of scientific gymnastics training.

[0026] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0027] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a structural diagram of a markerless image processing and intelligent motion capture system for vaulting horses based on computer vision technology, according to an embodiment of the present invention. Figure 2 This is a system hardware configuration diagram of a markerless image processing and intelligent motion capture system for vaulting horses based on computer vision technology, according to an embodiment of the present invention. Figure 3 This is a flowchart illustrating the workflow of a markerless image processing and intelligent motion capture system for vaulting horses based on computer vision technology, according to an embodiment of the present invention. Figure 4 This is a schematic diagram of the machine station layout according to an embodiment of the present invention; Figure 5 This is a diagram of the system startup interface according to an embodiment of the present invention; Figure 6 To draw a unit interface diagram for a region according to an embodiment of the present invention; Figure 7 This is a diagram of the analysis module - offline calibration visualization window according to an embodiment of the present invention; Figure 8 This is a diagram of the dynamic parameter analysis interface of the analysis module according to an embodiment of the present invention; Figure 9 This is a diagram of the key parameter analysis interface of the analysis module according to an embodiment of the present invention; Figure 10a and Figure 10bThese are, respectively, the recording interface from the right front view and the analysis module's interface for analyzing the motion trajectory of the recorded image, according to an embodiment of the present invention. Detailed Implementation

[0028] Embodiments of the present invention are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0029] like Figure 1 As shown, the vaulting horse unmarked image processing and intelligent motion capture system based on computer vision technology according to an embodiment of the present invention includes: multiple cameras 1, calibration rods 2, switches 3, and image processing and intelligent motion capture devices 4.

[0030] Specifically, such as Figure 2 As shown, multiple cameras 1 are placed on both sides of the vaulting arena to collect video images of the vaulting movement of the moving target at fixed points, and an external synchronous trigger circuit is used to ensure that all cameras 1 collect images simultaneously.

[0031] Multiple cameras 1 are calibrated to obtain their structural parameters, internal parameters, and distortion coefficients for 3D spatial positioning. The calibration of the structural parameters among the multiple cameras 1 requires that all cameras 1 simultaneously capture images of the same calibration rod 2. By synchronously acquiring images of the moving target, images of the moving target within different cameras 1 at the same time are obtained.

[0032] In this invention, the multiple cameras 1 include eight cameras 1, located at eight video information acquisition positions; the core acquisition area is the area from the take-off platform to the landing area; two cameras 1 are grouped together and placed at appropriate positions in the middle and rear section of the track, on both sides of the vaulting apparatus, in the middle of the landing area, and along the longitudinal extension line at the end of the field, forming a U-shaped arrangement around the area from the athlete's take-off to the landing. Each camera position can display images from at least three other camera positions.

[0033] It should be noted that the number and location of cameras can be adjusted adaptively according to the actual situation on site.

[0034] Specifically, the eight cameras 1 configured in this invention all adopt the Z CAM E2 model, and their core parameters are as follows: equipped with a 4 / 3-inch WDR CMOS image sensor (approximately 10.28 million effective pixels), supporting a dynamic range of 13–15 stops (up to 16 stops in WDR mode); the lens interface is an MFT mount. The Z-CAM cameras and their matching lenses, power cables, network cables and connecting cables for the cameras 1, tripods, and synchronizers together constitute a camera group. The eight cameras 1 correspond to eight camera groups.

[0035] In terms of shooting performance, it supports video recording at up to C4K / 4K resolution, with a frame rate range of 23.98–59.94fps. Under H.265 encoding, it can achieve high frame rate shooting of 120 fps in 4K resolution, 160 fps in 4K 2.4:1 mode, and 240 fps in Full HD mode. It supports ZRAW, Apple ProRes 422 series, and multiple encoding formats such as H.265 / H.264, with a maximum color bit depth of 12 bits.

[0036] In terms of input and output, it provides HDMI 2.0 Type A, USB 3.0 Type C, Gigabit Ethernet port, and supports LEMO I / O expansion interface for multi-device synchronization.

[0037] It is compatible with Sony NP-F series batteries and DC 12V 5A external power supplies. The body is made of aluminum alloy and weighs approximately 757g (excluding lens), making it suitable for multi-view simultaneous shooting needs. This invention uses eight high frame rate cameras 1 to capture video images of athletes vaulting at fixed points. The image data needs to be precisely aligned on the time axis to have analytical value and significance. Therefore, an external synchronization trigger circuit is used to ensure that all cameras 1 capture images simultaneously.

[0038] like Figure 4 As shown, the camera position arrangement for camera 1 is as follows: 1) Camera Position Layout: The vault intelligent motion capture system has a total of 8 video information acquisition camera positions, with the approach run direction being both forward and in the depth direction. Two cameras are grouped together and placed in the middle and rear section of the track, on both sides of the vault apparatus, on both sides of the middle section of the landing area, and at appropriate positions along the longitudinal extension line at the end of the field.

[0039] 2) Basic principles: Ensure that at least 3 other camera positions can be identified within the frame of each camera position; the center position of camera 1's frame should overlap with key positions on the site as much as possible.

[0040] The on-site setup is as follows: Three-dimensional reconstruction of limb movements is achieved using video images, requiring the use of images from multiple cameras for three-dimensional spatial point matching. In practical applications, to meet the high-speed and high-precision requirements of vaulting and improve the system's robustness, eight 120fps high-definition cameras are configured to simultaneously acquire video data.

[0041] In accordance with the competition requirements and to meet the deployment conditions within the venue, the hardware system was upgraded. After improvements to the camera 1 bracket, including aesthetic enhancements, safety protection, volume optimization, wiring optimization, and line protection, the camera 1 system was deployed to the competition venue, making it possible to capture complete and clear motion video images.

[0042] The calibration rod 2 is used to calibrate the swing on the vault field and the large field, so that the calibration ball on the calibration rod 2 can be identified and tracked in the field of view of multiple cameras 1, realizing the synchronous calibration of multiple cameras 1, automatically extracting the ball trajectory and calculating internal and external parameters, and realizing the construction of a unified coordinate system for multiple cameras 1.

[0043] Specifically, this invention employs a detachable purple and orange dual-ball calibration rod 2, which performs spatial geometric calibration through a "swing calibration" method. The calibration rod 2 consists of a purple ball and an orange ball, each 18 cm in diameter, with lengths selectable at 120 cm, 170 cm, or 220 cm depending on the site size to meet different shooting range requirements. During use, the operator holds the calibration rod 2 and swings it at large amplitudes and angles to ensure the calibration balls are identified and tracked in the fields of view of multiple cameras. This method allows for rapid synchronous calibration of multiple cameras 1 within five minutes, automatically extracting the ball trajectory and calculating intrinsic and extrinsic parameters, thus achieving the construction of a unified coordinate system for the multi-camera system.

[0044] This invention uses detachable purple and orange clubs for swing calibration on large courses. The entire calibration process is easy and convenient, and can simultaneously calibrate the intrinsic and extrinsic parameters of multiple cameras, completing the process in just 5 minutes. A high-performance computer host equipped with an NVIDIA RTX 3080 / 3090 series graphics card is used to analyze, calculate, output, and display the acquired motion images.

[0045] Switch 3 is used to enable data transmission between multiple cameras 1 and the image processing and intelligent motion capture device 4.

[0046] Specifically, this invention employs a high-performance network switch 3 and a dedicated host as core hardware support. Switch 3 has a switching capacity of 36Gbps and a forwarding capability of 26.8Mpps, providing 16 10 / 100 / 1000Base-T electrical ports and 2 1000Base-X SFP optical ports. It supports multiple modes such as standard switching, network cloning, aggregation uplink, and port isolation. It is configured with an 8K MAC table, adopts store-and-forward mode, has a power supply range of 100–240V AC, and features a fan cooling design to ensure the real-time transmission and stability of large-scale video streaming data.

[0047] The image processing and intelligent motion capture device 4 is used to extract feature points from images from different cameras 1, and to automatically identify the joints of the moving human body using machine vision and deep learning technologies. It tracks the movements and performs real-time analysis to reconstruct and obtain the three-dimensional spatial coordinate information of the moving target. The three-dimensional spatial coordinate information is then optimized to obtain motion capture data. This invention uses image reconstruction technology to reconstruct the three-dimensional spatial coordinate information of the moving target from 2D image data from eight cameras.

[0048] The image processing and intelligent motion capture device 4 of the present invention is implemented by a host computer. That is, the present invention can be configured with a movable integrated flight case, which contains a high-performance computer host computer as the image processing and intelligent motion capture device 4, which analyzes, calculates, outputs and displays the acquired motion images.

[0049] The following describes the host configuration of the image processing and intelligent motion capture device 4: The host system is equipped with an Intel Core i9-10900X processor (10 cores, 20 threads, 3.7GHz), 32GB of DDR4 memory (2400MHz), multiple high-speed NVMe SSDs (including a Samsung SSD 980 500GB system drive), and an NVIDIA GeForce RTX 3090 graphics card (24GB of VRAM). This configuration enables parallel processing of multiple high-resolution video streams and 3D reconstruction calculations. Combined with a 5GbE high-speed network card ensuring low-latency connectivity with Switch 3, it effectively supports the requirements of the markerless vault motion capture system in high frame rate acquisition and real-time pose estimation tasks.

[0050] Without the need for markers or sensors, machine vision and deep learning technologies are used to automatically identify human joints in motion, accurately track high-speed movement, tumbling, and rotation, and analyze and display various kinematic parameters in real time, such as board angle, push angle, height in the air, rotational angular velocity, center of mass, and the trajectory of each joint.

[0051] This invention can also realize the synchronous display or separate display of virtual skeleton models and recorded videos, support forward and backward slow motion playback or frame-by-frame playback, support multi-structure feedback such as fusion comparison of two-dimensional video images and three-dimensional models, on-screen comparison of three-dimensional models, and on-screen comparison of data curves, and provide evaluation and suggestions through sports biomechanics datasets and cluster analysis systems.

[0052] like Figure 3 As shown, this invention first requires the calibration of multiple cameras 1 to obtain the structural parameters, intrinsic parameters, and distortion coefficients of camera 1 for three-dimensional spatial positioning. The calibration of the structural parameters between cameras 1 requires that each camera 1 simultaneously capture images of the same calibration rod 2. By synchronously acquiring images of the moving target, images of the moving target within different cameras 1 at the same time are obtained. Further, computer graphics image processing is performed using image processing and an intelligent motion capture device 4 to extract feature points from images from different cameras 1. The three-dimensional spatial coordinate information of the moving target is obtained using a spatial positioning algorithm. Finally, after filtering and other optimizations, motion capture data (i.e., motion capture results) is obtained, presented in a visual manner, and supports the export of vaulting technique analysis reports. This invention employs image filtering technology to filter the three-dimensional spatial coordinate data to obtain the final motion capture data.

[0053] like Figure 5 As shown, this invention presents the startup interface of the AI-based vaulting intelligent motion capture system. The startup page includes three parameter settings: resolution, frame rate, and whether to restart. Users can select "1920*1080", "120", or "no restart" respectively, and then click the "Start" or "Offline Start" button.

[0054] The functions of the image processing and intelligent motion capture device 4 are explained below: 1) Automatic identification and tracking Automatically identify and track human joints without the need for markers or wearing sensors.

[0055] 2) Skeletal model reconstruction The 3D skeletal model was reconstructed in 360°, and the model has 25 joints.

[0056] 3) Rapid calibration The intrinsic and extrinsic parameters of camera 1 were calibrated within 5 minutes.

[0057] 4) Camera 1 model and distortion processing Distortion correction processing is performed on the photos taken by each camera 1 to obtain the true spatial coordinates and achieve high precision.

[0058] 5) Virtual display The virtual skeletal model can be displayed synchronously with or separately from the recorded video.

[0059] 6) Slow motion and frame-by-frame playback Motion videos and models support slow-motion playback, allowing for frame-by-frame playback forward and backward.

[0060] 7) Comparison function The system integrates and compares 2D video images with 3D models, compares 3D models on the same screen, and compares data curves on the same screen, achieving multi-structure feedback.

[0061] 8) Display of kinematic data For vaulting, the system automatically identifies key joints during training, such as the steps on the board, the first flight, and the landing, and displays the real-time changes in angles and angular velocities of the center of gravity, shoulder angle, elbow angle, hip angle, knee angle, and ankle angle.

[0062] 9) Simulated scoring By combining a motion knowledge base, the system simulates referee scoring and intelligently evaluates the completion of the motion.

[0063] 10) Results Report The system intelligently analyzes video images of athletes' daily training and competition techniques, interprets athletes' kinematic parameters, and generates a vault biomechanical technical analysis report, which is then exported in PDF format.

[0064] Specifically, the image processing and intelligent motion capture device 4 includes: a calibration module 41, an acquisition module 42, and an analysis module 43. The analysis module 43 includes: a region drawing unit 431, a two-dimensional pose analysis unit 432, a comparison unit 433, and a human pose filtering unit 434.

[0065] The calibration module 41 is used to collect calibration videos during the calibration of the calibration rod 2 at each camera position and verify whether the calibration results meet expectations. If they do not meet expectations, the calibration video is corrected. The calibration correction performed by the calibration module 41 includes: decomposing the calibration video into images, dragging and saving the calibration frame frame by frame, recalculating the calibration after each frame is modified, and verifying the calibration results again. If the calibration results are satisfactory, the calibration is confirmed to be complete.

[0066] Specifically, calibration involves assigning values ​​to the actual spatial coordinates of the video image. Calibration must be performed before kinematic data capture, typically using a calibration rod 2 (with spherical ends). System calibration is a crucial step in ensuring the accuracy of system calculations; calibration errors will affect all calculated coordinates. Camera 1 should not be moved or scaled during the initial calibration and video acquisition phases. The interface of calibration module 41 is as follows... Figure 9 As shown, it includes three buttons: "Start Calibration", "Calibration Verification", and "Calibration Correction".

[0067] The interface of calibration module 41 has a prompt at the bottom: If it is the first time using it or you are sure that the parameters of camera 1 have changed, click the "Start Calibration" button; if it is not the first time using it and you are unsure whether the parameters of camera 1 have changed, click the "Calibration Verification" button.

[0068] The initial calibration procedure for calibration module 41 is as follows: 1) Click the "Start Calibration" button in the toolbar at the bottom of the interface to enter the calibration process. Users can adjust the image parameters (shutter speed, ISO, white balance, focus) according to the actual needs of the scene. The adjustment principle is to ensure that the sharpness of all images is sufficient to distinguish the key features of calibration rod 2. Users can click the three buttons at the top of the module interface: "Display Acquisition Area," "Enable Automatic Refresh," and "Refresh," to respectively display the pre-drawn acquisition area, automatically adjust the parameter settings of all cameras 1, and manually refresh the parameter settings of camera 1.

[0069] 2) The calibration process is as follows: ① After ensuring there is no external interference in the area to be acquired (no one enters or moves within the area), click the "Background Acquisition" button and wait 1 to 5 seconds. Once the background acquisition is complete, the interface will automatically switch to X-axis acquisition. ② According to the actual situation, the user places the calibration rod 2 within the shooting frame of camera 1, clicks the "X-axis Acquisition" button, and waits 1 to 5 seconds. Once the X-axis acquisition is complete, the interface will automatically switch to Y-axis acquisition.

[0070] ③ According to the actual site conditions, the user places the calibration rod 2 within the field of view of the camera 1 (it should be perpendicular to the X-axis in the horizontal plane), clicks the "Y-axis acquisition" button, waits for 1 to 5 seconds, and after the Y-axis acquisition is completed, the interface will automatically jump to the swing acquisition.

[0071] ④ After the user clicks the "Start Swing" button, they hold the calibration club 2 in front of each camera 1 and swing it. The swing should ensure that the calibration ball is in the frame. After the swing is completed, click "End Swing".

[0072] ⑤ Click the “Calibration Calculation” button and wait for the calibration calculation to finish. After the calibration calculation is completed, the screen will display the newly created coordinate system, the area to be collected, and the calibration verification results of camera 1.

[0073] ⑥ If the calibration results are satisfactory, click the "End Calibration" button.

[0074] The calibration module 41 performs calibration verification and calibration correction, and the process is as follows: 1) Calibration Verification: After calibration is completed, or to verify the effectiveness of the previous calibration, click the "Calibration Verification" button. Capture images from each camera position in the field environment and check the calibration status. If the calibration is satisfactory, click the "No Calibration Required" button; if the calibration result is unsatisfactory, click the "Recalibrate" button and repeat the calibration operation, or click the "Calibration Correction" button to correct the calibration frame by frame.

[0075] 2) Calibration Correction: If calibration correction is required, the user can click the "Calibration Correction" button, wait for the system to decompose the calibration video into images, drag the calibration box (calibration ball) frame by frame and save it (the arrow keys "↑" and "↓" are the screen frame switching keys). After modifying frame by frame, click the "Calibration Calculation" button to verify the calibration result again. If the calibration result is ideal, click the "No Calibration Required" button.

[0076] It should be noted that calibration correction usually takes a lot of time. It is recommended to use the calibration correction function when the field environment does not allow recalibration, or when manual calibration correction is required.

[0077] The acquisition module 42 is used to acquire video data captured by camera 1, and to perform data cropping and calculation on the video data to obtain the kinematic data from the most recently acquired video data.

[0078] Specifically, after completing the calibration, click the "Acquire" button in the left navigation bar to enter the acquisition module 42. The interface of the video acquisition module 42 includes the following at the bottom: ① Camera 1 parameter options (shutter speed, ISO, white balance, focus), ② "Start Automatic Acquisition" button, ③ "Start Acquisition" button, ④ "Seconds Display" pane (displaying the duration of the currently acquired video), ⑤ "Data Cropping" button, ⑥ "Single Calculation" button, and ⑦ "Select Calculation" button.

[0079] Camera 1 parameter adjustment: Users can adjust the parameters of camera 1 according to the actual environment on site. The adjustment principle is to ensure that the clarity of all images is sufficient to distinguish the external features of the key points of the subject.

[0080] This invention provides both automatic and manual data acquisition methods, allowing users to choose the appropriate method based on their needs. It should be noted that during the first acquisition, the system needs to complete its initialization automatically (estimated to take 20 seconds). Once data acquisition begins, the duration of the currently acquired video will be automatically displayed in the seconds pane.

[0081] 1) Automatic Acquisition: If the user needs and trusts the software's automatic acquisition function and has completed the preset operations for the automatic acquisition area, the user can click the "Start Automatic Acquisition" button. The system will acquire and store the video to the specified folder according to the preset conditions and path (default storage to drive D), and sort it in chronological order from most recent to oldest in subsequent operations.

[0082] 2) Manual Acquisition: If the user chooses manual acquisition, they only need to click the "Start Acquisition" and "Stop Acquisition" buttons as needed. Similarly, the system will acquire and store the videos to the specified folder according to the preset conditions and path (default storage is to drive D), and sort them in chronological order from most recent to oldest in subsequent operations.

[0083] The acquisition module 42 performs data cropping on the video data, including: selecting the scene to be cropped in the interface, waiting for the video format to be converted, and then cropping the video segment according to the required precise data calculation.

[0084] To save system calculation time, users can click the "Data Cropping" button, then locate and select the footage to be cropped in the right-hand pane of the interface. After the software converts the video format, users can further refine the video segment for data calculation as needed, and then click the "OK" button. It should be noted that in addition to the system's real-time captured footage, the system also allows cropping of previously captured footage from folders.

[0085] The acquisition module 42 performs data calculations on the video data, including: after completing the data cropping operation, it will automatically calculate the kinematic data of the most recently acquired data or select the data file to be calculated according to the user's calculation needs.

[0086] After completing the data cropping operation, users can choose to click the "Single Calculation" button or the "Select Calculation" button according to their actual needs. Clicking the "Single Calculation" button will automatically calculate the most recently acquired kinematic data. If users have other calculation needs, they can click the "Select Calculation" button and select the data file to be calculated according to their actual needs. This process can be completed by clicking the "+" button on the right side of the folder to be calculated after locating it. In this frame, users can choose "Recalculate" or "Calculate 3D skeleton only" according to the actual situation. If there is already 2D image data in the data folder, the calculation time can be saved by selecting the "Calculate 3D skeleton only" calculation method. It should be noted that if only one set of data is collected and calculated in the acquisition module 42, the software will automatically jump to the analysis module 43 after the calculation is completed.

[0087] Analysis module 43 is used to analyze the dynamic parameters, key parameters and motion trajectory of video data.

[0088] After completing video capture and calculation, click the "Analyze" button in the left sidebar to enter analysis module 43. Analysis module 43's interface includes three analysis sections: dynamic parameters, key parameters, and motion trajectory. Additionally, this module includes an automatic scoring function. The left half of analysis module 43's interface displays the recorded video from the right-hand perspective, while the right half displays the dynamic parameter analysis interface for that video. The images in the left and right halves of analysis module 43 together constitute the interface used when performing dynamic parameter analysis on the right-hand perspective.

[0089] Note: Users can click the "Import Video" button to open the computer folder, select the recording folder to be analyzed, and click the "OK" button to complete the video import according to the actual analysis needs.

[0090] Offline calibration visualization: To check and confirm the reliability of the calibration during the acquisition process. The software has an "Offline Calibration Visualization" button. Users can click this button to view the calibration status of each camera position (the software automatically triggers a computer media player to open the image). Figure 7 This is a schematic diagram of the offline calibration visualization window for analysis module 43.

[0091] The specific parameters for vault include: dynamic parameters, key parameters, and movement trajectory.

[0092] 1. Dynamic parameters The analysis module 43 performs dynamic parameter analysis on the video data, including setting the truncation frequency, the start and end frames of the analysis segment, and analyzing the real-time changes in the kinematic data of the acquired object.

[0093] like Figure 8 As shown, in the Dynamic Parameter Analysis section, users can choose the following options based on their needs: cutoff frequency (filtering out high-frequency data components), start frame of the analysis segment (the beginning frame of the segment to be analyzed), and end frame (the end frame of the segment to be analyzed). Furthermore, users can analyze the real-time changes in the kinematic data of the collected object, including four items: joint points, joint angles, segment angles, and rotation angles. Each analysis section includes the following functions: 1) In the Joint Analysis section, users can select any key point of the object to be collected and analyze its actual motion speed and its decomposed velocity, acceleration and displacement on the X-axis, Y-axis and Z-axis.

[0094] 2) In the joint angle analysis section, in addition to selecting the combined angle of different joint angles for analysis, users can also select their respective projection angles on the athlete's own sagittal plane, coronal plane, and horizontal plane for analysis.

[0095] 3) In the segment angle analysis section, users can select different segment angles to analyze the projection angles of the athlete's own sagittal, coronal, and horizontal planes.

[0096] 4) The rotation angle analysis column is mainly used to analyze the rotation of the torso of the collected object.

[0097] 5) Users can zoom in or out on a portion of the line chart by moving the mouse over it and scrolling the mouse wheel.

[0098] 6) Users can click the "Export Report" button in the lower right corner of the interface, and the software will output a dynamic parameter analysis report.

[0099] 2. Key parameters Analysis module 43 performs key parameter analysis on video data, including: analyzing kinematic analysis indicators for the vault, adjusting the number of frames in which key parameters occur based on the calculation results and the actual situation of the video footage, confirming the time when key kinematic parameters occur, inputting the corresponding frame number into the corresponding wireframe, and the key parameters will change to the values ​​at the corresponding time, and finally outputting a key parameter analysis report.

[0100] like Figure 9 As shown, the Key Parameter Analysis section includes five sets of kinematic analysis indicators specific to the vault: take-off, first flight, second flight after hand push, and landing. The section interface is shown in the figure. This section includes the following functions: 1) Users can access the following functions via the buttons at the bottom of the module interface: play, pause, slow motion, single-frame forward, single-frame rewind, loop playback of video footage, and projection of virtual skeletons onto the video screen, totaling 7 functions. Users can click the video window at the top of the interface to switch between different viewing angles. 2) Users can adjust the number of frames in which key parameters occur based on the software's calculations and the actual video footage. Users can confirm the timing of key kinematic parameters by forwarding and rewinding frames, inputting the corresponding frame number into the corresponding frame frame; the key parameter will then change to the value at that moment. 3) After completing the calibration and analysis of the key parameters, users can click the "Export Report" button in the lower right corner of the interface; the software will then output a key parameter analysis report.

[0101] 3. Movement trajectory The analysis module 43 performs motion trajectory analysis on the video data, including: selecting key points of the human body to be analyzed, projecting their motion paths onto the screens of various viewpoints, supporting switching between different viewpoints to view the motion paths of key points, and outputting a motion trajectory coordinate report.

[0102] Figure 10a This is the video feed from the right front view, located in the left half of the analysis module 43 interface. Figure 10b This is the motion trajectory analysis interface for the recorded footage, located in the right half of the analysis module interface. Figure 10a and Figure 10b The images together constitute the interface of analysis module 43 when analyzing the motion trajectory from the right front view.

[0103] like Figure 10a and Figure 10b As shown, the main function of the motion trajectory analysis section is as follows: Users can click on the options under the "Joint Trajectory" section on the right side of the interface to select the key points of the human body to be analyzed (each joint point and center of gravity), and the software will project their motion path onto the screen from various viewpoints. Users can click on the video window at the top of the interface to switch between different viewpoints to view the motion path of the key points. Users can click on the "Export Report" button in the lower right corner of the interface, and the software will output a motion trajectory coordinate report.

[0104] The two-dimensional pose analysis unit 432 identifies the human body contour from the two-dimensional image of the moving target captured by the camera 1; detects the key points of the human skeleton, analyzes the coordinates of the key points of the human skeleton; and matches the pose data of the target athlete in the 2D poses of multiple perspectives at the same time. It further analyzes and calculates the joint positions of the human body with the three-dimensional fusion algorithm to generate the three-dimensional spatial coordinates of the key points of the human skeleton and obtain motion capture data.

[0105] The two-dimensional posture analysis unit 432 is the core module of the vaulting horse intelligent motion capture system. This module is responsible for analyzing the joint positions of the human body from the two-dimensional images captured by camera 1. The results of the human body joint analysis are then further analyzed and calculated by the three-dimensional fusion module to form the three-dimensional spatial coordinates of the key points of the human skeleton.

[0106] The two-dimensional pose parsing unit 432 employs deep learning technology to train two neural network models: a human target detection model and a human pose estimation model. The human target detection model is used to detect the human contours of athletes from images. Specifically, the human target detection model uses image classification, image semantic segmentation, and image detection techniques to detect and extract human contours from images, thereby achieving human target detection.

[0107] The human pose estimation model is used to analyze the detected athlete coordinates and calculate the image coordinates of 25 joints of the human body.

[0108] In the analysis of human posture in vault training scenarios, this invention addresses posture estimation under complex backgrounds and for complex movements. During the iterative process, the algorithm model selection referenced the most effective neural network architecture currently available to ensure the model's accuracy and robustness. The algorithm collected a large amount of vault video data from professional athletes and employed various data augmentation methods to train the neural network model, addressing the model's generalization ability. Ultimately, the human posture estimation model achieved a MAP of 0.9886 on the test data.

[0109] To address the computational time consumption issue of neural network models, the algorithm of this invention utilizes specialized hardware acceleration equipment and optimizes the system's computational efficiency through multi-threaded parallel processing. The computational efficiency of the human pose analysis module ultimately reaches 98.5 frames per second, ensuring the system's real-time performance.

[0110] The human posture filtering unit 434 is used to smooth the three-dimensional spatial coordinates of the calculated human skeleton key points, reduce the high-frequency components in the data, remove jitter and abrupt changes in the data, and obtain the filtered motion capture data.

[0111] The vault intelligent motion capture system analyzes motion images captured from multiple perspectives to determine the 3D coordinates of key human body points. However, the system is subject to various noise interferences during the measurement process. Specifically, during camera 1's acquisition, the athlete's high-speed movement can cause some blurring in the image; excessive shooting distance of camera 1 and video compression encoding can lead to image distortion; pixel errors exist in the human posture analysis results; and in cases of severe occlusion by the human body, the acquired data cannot accurately calculate joint coordinates. These various noise effects, along with inherent system errors, result in a certain deviation between the final motion capture data and the athlete's actual movement data, manifesting as a jittery posture in the final visualization.

[0112] To address the issue of data jitter, this invention's algorithm incorporates a human posture filtering unit 434 for optimization, reducing system errors. The vault intelligent motion capture system uses an exponentially weighted average method to smooth the calculated three-dimensional joint coordinates of the human body, reducing high-frequency components in the data and removing jitter and abrupt changes, thereby more accurately reflecting the athlete's physical movement state.

[0113] Data Acquisition Area Delineation and Automatic Data Acquisition: In practical use, the vault intelligent motion capture system often involves multiple athletes training simultaneously. The initial algorithm version randomly selected a target athlete in the field for motion capture. During algorithm iterations, the system added a data acquisition area selection function. After camera 1 is deployed, the user can select a rectangular area in the field; during motion capture, the software will ignore targets outside this rectangular area. Before collecting vault training data, the user can set a specific vault runway as the acquisition area, thereby achieving the goal of tracking a specific athlete in the scene.

[0114] In the initial version of the software, users needed to click the "Start Collection" button on the interface before the vault athlete began their run-up and the "Stop Collection" button after the athlete finished their vault. After collection, the software would calculate the collected motion data. During algorithm iterations, an automatic collection function was added. When the athlete enters the designated collection area, the software automatically triggers a video collection command to begin recording video data. After the athlete completes their vault and lands, the software automatically triggers a collection stop command to save the video data. After the system update, users no longer need to frequently operate the software. Once the software is deployed, simply enabling the automatic collection function allows for unattended recording of data from multiple vault training sessions over a period of time. All data can then be batch-calculated, significantly reducing the manual effort required for data collection and making the software more user-friendly.

[0115] In addition, the present invention also includes: a basic setting module 44, a region drawing unit 431, and a comparison unit 433.

[0116] The basic settings module 44 is used to set the preset parameters of camera 1, multiple calibration parameters of calibration module 41, multiple acquisition parameters of acquisition module 42, and multiple analysis parameters of analysis module 43. The basic settings module 44 includes settings for four parameters: "Camera 1," "Calibration," "Analysis," and "Acquisition." The resolution and frame rate under the Camera 1 section have already been set on the startup page. Users need to recalibrate the preset parameters of camera 1 to confirm that they are consistent with the presets on the startup page.

[0117] The input values ​​for the calibration ball scale factor, calibration parameter (hint, which is an approximate value of the calibration parameter, corresponding to the focal length of camera 1), and calibration rod length (cm) under the calibration section are 0.86, 850, and 120, respectively. The acquisition area is the horizontal plane (custom rectangle) of the space to be acquired. L1 (runway length), L2 (depth of the landing area after takeoff), L3 (width of the left side of the runway), and L4 (width of the right side of the runway) are the distances in the positive direction from the boundary of the acquisition area to the origin of the acquisition area (the 0 coordinate point of the acquired area, which is usually located on the horse).

[0118] The default sport under the Analysis section is "Vault," with the default vault height being 1.25m for women's vault. Users can adjust the vault height parameter by inputting a value as needed. Regarding the preset value for the number of camera 1 constraints, if the number of cameras 1 is even (n), the preset value should be n ÷ 2; if the number of cameras 1 is odd (n), the preset value should be (n - 1) ÷ 2. For example, if there are 7 cameras 1, the preset value here should be 3. The preset value for the filter window size should be 1 / 20 of the frame rate of camera 1. For example, if the frame rate of camera 1 is 120, the preset value here should be 6.

[0119] The gender (male or female) of the subject to be collected and the automatic collection duration (s) under the collection section can be entered as needed based on the actual situation. After entering the automatic collection duration, the system will automatically capture a fixed duration from the moment the subject enters the "collection (monitoring) area" until that moment. For example, if the fixed collection duration is entered as 10s, and the athlete enters the collection area at 11:11:11, the system will automatically capture and save the video from 11:11:11 to 11:11:20.

[0120] It should be noted that system parameter settings should generally remain at the default settings and should not be adjusted further. After completing all settings, the user needs to click the "Confirm Acquisition Area Parameters" button to proceed to the next step.

[0121] The area drawing unit 431 is used to customize the data acquisition area under the user's operation, including: custom drawing the acquisition and calculation area and setting the automatic acquisition area.

[0122] like Figure 6 As shown, after completing the input of basic settings parameters, users can click the "Region Drawing" module on the left to customize the data acquisition area. The region drawing unit 431 includes two functions: custom drawing of the acquisition and calculation area (two-dimensional) and automatic acquisition area editing (two-dimensional).

[0123] 1. Data Acquisition and Calculation Area Settings Users can click the "Custom Drawing" button under the Data Acquisition and Calculation Area section to begin drawing the data acquisition and calculation area. Users can switch between screens by clicking the left mouse button in the right-hand column of the module, and drag the endpoints of the red borders one by one within each screen to edit the data acquisition and calculation area.

[0124] 2. Automatic data collection area settings This function requires two steps to achieve the following: automatically collecting data from a specific area only when the athlete moves in the prescribed direction: 1) Users should click the "Region Drawing" button under the "Automatic Collection Area" section at the bottom of the module interface and draw the automatic collection trigger area.

[0125] 2) After completing the automatic acquisition area drawing, the user should click the "Direction Drawing" button under the Automatic Acquisition Area section. Within the area drawn in the previous step, click sequentially on the athlete's starting point (displayed as a red dot) and the athlete's direction relative to the starting point (displayed as a blue dot; the optimal position for this point is the center of the drawing area). After completing the above operations, the user can click the "Enter Calibration" button to jump to the calibration module 41 and begin calibration.

[0126] The comparison unit 433 is used to compare and analyze kinematic data and video footage of the same or different subjects in different videos. The main function of the comparison unit 433 is to compare and analyze key parameters (kinematic data) and video footage of the same or different subjects in two video clips. Users can select kinematic data for comparison and analysis according to their actual analysis needs. The operation method is as follows: 1) Users can click the "Select Video" button to select different video files in the two frames respectively.

[0127] 2) Users can drag the "Playback Progress" button (red inverted triangle) and the "Single Frame Switch" button (the button below the video frame) to adjust to the screen where analysis will begin. Then, click the "Play" button at the bottom of the interface to start playing the video.

[0128] 3) Users can click the "Camera" button on the right side of the interface to switch between different recording angles.

[0129] 4) Users can click the "Key Parameters" button in the lower right corner of the interface to select key kinematic data such as "center of gravity velocity" and "left hip angle" for comparative analysis as needed.

[0130] It should be noted that the joint angle kinematic data in this module are all resultant angles.

[0131] In this invention, the image processing and intelligent motion capture device 4 is also used to present the motion capture data in a visual manner and support the export of vaulting technique analysis reports.

[0132] This invention supports the synchronous or separate display of 3D reconstructed virtual skeleton models and recorded videos, supports forward and backward slow-motion playback or frame-by-frame playback, fusion and comparison of 2D video images and 3D models, on-screen comparison of 3D models, and on-screen comparison of data curves, and provides evaluation and suggestions through sports biomechanics datasets and cluster analysis.

[0133] The following example, using three competitions, illustrates the data and reports collected by this invention: A total of 1000 sets of video image data were collected from the three competitions, all of which have been delivered for subsequent analysis. Based on the characteristics of vaulting, the data was divided into stages, providing comprehensive and complete kinematic parameters and intelligent scoring, resulting in a vaulting biomechanical technical analysis report. The report includes: 1) Athlete Information: Name, Gender, Height, Weight; 2) Competition Information: Athlete's unit, number, action number, D score, E score; 3) Key parameters: top board angle, second takeoff height, bracing angle, landing angle, top board torso angle; 4) Parameters for each stage: ① Take-off phase: last step length before boarding, take-off time, boarding angle, boarding trunk angle, boarding horizontal / vertical speed, boarding angle, boarding trunk angle, boarding horizontal / vertical speed; ② First flight phase: time, center of gravity speed, shoulder joint rotation angle, hip joint rotation angle; ③ Pushing phase: pushing angle, time difference between left and right hand pushing, left / right shoulder angle at the moment of pushing, left / right elbow angle at the moment of pushing, left / right hip angle at the moment of pushing, pushing time, boarding horizontal / vertical speed, push-off angle, boarding horizontal / vertical speed; ④ Second flight phase: time, height (MAX), flip / rotation angular velocity, flight angle; ⑤ Landing phase: distance, landing angle, landing left / right hip / knee / ankle angle; 5) Center of gravity velocity curve; 6) Center of gravity trajectory curve; 7) Parameter division and technical action definition; 8) Diagnostic opinion.

[0134] The following is a description of the accuracy verification report for this invention: 1. Data Recording The recorded information includes: batch, acquisition time, acquisition time and camera configuration.

[0135] 2. Experimental Data Collection 2.1 Data Content The first batch of data was collected on-site at a gymnastics championship in Chengdu in a certain year, and included a total of 4 sets of data. During the analysis, only the motion capture data from the start of the exercise to the landing was calculated. The data content included: name (time, sequence number, athlete's name, gymnastics move, score), interval, and frame number.

[0136] 2.2 Data Generation of the Vault Intelligent Motion Capture System The system first detects the skeletal key points of all figures in the 2D image using a pose estimation algorithm; then it matches the pose data of the target athlete from multiple 2D poses at the same time; finally, it calculates the 3D spatial coordinates of the human skeletal key points using a 3D fusion algorithm to obtain motion capture data; and then it performs filtering operations on the 3D coordinates to obtain filtered motion capture data.

[0137] 2.3 Ground truth data generation Ground truth data is generated through manual annotation. First, annotators mark the locations of key skeletal points of the target athlete in all two-dimensional images. Then, a three-dimensional fusion algorithm is used to calculate the three-dimensional spatial coordinates of the key skeletal points.

[0138] 3. Analysis of Factors Affecting Error Model prediction error: The main source of data error is the prediction bias of the AI ​​model; Calibration parameter error and camera 1 synchronization error: The 3D fusion algorithm relies on the camera 1 system to calibrate the camera 1 parameters accurately and has good synchronization. Calibration parameter error and camera 1 synchronization error will affect both ground truth and motion capture data, resulting in a decrease in the reliability of evaluation indicators. Manual annotation bias: Manually annotated keypoints tend to fall within a small area near the center of the joint. For larger joints, the annotation bias will be greater, and the error will generally be larger as well. In addition, manual annotation results may occasionally contain annotation errors, which may slightly affect the analysis results. Filter parameter selection: The motion capture results are closely related to the selection of filter parameters. Currently, based on experience, the algorithm selects a filter parameter of 4. The selection of filter parameters is mainly related to the joint movement speed. Choosing more suitable filter parameters can reduce noise in the motion capture data and reduce errors.

[0139] 4. Data Error Analysis 4.1 Key point position error The vault intelligent motion capture system uses data containing the three-dimensional coordinates of 25 joints on the vaulting body. This data is compared with Groundtruth data to calculate the spatial distances, X-axis distances (direction of movement along the track), Y-axis distances (perpendicular to the track), and Z-axis distances (vertical direction) of each joint in each frame. The average result across all frames is then calculated. The first batch consists of four sets of data, showing the average error before and after filtering. The average error of the 25 joints is statistically analyzed, with the horizontal line representing the average value. The first two sets of data are from female athletes, the last two from male athletes, and the last set is the average of all data.

[0140] 4.2 Error of the Multiple Correlation Coefficient (CMC) The CMC metric was calculated using a set of data as an example. The results are shown in Table 1. Table 1 CMC Indicators Group 1 X-axis Y-axis Z-axis right shoulder 0.999985 0.978092 0.999760 right elbow 0.999973 0.993604 0.999662 right wrist 0.999919 0.987944 0.998455 left shoulder 0.999984 0.985103 0.999832 left elbow 0.999983 0.996466 0.999786 left wrist 0.999980 0.997846 0.999823 Right buttock 0.999987 0.959602 0.999842 right knee 0.999987 0.984821 0.999909 Right ankle 0.999986 0.988405 0.999908 Left buttock 0.999988 0.986161 0.999901 left knee 0.999991 0.993355 0.999873 left ankle 0.999994 0.996412 0.999926 left big toe 0.999972 0.994734 0.999875 left little toe 0.999983 0.994478 0.999879 left heel 0.999988 0.994381 0.999905 right big toe 0.999964 0.991828 0.999901 right little toe 0.999960 0.989091 0.999889 right heel 0.999976 0.989356 0.999884 nose 0.999984 0.964893 0.999841 Spine 1 0.999985 0.862839 0.999839 Spine 2 0.999994 0.973608 0.999873 right eye 0.999971 0.946588 0.999837 Left eye 0.999973 0.967840 0.999842 Right ear 0.999983 0.949441 0.999776 left ear 0.999972 0.921796 0.999778 average value 0.999978 0.975547 0.999792 Based on the data analysis results of the first batch of vault intelligent motion capture systems, the following conclusions were drawn: The proportion of test samples with a multiple correlation coefficient (CMC) greater than 0.95 between the automatically identified joint coordinates and the corresponding manually labeled joint coordinates of the AI-based vaulting intelligent motion capture system was 97.3% (292 / 300).

[0141] Aside from foot key points, filtered data generally shows smaller errors; however, foot key points show larger errors after filtering, particularly noticeable in the men's group. This is because the foot's movement speed is significantly greater than other joint points during the vault, and the filtering parameters chosen by the algorithm are more suitable for the movement of other body joints. For vaulting, more appropriate filtering parameters can be selected for foot key points to avoid increasing errors after filtering. Leg key points (hip, knee, ankle) show significantly smaller errors than other key points because lower body joints are generally less obstructed and easier to distinguish. By adding camera 1 to reduce arm obstruction and by improving the quality of acquired images to avoid blurring of the feet and facial features, theoretically, errors in arm, foot, and facial key points can be reduced.

[0142] 4.3 Scoring Error Because the vault movement databases for men and women are different and each has its own characteristics, the deduction module divides the movement recognition database into a men's movement database and a women's movement database. The vault deduction recognition module is divided into two parts: D score and E score. The D score part analyzes the three-dimensional posture features of the vault athlete at each stage (before mounting the board, mounting the board, leaving the board, mounting the horse, leaving the left and right hands on the horse, landing, etc.) (including the body orientation at the moment of mounting, the angle of self-rotation, the angle of body somersault, the type of somersault, tuck / bow / straight body, etc.), and then identifies the athlete's movement number and corresponding difficulty score. Table 2 shows the statistics of the calculation results of the D score part. Some data contain calculation errors. The main reasons for recognition failure are: 1) The athlete did not complete the movement; 2) The three-dimensional posture data is incorrect, resulting in a large deviation in the calculated rotation angle; 3) The understanding of the movement corresponding to the movement number is not complete; 4) The pattern recognition of each stage of the vault is not accurate enough; 5) Some postures do not meet the corresponding requirements, such as tuck and bend, which lead to recognition errors. The judgment criteria need to be relaxed.

[0143] Table 2. Statistics of D-score data Data Name Number of groups Correctly identify the number of groups Accuracy (%) man 121 106 87.6 woman 163 120 73.6 E-score recognition is highly subjective. Although points can be deducted according to the rules in the 2025-2028 MAG CoP based on criteria such as major and minor errors, the algorithm requires more explicit and quantifiable standards. In the E-score data, the deduction items are not clearly defined, and the deduction points and requirements cannot be determined at present, therefore calculation is not possible. The program takes the average score of the statistical data as the E-score output.

[0144] The main performance indicators of the vault markerless image processing and intelligent motion capture system based on computer vision technology in this invention are as follows: 1) Achieve automatic identification and tracking of key points and limb segments of the human body. 2) Reconstruct a 3D human skeleton model, with the reconstructed coordinate system supporting translation and rotation. 3) Enable simultaneous display or export of skeleton diagrams, limb diagrams, data curves, and source video images. 4) Display the trajectory of points, limbs, and the center of gravity of the human body. 5) For vaulting, achieve automatic identification of key joints during training, such as the steps onto the board, the first flight, and landing, displaying the trajectory of the center of gravity, and real-time changes in angles and angular velocities of shoulder, elbow, hip, knee, and ankle angles. 6) Implement a system scoring function. 7) The proportion of test samples with a multiple correlation coefficient greater than 0.95 between the software-automated analysis curve and the corresponding manually analyzed average curve is no less than 90%. 8) The overlap between system scores and manual scores is greater than 90%.

[0145] In summary, this invention, without the need for marker points or worn sensors, utilizes machine vision and deep learning technologies to automatically identify joints in the human body during movement. It accurately tracks high-speed movements, tumbling, and rotations, and analyzes and displays various kinematic parameters in real time, such as the board angle, push angle, takeoff height, rotational angular velocity, center of mass, and the trajectory of each joint. The system can also display the virtual skeletal model synchronously or separately from the recorded video, supports forward and backward slow-motion playback or frame-by-frame playback, and supports multi-structure feedback such as fusion comparison of 2D video images and 3D models, simultaneous comparison of 3D models on the same screen, and simultaneous comparison of data curves on the same screen. Furthermore, it provides evaluation and suggestions through a sports biomechanics dataset and cluster analysis system.

[0146] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0147] It will be readily understood by those skilled in the art that this invention includes any combination of the inventive description and specific embodiments outlined in the foregoing specification, as well as the various parts shown in the accompanying drawings. Due to space limitations and for the sake of brevity, not all of these combinations have been described in detail. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

[0148] Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A markerless image processing and intelligent motion capture system for vaulting horses based on computer vision technology, characterized in that, include: Multiple cameras and image processing and intelligent motion capture devices, among which, The multiple cameras are used to capture video images of the vaulting motion of the moving target at fixed points; The image processing and intelligent motion capture device is used to extract feature points from video images from different cameras, automatically identify human joints in motion using machine vision and deep learning technologies, track and analyze the motion in real time, reconstruct and obtain the three-dimensional spatial coordinate information of the moving target, and optimize the three-dimensional spatial coordinate information to obtain motion capture data. The image processing and intelligent motion capture device includes: a calibration module, an acquisition module, an analysis module, a two-dimensional posture analysis unit, and a human posture filtering unit; The calibration module is used to collect calibration videos of the calibration rod swinging at each camera position and verify whether the calibration results meet expectations; if not, the calibration video is corrected. The acquisition module is used to acquire video data captured by the camera, and to perform data cropping and calculation on the video data to obtain the kinematic data from the most recently acquired video data; The analysis module is used to analyze the dynamic parameters, key parameters, and motion trajectories of video data; The two-dimensional posture analysis unit identifies the human body outline from the two-dimensional image of the moving target captured by the camera; detects the key points of the human skeleton, analyzes the coordinates of the key points of the human skeleton; and matches the posture data of the target athlete in the 2D posture of multiple perspectives at the same time, further analyzes and calculates the joint position of the human body with the three-dimensional fusion algorithm, generates the three-dimensional spatial coordinates of the key points of the human skeleton, and obtains motion capture data. The human posture filtering unit is used to smooth the calculated three-dimensional spatial coordinates of key points of the human skeleton, reduce high-frequency components in the data, remove jitter and abrupt changes in the data, and obtain filtered motion capture data.

2. The vaulting horse markerless image processing and intelligent motion capture system based on computer vision technology according to claim 1, characterized in that, The multiple cameras are calibrated to obtain their structural parameters, internal parameters, and distortion coefficients for three-dimensional spatial positioning. The calibration of the structural parameters among the multiple cameras needs to ensure that the multiple cameras take pictures of the same calibration rod at the same time; by synchronously acquiring images of the moving target, the images of the moving target in different cameras at the same time are obtained.

3. The vaulting markerless image processing and intelligent motion capture system based on computer vision technology according to claim 1, characterized in that, The multiple cameras include 8 cameras, located at 8 video information acquisition positions; the core acquisition area is the area from the take-off board to the landing area; two cameras are placed in a group, respectively at the middle and rear section of the track, the vaulting apparatus, the middle of the landing area on both sides, and at appropriate positions along the longitudinal extension line at the end of the field; each camera position can identify at least 3 other camera positions within its frame.

4. The vaulting markerless image processing and intelligent motion capture system based on computer vision technology according to claim 1, characterized in that, The calibration module performs calibration correction by: decomposing the calibration video into images, dragging and saving the calibration box frame by frame, recalculating the calibration after each frame is modified, verifying the calibration result again, and confirming the calibration is complete if the calibration result is ideal.

5. The vaulting markerless image processing and intelligent motion capture system based on computer vision technology according to claim 1, characterized in that, The acquisition module performs data cropping on the video data, including: selecting the scene to be cropped in the interface, waiting for the video format conversion to be completed, and then cropping the video segment according to the required precise data calculation. The acquisition module performs data calculations on the video data, including: after completing the data cropping operation, it will automatically calculate the kinematic data of the most recently acquired data or select the data file to be calculated according to the user's calculation needs.

6. The vaulting horse unmarked image processing and intelligent motion capture system based on computer vision technology according to claim 1, characterized in that, The two-dimensional pose parsing unit uses deep learning technology to train two neural network models, including a human target detection model and a human pose estimation model. The human target detection model is used to detect the human body contour of the athlete from the image; the human pose estimation model is used to analyze the coordinates of the detected athlete and calculate the image coordinates of 25 joints of the human body.

7. The vaulting markerless image processing and intelligent motion capture system based on computer vision technology according to claim 1, characterized in that, The analysis module performs dynamic parameter analysis on video data, including setting the truncation frequency, the start and end frames of the analysis segment, and analyzing the real-time changes in the kinematic data of the acquired object. The analysis module performs key parameter analysis on the video data, including: analyzing the kinematic analysis indicators of the vault, adjusting the number of frames in which the key parameters occur based on the calculation results and the actual situation of the video, confirming the time when the key kinematic parameters occur, inputting the corresponding frame number into the corresponding wireframe, and the key parameters will change to the value at the corresponding time, and finally outputting a key parameter analysis report. The analysis module performs motion trajectory analysis on video data, including: selecting key human points to be analyzed, projecting their motion paths onto various viewpoints, supporting switching between different viewpoints to view the motion paths of key points, and outputting a motion trajectory coordinate report.

8. The vaulting markerless image processing and intelligent motion capture system based on computer vision technology according to claim 1, characterized in that, Also includes: The basic settings module and the area drawing unit, among which, The basic settings module is used to set the preset parameters of the camera, multiple calibration parameters of the calibration module, multiple acquisition parameters of the acquisition module, and multiple analysis parameters of the analysis module. The region drawing unit is used to customize the data acquisition region under the user's operation, including: custom drawing the acquisition and calculation region and setting the automatic acquisition region.

9. The vaulting markerless image processing and intelligent motion capture system based on computer vision technology according to claim 1, characterized in that, Also includes: The comparison unit is used to compare and analyze the kinematic data and video footage of the same or different subjects in different videos.

10. The vaulting markerless image processing and intelligent motion capture system based on computer vision technology according to claim 1, characterized in that, The image processing and intelligent motion capture device is also used to present motion capture data in a visual manner and supports the export of vaulting technique analysis reports, including: displaying the three-dimensional reconstructed virtual skeleton model and the recorded video synchronously or separately, supporting forward and backward slow motion playback or frame-by-frame playback, fusion and comparison of two-dimensional video images and three-dimensional models, comparison of three-dimensional models on the same screen, comparison of data curves on the same screen, and providing evaluation and suggestions through sports biomechanics datasets and cluster analysis.

Citation Information

Patent Citations

  • Calibration method and system of multi-camera system

    CN114283203A

Cited By

  • A motion data acquisition system, a swimming trajectory analysis method and a storage medium

    CN122199620A