A motion capture method based on sports equipment training
By combining the temporal mapping of training event data streams and image data streams with multi-view geometric constraints, the problems of unstable temporal correspondence between human body and equipment movements and insufficient continuity of motion reconstruction during high-speed motion phases are solved, achieving high-precision motion capture results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUIZHOU BUSINESS SCHOOL
- Filing Date
- 2026-01-22
- Publication Date
- 2026-05-01
AI Technical Summary
Existing motion capture methods suffer from unstable temporal correspondence between human body and equipment movements in sports training scenarios, and insufficient continuity in motion reconstruction during high-speed movements, resulting in low capture accuracy and continuity.
By acquiring training event data streams, training field image data streams, and camera pose parameter information, human keypoint detection and target segmentation algorithms are used to extract skeleton keypoints and equipment 2D contours, establish temporal mapping relationships, perform intermediate frame reconstruction and multi-view geometric constraints, and recover 3D pose data.
It achieves continuous completion of action sequences and improves temporal accuracy, enhances the recognition and analysis capabilities of human-equipment interaction actions, and strengthens spatiotemporal consistency and reconstruction accuracy in high-speed dynamic training scenarios.
Smart Images

Figure CN121545230B_ABST
Abstract
Description
A motion capture method based on sports equipment training Technical Field
[0001] This invention relates to the field of intelligent motion capture technology, and in particular to a motion capture method based on sports equipment training. Background Technology
[0002] With the development of scientific and digital sports training, motion capture technology has gradually become an important means of sports performance analysis, movement evaluation, and training feedback. Existing motion capture methods mainly include marker-based optical capture systems, inertial sensor-based posture capture systems, and markerless capture technology based on visual algorithms. Optical systems, by attaching reflective markers to the subject's body and using a multi-camera array to reconstruct the three-dimensional motion trajectory, offer high accuracy but have strict requirements for the venue environment and equipment layout, resulting in high costs. In recent years, with the advancement of deep learning and computer vision, image-based markerless motion capture has gradually emerged. This method identifies key point skeleton structures in images through human pose estimation algorithms and combines them with temporal information to achieve motion reconstruction. This type of method significantly reduces experimental costs and improves flexibility.
[0003] In sports training scenarios, the interactions between athletes and equipment (such as rackets, barbells, dumbbells, and balls) are characterized by high speed, complexity, and multiple perspectives. Although attempts have been made to use video sequences for joint recognition, the lack of comprehensive utilization of training event data and camera pose information leads to unstable temporal correspondences between human key points and equipment movements, making it difficult to achieve high spatiotemporal accuracy in 3D reconstruction. Furthermore, during high-speed motion phases, image frame blurring and missing key points result in gaps in intermediate frames during motion reconstruction, affecting the overall capture continuity and accuracy. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a motion capture method based on sports equipment training to solve the problems of unstable temporal correspondence between human body and equipment movement and insufficient continuity of motion reconstruction during high-speed movement.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] This invention provides a motion capture method based on sports equipment training, which includes acquiring training event data stream, training field image data stream, and camera pose parameter information;
[0008] The human key point detection algorithm is used to extract the skeleton key point set of the test object from the training field image data stream, and the target segmentation algorithm is used to extract the two-dimensional contour region of the sports equipment. At the same time, the time index information of the corresponding image frame is recorded, and the timestamp of the training event data stream is used for synchronous comparison to establish a temporal mapping relationship and obtain temporal action data.
[0009] Using temporal motion data as spatial anchors, and based on temporal mapping relationships, intermediate frames are reconstructed for the high-speed motion phase to restore the equipment's motion trajectory. Joint matching is then performed to generate a temporal fusion sequence.
[0010] By establishing multi-view geometric constraints through temporal fusion sequence and camera pose parameter information, the point cloud of the human three-dimensional skeleton and the initial six-dimensional pose of the equipment are restored.
[0011] Noise filtering, temporal smoothing, and feature normalization are performed on the human body 3D skeleton point cloud and the initial six-dimensional pose of the equipment to output a 3D pose dataset.
[0012] Backpropagation optimization is performed on the 3D pose dataset to output the training motion capture results.
[0013] As a preferred embodiment of the motion capture method based on sports equipment training described in this invention, the specific steps for extracting the set of skeletal key points of the tested object from the training field image data stream using a human key point detection algorithm are as follows:
[0014] Using consecutive image frames in the training field image data stream as input, the timestamps of each image frame are extracted, and the training event data stream is synchronized and registered based on the timestamps to generate a synchronized time series dataset.
[0015] Abnormal frames and redundant data are removed from the synchronous time-series dataset, and interpolation is performed based on the time interval between image frames to form a standardized image sequence.
[0016] In a standardized image sequence, a human key point detection algorithm is used to identify the main body parts of the subject and extract the two-dimensional skeleton structure information of the human body in each image frame to obtain a set of skeleton key points.
[0017] As a preferred embodiment of the motion capture method based on sports equipment training described in this invention, the specific steps for extracting the two-dimensional contour region of the sports equipment using a target segmentation algorithm are as follows:
[0018] The skeleton keypoint set is matched in chronological order, and the motion trend is calculated by the relative displacement relationship of the keypoint trajectories in adjacent image frames to form a skeleton temporal chain;
[0019] Dynamically smooth the skeleton temporal chain to output a dynamic skeleton sequence;
[0020] Using the skeleton dynamic sequence as a reference, the corresponding frames are extracted from the standardized image sequence, and the pixel regions of sports equipment are detected and separated near the skeleton region using the target segmentation algorithm to obtain the sports equipment segmentation result set.
[0021] Morphological analysis was performed on the sports equipment segmentation result set to extract the two-dimensional contour boundary points of the sports equipment in each image frame, and the equipment contour sequence was generated by judging the boundary connectivity.
[0022] As a preferred embodiment of the motion capture method based on sports equipment training described in this invention, the steps of simultaneously recording the time index information of corresponding image frames and synchronously comparing them using the timestamps of the training event data stream to establish a temporal mapping relationship and obtain temporal motion data are as follows.
[0023] Based on the time index of the image frame, the event timestamp information in the training event data stream is synchronized and compared with the image frame time. With the time alignment result as a reference, the skeleton dynamic sequence and the equipment contour sequence are bound to each other on the same time axis to establish a time-series mapping relationship.
[0024] Perform inter-frame association verification on the time frame pairing data in the time-series mapping relationship to generate a valid mapping relationship set;
[0025] Based on the effective mapping relationship set, the dynamic sequence of the skeleton and the contour sequence of the equipment under the corresponding time frame are fused with temporal features to output temporal motion data.
[0026] As a preferred embodiment of the motion capture method based on sports equipment training described in this invention, the specific steps of using temporal motion data as spatial anchor points and reconstructing intermediate frames for the high-speed motion phase based on temporal mapping relationships are as follows.
[0027] Using temporal motion data as spatial anchors and identifying key time intervals in the high-speed motion phase based on temporal mapping relationships, the time gap between two image frames is determined.
[0028] Interpolation frames are used to fill in the time gaps, generating intermediate frames, which are then arranged in chronological order to form an interpolated frame sequence.
[0029] As a preferred embodiment of the motion capture method based on sports equipment training according to the present invention, the specific steps for recovering the motion trajectory of the equipment are as follows:
[0030] By utilizing the changing characteristics of the skeleton position and the outline of the sports equipment in the interpolated frame sequence, the transient motion path of the sports equipment in the high-speed motion phase can be calculated, and a preliminary transient trajectory can be obtained.
[0031] The initial transient trajectory is time-aligned with the training event data stream, and the trajectory continuity is corrected based on the high temporal resolution response characteristics in the training event data stream to generate the equipment motion trajectory.
[0032] In a preferred embodiment of the motion capture method based on sports equipment training described in this invention, the specific steps for generating the temporal fusion sequence are as follows:
[0033] Spatial registration is performed between the spatial texture features of the corresponding image frames in the training field image data stream and the motion trajectory of the equipment to form a spatiotemporal matching set.
[0034] Based on the spatiotemporal matching set, the dynamic information of the human skeleton and the motion trajectory of the equipment are jointly fused in time to generate a time-series fusion sequence.
[0035] As a preferred embodiment of the motion capture method based on sports equipment training described in this invention, the steps for establishing multi-view geometric constraints by combining temporal fusion sequences and camera pose parameter information to recover the human body's three-dimensional skeleton point cloud and the equipment's initial six-dimensional pose are as follows:
[0036] Using the skeleton key points and equipment contour points in the temporal fusion sequence as corresponding targets, two-dimensional coordinate positions are extracted in image frames from each viewpoint to form a multi-view corresponding point set. Based on the camera pose parameter information, multi-view geometric constraints are established.
[0037] Based on the projection positions under different camera views, the reprojection deviation is calculated, and the multi-view geometric constraints are adjusted by an iterative optimization method. The optimized multi-view geometric constraint relationship is used to perform triangulation calculation on the skeleton key points in the temporal fusion sequence to generate a human three-dimensional skeleton point cloud.
[0038] Based on the contour direction, position center, and motion trend of the equipment in each viewpoint, the three-dimensional translation and three-dimensional rotation of the initial attitude are calculated to obtain the initial six-dimensional attitude of the equipment.
[0039] As a preferred embodiment of the motion capture method based on sports equipment training described in this invention, the steps of performing noise filtering, temporal smoothing, and feature normalization on the human body's three-dimensional skeleton point cloud and the initial six-dimensional posture of the equipment to output a three-dimensional posture dataset are as follows.
[0040] Noise is removed from the coordinate sequence of the human 3D skeleton point cloud while maintaining the integrity of the spatial structure, resulting in a smooth skeleton point cloud sequence.
[0041] Based on the smooth skeleton point cloud sequence, the abrupt changes and discontinuities in the time dimension of the initial six-dimensional attitude of the equipment are corrected to generate a time-stable attitude sequence.
[0042] The smooth skeleton point cloud sequence and the time-stable pose sequence are normalized based on human body proportions and equipment size information to output a 3D pose dataset.
[0043] As a preferred embodiment of the motion capture method based on sports equipment training described in this invention, the output training motion capture result is calculated based on the spatial continuity and motion constraint relationship between each temporal frame in the three-dimensional pose dataset, the deviation of the pose change is calculated, and the spatial position and pose direction of the skeleton node are gradually corrected through error back-iteration until the overall error converges.
[0044] The beneficial effects of this invention are as follows: By establishing a temporal mapping relationship and using temporal motion data as spatial anchors, intermediate frame reconstruction is performed on the high-speed motion phase, achieving continuous completion of motion sequences and improving temporal accuracy; at the same time, multi-view geometric constraints are established based on temporal fusion sequences and camera pose parameters to restore the human body's three-dimensional skeleton point cloud and the equipment's six-dimensional posture, realizing collaborative three-dimensional reconstruction of the human body and equipment under multi-view conditions, effectively improving the spatiotemporal consistency and reconstruction accuracy of motion capture in high-speed dynamic training scenarios, and enhancing the system's ability to recognize and analyze human-equipment interactive actions. Attached Figure Description
[0045] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 is a flowchart of a motion capture method based on sports equipment training.
[0047] Figure 2 is a flowchart of the equipment contour extraction based on the skeleton time chain.
[0048] Figure 3 is a flowchart of the 3D reconstruction process under multi-view geometric constraints.
[0049] Figure 4 is a flowchart of the post-processing of 3D attitude data. Detailed Implementation
[0050] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0051] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0052] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0053] Referring to Figures 1-4, an embodiment of the present invention is provided, which offers a motion capture method based on sports equipment training, comprising the following steps:
[0054] S1. Acquire training event data stream, training field image data stream, and camera pose parameter information.
[0055] It should be noted that the training event data stream refers to the pixel-level brightness change event sequence indexed by timestamps collected by the event camera, which includes the event trigger time, pixel coordinates and polarity information.
[0056] The training field image data stream refers to the sequence of image frames continuously acquired by a regular image sensor in a training field environment, including the image data of each frame and the corresponding acquisition timestamp.
[0057] Camera pose parameter information refers to the external parameters (rotation matrix and translation vector) used to characterize the camera's position and orientation in the world coordinate system, as well as the internal parameters (focal length, principal point coordinates, distortion coefficients, etc.) of the camera imaging model.
[0058] S2. Use the human body key point detection algorithm to extract the skeleton key point set of the tested object from the training field image data stream, and use the target segmentation algorithm to extract the two-dimensional contour region of the sports equipment. At the same time, record the time index information of the corresponding image frame, and use the timestamp of the training event data stream for synchronous comparison to establish a temporal mapping relationship and obtain temporal action data.
[0059] S2.1. Using consecutive image frames in the training field image data stream as input, extract the timestamp of each image frame, and use the timestamp as a reference to perform synchronous registration on the training event data stream to generate a synchronous time-series dataset.
[0060] Furthermore, the acquisition timestamp of each image frame is read from the frame header information of the training field image data stream, and an image frame time index table is established. Then, based on the timestamps in the image frame time index table, event data segments whose timestamps fall within the corresponding range are retrieved in the training event data stream. Each image frame is paired and integrated with the matched event data segments in chronological order to generate a synchronized time-series dataset.
[0061] S2.2 Remove abnormal frames and redundant data from the synchronous time-series dataset, and interpolate the frames according to the time interval between image frames to form a standardized image sequence.
[0062] Furthermore, integrity checks are performed on each image frame in the synchronous time-series dataset. Image frames with abnormal exposure, blurring, out-of-focus, or abnormal timestamps are removed based on image frame quality parameters and timestamp continuity. Subsequently, redundant image frames with duplicate timestamps or excessively close adjacency are removed. Finally, missing frames are generated using a time interpolation algorithm based on the time interval between adjacent valid image frames, and then rearranged in chronological order to form a standardized image sequence.
[0063] S2.3 In the standardized image sequence, the main body parts of the tested object are identified using the human body key point detection algorithm, and the two-dimensional skeleton structure information of the human body in each image frame is extracted to obtain the skeleton key point set.
[0064] Furthermore, using each frame of a standardized image sequence as input, a human keypoint detection algorithm is used to detect and locate the subject in each frame, identifying the main body parts of the subject. Then, based on the detected body parts, the two-dimensional skeleton structure information of the human body is extracted, including the two-dimensional coordinate positions of each body part node and their connection relationships. The two-dimensional skeleton structure information in all image frames is summarized in chronological order to generate a skeleton keypoint set.
[0065] It should be noted that human keypoint detection algorithms are a computer vision technique primarily used to automatically identify and locate key parts of the human body from images or videos, such as the head, shoulders, elbows, knees, wrists, and ankles. Human keypoint detection algorithms analyze the human posture in an image, extract the two-dimensional coordinates of each keypoint, and construct the human skeleton structure based on these coordinates.
[0066] S2.4. The skeleton key point set is matched in chronological order, and the motion trend is calculated by the relative displacement relationship of the key point trajectories of adjacent image frames to form a skeleton temporal chain.
[0067] Furthermore, according to the temporal order of the image frames, the skeleton key points in each frame are numbered and matched with the skeleton key points of the previous frame. Based on the two-dimensional coordinate positions of the corresponding skeleton key points in adjacent image frames, the relative displacement vector is calculated to obtain the motion change information of each skeleton key point (including the relative displacement, velocity change and direction change of the motion trajectory of the key points). Based on the temporal continuity of the skeleton key points, the displacement trajectories of adjacent frames are connected to form a complete skeleton temporal chain.
[0068] S2.5. Perform dynamic smoothing on the skeleton timing chain and output the skeleton dynamic sequence.
[0069] Furthermore, taking the skeleton time sequence chain as input, the system first performs abnormal fluctuation detection (such as velocity mutation and frame missing marker) on the two-dimensional coordinate sequence of each skeleton node according to the timestamp order to determine the time period that needs to be smoothed; the system then performs dynamic smoothing on the two-dimensional coordinate sequence of each skeleton node using bidirectional Savitzky-Golay filtering or Kalman filtering, and applies soft constraints on the bone segment length to keep the skeleton topology and relative length unchanged, and outputs the smoothed skeleton dynamic sequence in time order.
[0070] It should be noted that the soft constraint on bone segment length is a constraint used to maintain the consistency of the skeleton topology. It ensures that the relative bone segment lengths between skeleton nodes do not change significantly during the smoothing process. The soft constraint on bone segment length is set by keeping the relative lengths between adjacent skeleton nodes constant.
[0071] S2.6. Using the skeleton dynamic sequence as a reference, extract the corresponding frames from the standardized image sequence, and use the target segmentation algorithm to detect and separate the pixel regions of sports equipment near the skeleton region to obtain the sports equipment segmentation result set.
[0072] Furthermore, a one-to-one correspondence is established between the timestamps in the dynamic skeleton sequence and the timestamps in the standardized image sequence to extract the image frames corresponding to each timestamp from the standardized image sequence. Then, using the two-dimensional coordinates of each skeleton node in the dynamic skeleton sequence, a skeleton neighborhood search region is constructed on each corresponding image frame. Only equipment that has a relatively stable outline during movement within the skeleton neighborhood search region, can form a continuous and identifiable boundary in the image, and is at least partially exposed within the skeleton neighborhood search region, and meets the minimum area / minimum width or diameter requirements, such as slender or rod-shaped equipment like bars / sticks (meeting the minimum visible width threshold) and training aids like kettlebell handles and training... The process involves extracting candidate pixel regions of sports equipment within the skeleton neighborhood search area using a target segmentation algorithm based on foreground extraction and edge detection. All independent regions in the image are identified through connected component analysis. Morphological opening operations (erosion + dilation) are then used to remove small noise points and isolated regions. Edge defects are repaired using closing operations. Noisy connected components that do not meet the criteria are filtered and removed based on features such as region area, aspect ratio, and boundary continuity. Finally, candidate pixel regions are filtered and retained according to the connected component area range, aspect ratio, and boundary continuity. The selected sports equipment pixel regions from each corresponding image frame are then compiled into a sports equipment segmentation result set in chronological order.
[0073] It should be noted that the minimum visible width threshold is used to limit the minimum pixel width at which slender or rod-shaped objects can be stably segmented in an image;
[0074] The minimum visible width threshold is determined based on the imaging resolution and the skeleton scale normalization result. It is usually based on a fixed ratio of the distance between key skeleton nodes (such as shoulder width or forearm length) and combined with the ability of edge detection to distinguish noise from real boundaries, so as to avoid misidentifying thin line noise or motion blur artifacts below this width as equipment.
[0075] The area threshold is used to limit the minimum effective pixel area of candidate connected regions. The area threshold is based on the expected visible size of the equipment at the current shooting angle and distance, and can be normalized by combining human scale or camera calibration parameters. Connected regions below the area threshold are regarded as noise or incomplete visible areas and are eliminated, thereby ensuring that the retained equipment areas have the reliability of subsequent analysis in terms of morphology and stability.
[0076] S2.7 Perform morphological analysis on the sports equipment segmentation result set, extract the two-dimensional contour boundary points of the sports equipment in each image frame, and generate the equipment contour sequence by judging the boundary connectivity.
[0077] Furthermore, morphological opening and closing operations are performed on the sports equipment segmentation mask in each image frame to remove isolated small regions and repair boundary gaps. Then, the morphological gradient is calculated to obtain the sports equipment segmentation boundary, and a boundary tracking algorithm is used to extract the two-dimensional contour boundary points. Subsequently, the connectivity of the two-dimensional contour boundary points in each image frame is judged according to the four-neighbor or eight-neighbor connectivity rules. The continuous boundary point set is divided into independent contours by the connectivity component label, and noisy contours are removed according to the boundary length and geometric continuity. The two-dimensional contour boundary points in each image frame that have been judged by connectivity are organized into a sequence of equipment contours in chronological order.
[0078] It should be noted that the boundary tracking algorithm is used to extract the contour boundaries of objects in an image. It extracts the closed boundary of the object by traversing the edge pixels in the binary image and connecting adjacent edge points according to certain rules (such as clockwise or counterclockwise). During the extraction process, the boundary tracking algorithm can effectively identify the complete boundary of the object and provide an accurate set of contour points for subsequent analysis.
[0079] Four-neighbor and eight-neighbor are two pixel connection rules in image processing. The four-neighbor rule means that a pixel is adjacent to the pixels in the four directions of up, down, left, and right. The eight-neighbor rule includes the pixels adjacent in the four directions of up, down, left, right, and four diagonals. By using these two pixel connection rules, we can determine whether adjacent pixels in an image belong to the same connected region, which helps to determine the boundaries and connectivity of objects.
[0080] S2.8. Based on the time index of the image frame, the event timestamp information in the training event data stream is synchronized and compared with the image frame time. With the time alignment result as a reference, the skeleton dynamic sequence and the equipment contour sequence are bound to each other on the same time axis to establish a time-series mapping relationship.
[0081] Furthermore, based on the image frame time index table, the event timestamps in the training event data stream are retrieved one by one and matched with nearest neighbor / linear interpolation to generate a time alignment result table of the correspondence between image frame time and event time. With the time alignment result table as a reference, the corresponding skeleton frame in the skeleton dynamic sequence and the corresponding equipment contour frame in the equipment contour sequence are located for each time record on a unified time axis, and consistency is checked with a time error threshold (e.g., the time difference does not exceed 5 milliseconds in the example). The skeleton frames and equipment contour frames are paired and bound in order according to the unified time axis to establish a time sequence mapping relationship.
[0082] It should be noted that the time error threshold is determined by calculating the average time interval between image frames and combining it with the minimum event trigger interval of the event data stream, taking a certain proportion (e.g., 0.1 to 0.2 times) of the smaller of the two values.
[0083] The time error threshold is usually set between 1 millisecond and 10 milliseconds. The specific value range is dynamically adjusted according to the image frame rate, event stream trigger density and camera synchronization accuracy. When the image frame rate is high or the event trigger frequency is dense, the time error threshold is taken as a smaller value.
[0084] S2.9 Perform inter-frame association verification on the paired data of each time frame in the time-series mapping relationship to generate a valid mapping relationship set.
[0085] Furthermore, each pair of skeleton frames and equipment contour frames is paired and verified in chronological order. The spatial displacement difference and directional change between the skeleton key point trajectory and the equipment contour center point in adjacent time frames are calculated. Then, the validity of the pairing relationship is judged based on the continuity of spatial displacement and the consistency of direction (for example, if the displacement difference between the skeleton key point and the equipment contour center point in adjacent time frames is large, or their motion direction deviates significantly (such as the angle exceeding the direction consistency threshold, or the displacement difference exceeding the spatial displacement continuity threshold), the pairing relationship is considered invalid and is removed). Abnormal pairings are removed, and the pairing data that have passed the inter-frame association verification are organized in chronological order to generate a set of valid mapping relationships.
[0086] It should be noted that the spatial displacement continuity threshold and orientation consistency threshold are usually determined based on the image frame rate, motion speed, and characteristics of the target object.
[0087] The spatial displacement continuity threshold is set to a range of 5 to 30 pixels based on the image resolution and the speed of motion.
[0088] The directional consistency threshold is set based on the motion trajectory characteristics of the target object, such as the type of motion (e.g., rapid rotation or smooth movement).
[0089] The directional consistency threshold is set between 10° and 20° based on the smoothness of the movement. If the movement is relatively smooth, the directional consistency threshold is larger; if the movement is more vigorous, the directional consistency threshold is smaller.
[0090] S2.10. Based on the effective mapping relationship set, the skeleton dynamic sequence and the equipment contour sequence under the corresponding time frame are fused with temporal features to output temporal motion data.
[0091] Furthermore, using the effective mapping relationship set as an index, for each timestamp, the time-series features such as the two-dimensional coordinates of skeleton nodes, the line vectors between skeleton nodes, and the displacement vectors of adjacent frames are first extracted from the skeleton dynamic sequence. The time-series features such as the coordinates of the center point of the contour, the main axis direction vector, and the statistical parameters of the boundary point set are extracted from the equipment contour sequence. Then, based on the pairing index in the effective mapping relationship set, the time-series features of the skeleton dynamic sequence and the equipment contour sequence are aligned and spliced one by one on a unified time axis. The continuity is checked by the Euclidean distance and the direction angle between the fused features of the two frames. The skeleton and equipment fusion feature sequence that has passed the continuity check is output in chronological order as the time-series motion data.
[0092] In summary, this invention establishes a time index table for training field image data streams and utilizes high-time-resolution timestamps from training event data streams for precise synchronous comparison, achieving dynamic temporal registration between image frames and event data. Simultaneously, it binds the skeleton key point set to the two-dimensional contour region of sports equipment on the same time axis, forming a temporal mapping relationship, ensuring that human movement and equipment movement remain consistent in the time dimension, thereby significantly improving the registration accuracy and continuity of movement correspondence in high-speed motion capture.
[0093] Existing methods typically treat human keypoint detection and equipment segmentation as two independent processes. First, a single-frame human keypoint detection algorithm is used to extract the skeleton nodes of the object being tested. Then, equipment detection based on semantic segmentation or instance segmentation is performed separately. The two methods perform simple time alignment or frame index matching at the result level. They do not perform high-precision time synchronization between image frames and event data streams, and they also lack time constraints based on dynamic trajectories. Therefore, in cases of high-speed movement or occlusion, errors in the correspondence between keypoints and equipment or time drift problems are likely to occur.
[0094] S3. Using temporal motion data as spatial anchors, and based on temporal mapping relationships, perform intermediate frame reconstruction for the high-speed motion phase to restore the equipment's motion trajectory, and perform joint matching to generate a temporal fusion sequence.
[0095] S3.1 Using temporal motion data as spatial anchors and identifying key time intervals in the high-speed motion phase based on temporal mapping relationships, the time gap between two image frames is determined.
[0096] Furthermore, using temporal motion data as spatial anchors, and calculating the motion velocity and acceleration change rate of adjacent time frames based on the timestamp difference between the skeleton dynamic sequence and the equipment contour sequence in the temporal mapping relationship; then determining the high-speed motion phase based on the motion velocity change rate threshold and the acceleration change rate threshold, and identifying time periods in the high-speed motion phase where the continuous timestamp interval exceeds the upper limit of the average inter-frame time interval, marking these interval segments as time missing segments;
[0097] The rate of change of motion velocity between adjacent time frames is calculated using the following expression:
[0098] ;
[0099] In the formula, Adjacent time frames The rate of change of motion speed between the two is used to describe the rate of change of the relative motion speed of the skeleton dynamic sequence and the equipment contour sequence over time, and the unit is meters per square second (m / s²) or pixels per square second (px / s²). In time frame The relative motion velocity magnitude between the dynamic sequence of the skeleton and the contour sequence of the equipment, in meters per second (m / s) or pixels per second (px / s). In time frame The relative motion velocity magnitude between the dynamic sequence of the skeleton and the contour sequence of the equipment, in meters per second (m / s) or pixels per second (px / s). It is a time frame The timestamp is in seconds (s). It is a time frame The timestamp is in seconds (s). An index variable representing a time frame;
[0100] The rate of change of acceleration between adjacent time frames is calculated using the following expression:
[0101] ;
[0102] In the formula, Adjacent time frames The rate of change of acceleration between the two is used to describe the rate of change of the relative acceleration of the skeleton dynamic sequence and the equipment contour sequence over time, and the unit is meters per cubic second (m / s³) or pixels per cubic second (px / s³). It is a time frame The relative acceleration magnitude between the skeleton dynamic sequence and the equipment contour sequence, in meters per square second (m / s²) or pixels per square second (px / s²); is within a time frame. The relative acceleration magnitude between the time skeleton dynamic sequence and the equipment contour sequence, in meters per square second (m / s²) or pixels per square second (px / s²). It is a time frame The timestamp is in seconds (s). It is a time frame The timestamp is in seconds (s).
[0103] It should be noted that the threshold for the rate of change of motion speed is set based on the average motion speed change characteristics of the skeletal dynamic sequence and the equipment contour sequence, and the value ranges from 0.05 m / s² to 0.3 m / s², depending on the speed fluctuation amplitude of different motion types;
[0104] The acceleration change rate threshold is set according to the stability requirements of acceleration during the movement, and the value ranges from 0.5 m / s³ to 2.0 m / s³, determined based on the acceleration change range of the frame and equipment during the high-speed movement phase.
[0105] S3.2. Interpolate and fill in the missing time segments to generate intermediate frames, and arrange them in chronological order to form an interpolated frame sequence.
[0106] Furthermore, the system calculates the target timestamps for each time gap segment according to the target output frame rate, determines the indexes of the preceding and following image frames corresponding to each target timestamp, applies linear time interpolation or motion compensation interpolation based on block matching to the pixels of the entire frame, applies linear or spline interpolation to the two-dimensional coordinates of skeleton nodes in the skeleton dynamic sequence and the two-dimensional coordinates of contour boundary points in the equipment contour sequence, synchronously generates image and structural annotations for intermediate frames, sorts all intermediate frames by timestamp from smallest to largest, and outputs an interpolated frame sequence arranged in chronological order.
[0107] S3.3. By utilizing the changing characteristics of the skeleton position and the outline of the sports equipment in the interpolated frame sequence, the transient motion path of the sports equipment in the high-speed motion phase is calculated, and a preliminary transient trajectory is obtained.
[0108] Furthermore, following the timestamp order of the interpolation frame sequence, the two-dimensional coordinates of the skeleton position and the center point and principal axis direction of the two-dimensional contour of the sports equipment are extracted from each interpolation frame. The displacement vector and direction change of the center point of the sports equipment contour between adjacent interpolation frames are calculated, and the motion trend of the skeleton position is compared with the displacement vector of the sports equipment contour in time according to the temporal mapping relationship. Subsequently, during the high-speed motion phase, the coordinates of the center point of the sports equipment contour ordered by time are accumulated and connected, and spline curve smoothing or least squares curve fitting is used to perform temporal smoothing on the discrete trajectory points. The motion trend of the skeleton position is used to constrain the consistency of the curve direction (for example, when the angle between the motion direction vector of the upper limb node in the skeleton dynamic sequence and the displacement direction of the center point of the sports equipment contour is less than the example value of 10°, the curve direction remains unchanged; when the angle exceeds the example value of 10°, the trajectory segment that deviates from the direction is corrected to ensure that the overall direction change of the transient motion path of the sports equipment is consistent with the motion trend of the skeleton position). Finally, the smoothed continuous path is output as the transient motion path.
[0109] S3.4. Align the initial transient trajectory with the training event data stream in time, and correct the trajectory continuity based on the high time resolution response characteristics in the training event data stream to generate the equipment motion trajectory.
[0110] Furthermore, the timestamps of the initial transient trajectory are matched with those of the training event data stream. The time offset is determined and time alignment is completed by the correlation between the event trigger density sequence and the velocity change trend of the initial transient trajectory. On the time axis after time alignment, the initial transient trajectory is continuously corrected based on the high response time interval corresponding to the event trigger peak in the training event data stream. Piecewise spline smoothing and direction consistency constraints are used to eliminate abrupt changes and discontinuities in the trajectory, and the output is a continuous motion trajectory of the equipment with the same time resolution as the training event data stream.
[0111] S3.5. Spatial registration is performed between the spatial texture features of the corresponding image frames in the training field image data stream and the motion trajectory of the equipment to form a spatiotemporal matching set.
[0112] Furthermore, corresponding image frames are extracted from the training field image data stream, and spatial texture feature sets are extracted from the corresponding image frames using existing feature point and feature descriptor methods (such as ORB or SIFT). A registration search region is established based on the position of the equipment's motion trajectory in the corresponding image frame. Feature matching is performed between the spatial texture feature set and the position of the equipment's motion trajectory. Through progressive optimization and error detection, feature points that are inconsistent with the position of the equipment's motion trajectory are eliminated, ensuring that only accurately matched feature points are retained, thus achieving high-precision spatial registration. Following the temporal order of the equipment's motion trajectory, the spatially registered spatial texture features are bound to the temporal index and spatial position of the equipment's motion trajectory, outputting a spatiotemporal matching set.
[0113] S3.6. Based on the spatiotemporal matching set, perform joint temporal fusion of the dynamic information of the human skeleton and the motion trajectory of the equipment to generate a temporal fusion sequence.
[0114] Furthermore, based on the time index table, the dynamic information of the human skeleton and the motion trajectory of the equipment are matched on the same time axis. The two-dimensional coordinate information of the skeleton node in each time frame is time-aligned with the spatial position of the equipment motion trajectory. Based on the spatial texture feature matching results in the spatiotemporal matching set, spatial fusion is performed on the skeleton node position and the equipment trajectory position. The spatial deviation between the skeleton dynamic information and the equipment motion trajectory is balanced by a weighted average fusion method. The results after time alignment and spatial fusion are integrated in chronological order to generate a temporal fusion sequence containing information on the coordinated motion of the skeleton and the equipment.
[0115] S4. Establish multi-view geometric constraints by using temporal fusion sequence and camera pose parameter information to restore the human body 3D skeleton point cloud and the initial six-dimensional pose of the equipment.
[0116] S4.1. Using the skeleton key points and equipment contour points in the temporal fusion sequence as corresponding targets, extract the two-dimensional coordinate positions in the image frames of each viewpoint to form a multi-viewpoint corresponding point set, and establish multi-viewpoint geometric constraints based on the camera pose parameter information.
[0117] Furthermore, the two-dimensional coordinate positions of the skeleton key points and equipment contour points are extracted one by one in each viewpoint image frame, and a cross-viewpoint correspondence is established based on the timestamp and target index to form a multi-viewpoint corresponding point set. Based on the camera pose parameter information, the camera intrinsic parameters (including focal length, principal point coordinates and distortion coefficients) and camera extrinsic parameters (including rotation matrix and translation vector, used to describe the camera's attitude and position) are used as the basis for geometric constraints. The consistency of the multi-viewpoint corresponding point set is verified and bound according to the projection consistency constraint and epipolar consistency constraint. The multi-viewpoint corresponding point set that has completed the consistency verification is output in time order, and together with the camera pose parameter information, it constitutes the multi-view geometric constraint.
[0118] S4.2 Based on the projection positions under different camera views, calculate the reprojection deviation, and use an iterative optimization method to adjust the multi-view geometric constraints. Then, use the optimized multi-view geometric constraint relationship to perform triangulation calculation on the skeleton key points in the temporal fusion sequence to generate a human three-dimensional skeleton point cloud.
[0119] Furthermore, in each viewpoint image frame, the two-dimensional coordinates of the skeleton key points are subtracted point-by-point from the predicted projection positions obtained by forward projection of the camera intrinsic and extrinsic parameters to obtain the reprojection bias. An iterative optimization method is used to update the weights of the corresponding points in the multi-view and the geometric consistency parameters based on the reprojection bias in the least squares sense until convergence, thereby obtaining the optimized multi-view geometric constraints. The optimized multi-view geometric constraints are used to perform triangulation calculations on the skeleton key points in the temporal fusion sequence and to aggregate the three-dimensional coordinates in chronological order to generate a human three-dimensional skeleton point cloud.
[0120] S4.3. Based on the contour direction, position center and motion trend of the equipment in each viewpoint, calculate the three-dimensional translation and three-dimensional rotation of the initial posture to obtain the initial six-dimensional posture of the equipment.
[0121] Furthermore, based on the camera's intrinsic parameters (such as focal length and principal point coordinates), the two-dimensional points in the image are transformed into direction vectors in the camera coordinate system. Combined with the camera's extrinsic parameters (such as rotation matrix and translation vector), the direction vectors are transformed from the camera coordinate system to the world coordinate system to obtain the line-of-sight ray from the camera to the center of the equipment. The three-dimensional translation of the equipment is obtained by the minimum distance intersection of the multi-view rays. Then, the contour direction of the equipment in each view is used as the principal axis direction for observation. The three-dimensional principal axis direction is solved under the constraint of multi-view projection consistency. The motion trend of the equipment in each view constrains the principal axis orientation to eliminate directional ambiguity, thereby obtaining the three-dimensional rotation. The three-dimensional translation and three-dimensional rotation are bound to the initial six-dimensional attitude of the equipment by time index.
[0122] S5. Perform noise filtering, temporal smoothing, and feature normalization on the human body 3D skeleton point cloud and the initial six-dimensional pose of the equipment, and output a 3D pose dataset.
[0123] S5.1. Remove noise from the coordinate sequence of the human body 3D skeleton point cloud while maintaining the integrity of the spatial structure to obtain a smooth skeleton point cloud sequence.
[0124] Furthermore, anomaly detection and removal are performed on the 3D coordinate sequence of each skeleton node based on the time index. Anomaly criteria include velocity abrupt changes, acceleration abrupt changes, and consistency verification of relative reprojection residual thresholds between adjacent frames. Subsequently, temporal smoothing is performed on the 3D coordinate sequence of each skeleton node using bidirectional Savitzky-Golay filtering or Kalman filtering. During the smoothing process, soft constraints on bone segment length and joint topological constraints are applied to maintain the integrity of the spatial structure. Finally, the human 3D skeleton point cloud with noise removal and structural constraint smoothing is output in chronological order, generating a smooth skeleton point cloud sequence.
[0125] S5.2 Based on the smooth skeleton point cloud sequence, correct the abrupt changes and discontinuities in the time dimension of the initial six-dimensional attitude of the equipment, and generate a time-stable attitude sequence.
[0126] Furthermore, based on the time index table, the three-dimensional translation and three-dimensional rotation of the initial six-dimensional attitude of the equipment are sorted in time series. By observing whether the position and orientation changes of the equipment in adjacent time frames are stable, if the changes are too large or the directions are inconsistent, it indicates that there are abrupt changes or discontinuities, which need to be smoothed. Subsequently, quaternion spherical linear interpolation is used to smooth the three-dimensional rotation for time periods with abrupt changes or discontinuous directions, and low-pass filtering is used to smooth the three-dimensional translation to eliminate jumps. Finally, under the constraint of keeping the attitude trajectory consistent with the motion trend of the smooth skeleton point cloud sequence, a temporally stable attitude sequence with continuous time and no abrupt changes is output.
[0127] S5.3. The smooth skeleton point cloud sequence and the time-stable pose sequence are normalized according to human body proportion features and equipment size information to output a three-dimensional pose dataset.
[0128] Furthermore, reference scale parameters are determined based on human proportions and equipment dimensions (e.g., human torso length as human reference scale, and measured length, width, and height of equipment as equipment reference scale), and a time-indexed reference coordinate benchmark is established. Subsequently, center alignment and linear scale normalization are performed on the 3D coordinates of the smooth skeleton point cloud sequence, so that the length of each bone segment is scaled to the reference length by a scale factor while keeping the bone segment topology and joint relative angles unchanged. At the same time, the 3D translation of the time-stable pose sequence is scaled by the same scale factor and the 3D rotation remains unchanged. Finally, the normalized skeleton 3D coordinates and equipment pose are merged according to the time index, and the reference scale parameters and units are recorded in the output to form a 3D pose dataset.
[0129] S6. Perform backpropagation optimization on the 3D pose dataset and output the training motion capture results.
[0130] S6.1 Based on the spatial continuity and motion constraint relationship between each temporal frame in the 3D pose dataset, calculate the deviation of pose change, and gradually correct the spatial position and pose direction of the skeleton node through error back-to-back iteration until the overall error converges, and output the training motion capture results.
[0131] Furthermore, based on the spatial continuity constraints and kinematic constraints of adjacent time frames, posture change deviations are constructed, including the 3D coordinate differences of skeleton nodes, bone segment length deviations, and joint orientation continuity deviations. These deviations are then aggregated frame by frame in a time index table to form a target deviation set. Subsequently, error backward iteration is performed in the least squares sense, updating the 3D coordinates of skeleton nodes and joint orientation parameters according to the gradient direction. Joint orientation is updated in quaternion form and quaternion normalization is performed after each update, while keeping the soft constraints of bone segment length and joint topological constraints unchanged. Finally, the error convergence criterion is used (e.g., the relative decrease of the target deviation set is less than the example value of 1×10⁻). 4 If the updated step size norm is lower than the example value of 1×10⁻³ and the number of iterations does not exceed the example value of 100 times, the error is terminated and the reverse iteration is performed. The converged skeleton node 3D coordinates and joint orientation are output in time order to form the training motion capture results.
[0132] It should be noted that the error convergence criterion is set based on the convergence trend of the target deviation set during the error reverse iteration process. By statistically analyzing the relationship between the overall error decrease and the parameter update magnitude in multiple iterations, the threshold range in which the error decrease rate is stable and the change tends to be gradual is determined.
[0133] In summary, this invention achieves continuous completion of motion sequences and improves temporal accuracy by establishing a temporal mapping relationship and using temporal motion data as spatial anchors to reconstruct intermediate frames during high-speed motion phases. Simultaneously, it establishes multi-view geometric constraints based on temporal fusion sequences and camera pose parameters to recover the human body's 3D skeleton point cloud and the equipment's six-dimensional pose, enabling collaborative 3D reconstruction of the human body and equipment under multi-view conditions. This effectively improves the spatiotemporal consistency and reconstruction accuracy of motion capture in high-speed dynamic training scenarios and enhances the system's ability to recognize and analyze human-equipment interaction actions.
[0134] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A motion capture method based on sports equipment training, characterized in that: This includes acquiring training event data streams, training field image data streams, and camera pose parameter information; A human keypoint detection algorithm is used to extract the skeleton keypoint set of the tested object from the training field image data stream, and a target segmentation algorithm is used to extract the two-dimensional contour region of the sports equipment. Simultaneously, the time index information of the corresponding image frames is recorded, and the timestamps of the training event data stream are used for synchronous comparison to establish a temporal mapping relationship and obtain temporal motion data. The specific steps for extracting the two-dimensional contour region of the sports equipment using the target segmentation algorithm are as follows: the skeleton keypoint set is matched according to time sequence; the motion trend is calculated through the relative displacement relationship of the keypoint trajectories of adjacent image frames to form a skeleton temporal chain; the skeleton temporal chain is dynamically smoothed to output a skeleton dynamic sequence; using the skeleton dynamic sequence as a reference, from the standardized image sequence... Extract the corresponding frames, and use a target segmentation algorithm to detect and separate the pixel regions of sports equipment near the skeleton region to obtain a sports equipment segmentation result set; perform morphological analysis on the sports equipment segmentation result set to extract the two-dimensional contour boundary points of the sports equipment in each image frame, and generate a sports equipment contour sequence by judging boundary connectivity; use temporal motion data as spatial anchor points, and reconstruct intermediate frames for the high-speed motion stage based on temporal mapping relationships to restore the motion trajectory of the equipment, and perform joint matching to generate a temporal fusion sequence; the specific steps of using temporal motion data as spatial anchor points and reconstructing intermediate frames for the high-speed motion stage based on temporal mapping relationships are as follows: using temporal motion data as spatial anchor points, and based on temporal mapping relationships to identify... The key time intervals of the high-speed motion phase are identified to determine the time gaps between two image frames; interpolation frames are used to fill in the time gaps, generating intermediate frames, which are then arranged sequentially in chronological order to form an interpolated frame sequence; the specific steps for restoring the equipment's motion trajectory are as follows: using the changing characteristics of the skeleton position and the sports equipment outline in the interpolated frame sequence, the transient motion path of the sports equipment during the high-speed motion phase is calculated to obtain a preliminary transient trajectory; the preliminary transient trajectory is time-aligned with the training event data stream, and the trajectory continuity is corrected based on the high temporal resolution response characteristics in the training event data stream to generate the equipment's motion trajectory; multi-view geometric constraints are established through temporal fusion sequence and camera pose parameter information to restore the human three-dimensional skeleton point cloud and the initial transient trajectory of the equipment. The initial six-dimensional pose is determined through the following steps: Using the skeleton keypoints and equipment contour points in the temporal fusion sequence as corresponding targets, two-dimensional coordinate positions are extracted from image frames at each viewpoint to form a multi-viewpoint corresponding point set. Multi-view geometric constraints are established based on camera pose parameters. Reprojection deviations are calculated based on the projection positions under different camera views, and iterative optimization methods are used to adjust the multi-view geometric constraints. The optimized multi-view geometric constraint relationships are then used to triangulate the skeleton keypoints in the temporal fusion sequence, generating a three-dimensional human skeleton point cloud. Based on the equipment's contour direction, position center, and motion trend in each viewpoint, the three-dimensional translation and three-dimensional rotation of the initial pose are calculated to obtain the initial six-dimensional pose of the equipment. The three-dimensional human skeleton point cloud and the initial six-dimensional pose of the equipment are processed by noise filtering, temporal smoothing and feature normalization to output a three-dimensional pose dataset; the three-dimensional pose dataset is optimized by backpropagation to output the training motion capture results.
2. The motion capture method based on sports equipment training as described in claim 1, characterized in that: The specific steps for extracting the skeleton key point set of the test object from the training field image data stream using the human key point detection algorithm are as follows: taking continuous image frames in the training field image data stream as input, extracting the timestamp of each image frame, and using the timestamp as a reference to synchronize and register the training event data stream to generate a synchronized time series dataset; removing abnormal frames and redundant data from the synchronized time series dataset, and interpolating and supplementing frames according to the time interval between image frames to form a standardized image sequence; in the standardized image sequence, using the human key point detection algorithm to identify the main body parts of the test object, and extracting the two-dimensional skeleton structure information of the human body in each image frame to obtain the skeleton key point set.
3. The motion capture method based on sports equipment training as described in claim 1, characterized in that: The process involves simultaneously recording the time index information of the corresponding image frame, and using the timestamp of the training event data stream for synchronous comparison to establish a temporal mapping relationship and obtain temporal action data. The specific steps are as follows: Based on the time index of the image frame, the event timestamp information in the training event data stream is synchronously compared with the image frame time, and the time alignment result is used as a reference to bind the skeleton dynamic sequence and the equipment contour sequence on the same time axis to establish a temporal mapping relationship. Perform inter-frame association verification on the time frame pairing data in the time-series mapping relationship to generate a valid mapping relationship set; Based on the effective mapping relationship set, the dynamic sequence of the skeleton and the contour sequence of the equipment under the corresponding time frame are fused with temporal features to output temporal motion data.
4. The motion capture method based on sports equipment training as described in claim 1, characterized in that: The specific steps for generating the temporal fusion sequence are as follows: spatially register the spatial texture features of the corresponding image frames in the training field image data stream with the motion trajectory of the equipment to form a spatiotemporal matching set; based on the spatiotemporal matching set, perform joint temporal fusion of the dynamic information of the human skeleton and the motion trajectory of the equipment to generate a temporal fusion sequence.
5. The motion capture method based on sports equipment training as described in claim 1, characterized in that: The process of noise filtering, temporal smoothing, and feature normalization of the human body's 3D skeleton point cloud and the initial six-dimensional pose of the equipment to output a 3D pose dataset involves the following steps: Noise is removed from the coordinate sequence of the human body's 3D skeleton point cloud while maintaining spatial structural integrity to obtain a smooth skeleton point cloud sequence; based on the smooth skeleton point cloud sequence, abrupt changes and discontinuities in the temporal dimension of the initial six-dimensional pose of the equipment are corrected to generate a temporally stable pose sequence; the smooth skeleton point cloud sequence and the temporally stable pose sequence are then subjected to feature normalization based on human body proportions and equipment size information to output a 3D pose dataset.
6. The motion capture method based on sports equipment training as described in claim 1, characterized in that: The output training motion capture result is based on the spatial continuity and motion constraint relationship between each temporal frame in the 3D pose dataset. The deviation of the pose change is calculated, and the spatial position and pose direction of the skeleton node are gradually corrected through error back-to-back iteration until the overall error converges.
Citation Information
Patent Citations
Tunnel construction monitoring analysis method and system based on AI video monitoring
CN120108161A
Rehabilitation training action evaluation method and device based on multi-view vision
CN120183042A
Method and system for 3D scanning and dynamic posture capturing of figure model
CN120599132A