A method and system for physical fitness assessment based on human skeletal trajectory tracking

By using multi-view video acquisition and a multimodal spatiotemporal pyramid network, the problem of high cost and inaccurate assessment of existing physical fitness testing equipment has been solved, achieving high-precision and automated physical fitness assessment, providing personalized feedback, and applicable to fields such as public health and sports training.

CN120837061BActive Publication Date: 2026-01-06BEIJING KINGTOP TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510945587.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2026-01-06
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

Existing physical fitness assessment methods rely on a single perspective and traditional image processing techniques. The equipment is expensive, has poor adaptability to different venues, and is difficult to promote on a large scale. Furthermore, they lack the utilization of multimodal spatiotemporal dynamic characteristics, resulting in inaccurate motion assessment and a lack of personalized feedback.

Method used

Employing multi-view video acquisition, dynamic skeletal analysis, and a multimodal spatiotemporal pyramid network, the system performs quantitative evaluation by tracking skeletal key points, verifying identity, recognizing actions, and completing anomalies, combined with biomechanical parameters, to generate personalized feedback.

Benefits of technology

It achieves high-precision and automated physical fitness assessment, improves the scientific nature and reliability of the assessment, and has low cost and high applicability, making it suitable for fields such as public health and sports training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120837061B_ABST
    Figure CN120837061B_ABST
Patent Text Reader

Abstract

The application discloses a kind of physical ability evaluation methods based on human skeleton trajectory tracking, comprising: S1, the whole process video data of physical ability test is obtained by video acquisition equipment, and video frame is generated according to acquisition frequency;S2, video frame is handled using skeletal dynamic analysis algorithm, and human skeleton key point is extracted;S3, tester identity feature is extracted and tracked confirmation by motion atlas identity recognition network;S4, action recognition, segment segmentation and compliance determination are carried out to skeletal key point time sequence trajectory;S5, for abnormal skeletal key point, reverse kinematics inference method is used to complete and time sequence smooth;S6, complete skeletal key point time sequence data is input to human action evaluation model and is handled, and physical ability evaluation index is quantitatively output;S7, physical ability evaluation report is generated and action feedback and optimization suggestion are given.The application effectively improves the objectivity, accuracy and intelligent level of physical ability evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision and artificial intelligence, in particular to a physical ability evaluation method and system based on human skeleton trajectory tracking. BACKGROUND

[0002] With the rapid development of intelligent hardware, computer vision and artificial intelligence technology, human motion recognition and physical ability evaluation have become a research hotspot in sports, health and rehabilitation fields. At present, most of the physical ability evaluation methods based on video adopt single-view video acquisition and traditional image processing technology, and can only preliminarily evaluate the basic motion state of the measured person through manual scoring or pose estimation algorithm. In order to improve the accuracy of motion capture, high-precision sensors, wearable devices or three-dimensional acquisition technology with marker points are introduced in some research fields, but these systems and devices are expensive, have poor site adaptability and are complicated to operate, which makes it difficult to meet the large-scale popularization and application in the scenes of national fitness and youth physical fitness testing.

[0003] In recent years, the skeleton point detection technology based on deep learning has been gradually applied to the field of human pose recognition, and has broken through the bottleneck of traditional two-dimensional image recognition, realizing the automatic detection and tracking of the main skeleton joints of the human body. At the same time, the spatio-temporal graph convolutional neural network can mine the dynamic structure features of the continuous action sequence of the human body, and improve the accuracy and automation level of action recognition. However, the existing technology still has deficiencies in continuous action time sequence analysis, occlusion processing, multi-person interference and accurate tracking of skeleton points in complex natural environment. Most systems lack an intelligent completion mechanism for abnormal skeleton points, which leads to missing and misidentification of action data during intense exercise or local occlusion, thereby affecting the accuracy and reliability of the physical ability evaluation results.

[0004] Traditional physical ability evaluation methods and systems mainly rely on single modal or static feature data, and lack the use of multi-modal spatio-temporal dynamic features. Existing action evaluation algorithms cannot effectively combine biomechanical indicators and sports science parameters for quantitative analysis, and have limitations in detecting key physical ability indicators such as action standardization, amplitude, speed and duration, and lack the ability to recognize fine-grained action differences and abnormal motion patterns. In addition, existing physical ability evaluation reports mainly focus on result display, lack automatic feedback and optimization suggestions for individual differences, which is not conducive to the measured person to carry out targeted physical ability improvement training, and is also inconvenient for subsequent tracking and dynamic optimization.

[0005] Therefore, how to provide a physical ability evaluation method and system based on human skeleton trajectory tracking is a problem that needs to be solved by those skilled in the art. SUMMARY

[0006] One objective of this invention is to propose a physical fitness assessment method and system based on human skeletal trajectory tracking. This invention fully utilizes skeletal dynamic analysis, intelligent motion tracking, and multimodal spatiotemporal pyramid networks, and describes in detail the intelligent physical fitness assessment process of full-process video acquisition, human skeletal point extraction, motion evaluation, and personalized feedback generation. It has the advantages of high automation, accurate assessment, and strong applicability.

[0007] A physical fitness assessment method based on human skeletal trajectory tracking according to an embodiment of the present invention includes the following steps:

[0008] S1. Acquire video data of the entire physical fitness test process through video acquisition equipment, and generate video frames according to the acquisition frequency;

[0009] S2. Process video frames using a skeletal dynamic analysis algorithm to construct a set of temporal motion trajectories of skeletal key points;

[0010] S3. Based on the temporal motion trajectory set of skeletal key points, extract the identity features of the tester through the motion graph identity recognition network throughout the test process, and perform identity tracking and confirmation.

[0011] S4. Perform preliminary motion recognition on the time-series motion trajectory set of the tracked skeletal key points, segment and extract motion segments, and determine the compliance of each motion segment.

[0012] S5. For abnormal skeletal key points that do not pass identity tracking confirmation and compliance judgment, the inverse kinematics reasoning method is used to complete the position, the completion result is processed by temporal smoothing, and then synthesized with the temporal data of skeletal key points that pass the compliance judgment to form complete skeletal key point temporal data.

[0013] S6. Input the complete skeletal key point temporal data into the human motion assessment model based on the multimodal spatiotemporal pyramid network for processing, and quantify and output physical fitness assessment indicators.

[0014] S7. Automatically generate a physical fitness assessment report based on the assessment results, and output feedback on movement performance and optimization suggestions.

[0015] Optionally, step S1 further includes:

[0016] S11. Set up one or more video acquisition devices in a preset natural environment and initialize the parameters of the video acquisition devices. The video acquisition devices include fixed-point cameras, mobile cameras or depth cameras. The parameters of the video acquisition devices include camera angle θ, resolution R and exposure parameter E.

[0017] S12. The video acquisition device continuously acquires video data of the test subject throughout the entire process of the test preparation stage, the formal test stage and the test end stage according to the set time frequency f. The acquisition frequency f meets the requirements of the required action detail resolution and the acquisition frequency is not lower than the preset lower limit.

[0018] S13. Real-time recording of each video frame I(t) in the entire process video data. i The timestamp t i and all video frames The data is saved to a storage device in chronological order, where n represents the number of frames in the entire video data process. The storage device can be the built-in storage of the video acquisition device or an externally connected storage medium.

[0019] S14. For scenarios where multiple video acquisition devices work collaboratively, multiple video frames acquired at the same time are sorted according to their timestamps t. i Perform synchronization processing to obtain a time-aligned set of multi-view video frames {I} k (t i Let |k = 1, 2, ..., m}, where m is the number of cameras.

[0020] Optionally, step S2 further includes:

[0021] S21, transfer video frame I(t) i The algorithm detects key points of bones through a dynamic bone analysis algorithm, which is implemented by an improved spatiotemporal graph convolutional neural network. The improved spatiotemporal graph convolutional neural network uses a deformable three-dimensional convolutional neural network as the backbone and integrates a spatiotemporal graph convolutional module based on human kinematic chain topology as an anti-occlusion collaborative detection mechanism. Furthermore, the backbone network further incorporates a three-dimensional coordinate joint inference mechanism to generate a two-dimensional heat map, depth value, and confidence score of key points of bones.

[0022] S22. Process video frames using a deformable 3D convolutional neural network:

[0023] Deformable convolution kernels are introduced in the spatial dimension, and the sampling position of the convolution is dynamically adjusted according to the skeletal joint heatmap automatically generated by the network.

[0024] Introducing an inter-frame offset gating mechanism in the time dimension, specifically including using a Sigmoid gating function with a forget gate structure, dynamically constraining the sampling range of adjacent frames to within ±2 frames, and fusing real-time calculated optical flow features as motion sensing input;

[0025] A four-level feature pyramid structure is used to perform spatially sensitive region pooling and temporal difference feature extraction, generating an intermediate feature tensor that fuses spatiotemporal features.

[0026] S23. The intermediate feature tensor is processed by the spatiotemporal graph convolution module. When the skeletal joint is detected at t... i When occlusion occurs or the confidence level of the detection result is lower than the threshold, the anti-occlusion collaborative detection mechanism is activated according to the preset human motion chain topology. The spatial and temporal features of adjacent joints are communicated and fused, the spatial position of the occluded joint is automatically inferred, a virtual feature map is constructed and the features of the missing joints are filled in, and a corrected skeletal feature tensor is generated.

[0027] S24. The corrected skeletal feature tensor is processed through a three-dimensional coordinate joint calculation mechanism. Specifically, a heat map of each skeletal key point on the image plane is generated through a two-dimensional coordinate branch, and the two-dimensional coordinates (x, y, z) are output. j (t i ),y j (t i The depth value z of each skeletal keypoint relative to the torso reference plane is predicted using a depth regression branch. j (t i The confidence score C for the 3D position of each skeletal keypoint in each frame is output through the confidence evaluation branch. j (t i );

[0028] S25. Obtain and output each time frame t. i The corresponding 3D coordinates P of the skeletal key points j (t i )=(x j (t i ),y j (t i ),z j (t i )) and confidence level C j (t i );

[0029] S26. Construct a set of temporal motion trajectories for skeletal key points. Where j = 1, 2, ..., J, J represents the number of skeletal key points, and n represents the number of frames captured in the entire video data process;

[0030] S27. During the tester registration or initialization process, extract basic identity features F based on the temporal motion trajectory of skeletal key points. base (u), and the basic identity features F base (u) is stored in the identity feature database corresponding to the tester's identity information, where u = 1, 2, ..., N, and N represents the number of registered users.

[0031] Optionally, step S3 further includes:

[0032] S31. Set of temporal motion trajectories based on skeletal key points Identity features are extracted and tracked using a motion graph identity recognition network. The motion graph identity recognition network includes a dual-channel feature extraction architecture and a biometric feature fusion module. The dual-channel feature extraction architecture includes a spatial topology analysis channel and a temporal dynamic modeling channel.

[0033] S32. A hierarchical graph attention convolution is used to construct a spatial topology analysis channel. The set of temporal motion trajectories of skeletal key points is input. The hierarchical graph attention convolution dynamically adjusts the connection weights between skeletal joints through a learnable adjacency matrix and introduces a local-global feature hierarchical aggregation mechanism. The first layer processes anatomical connections, and the second layer automatically learns functional action associations to generate spatial topology features characterized by spatial skeletal topology relationships.

[0034] S33. A temporal dynamic modeling channel is constructed through a causal dilation gating unit. The set of temporal motion trajectories of skeletal key points is input. The causal dilation gating unit embeds a dilated temporal convolutional layer in the gating unit network. During the dilation temporal convolution operation, a forward mask is used to retain current and historical frame information while masking future frame feature inputs, thereby generating temporal dynamic features.

[0035] S34. Through the biometric fusion module, the spatial topological features F topo (t i ) and temporal dynamic characteristics F seq (t i The features are then merged to generate a merged identity feature. The fusion method is

[0036]

[0037] Among them, f att (·) denotes the attention fusion function, f crs (·) denotes the feature cross function, which is performed element-wise by the Hadamard product. Indicates feature splicing;

[0038] S35, Calculate and fuse identity features F dyn (t i ) and basic identity features F base The cosine similarity between (u) and (u) is used to determine if the identities are consistent when the similarity exceeds a set threshold.

[0039] S36. If all similarities are below the set threshold within a consecutive preset number of frames, they are determined to be abnormal skeletal key points, and abnormal handling or re-initialization of the tracking confirmation process is initiated.

[0040] S37. Output the set of temporal motion trajectories of skeletal key points that are determined to be consistent with the identity and the set of skeletal key points that are determined to be abnormal.

[0041] Optionally, step S4 further includes:

[0042] S41. The set of skeletal key points with consistent identities is segmented into action segments, and skeletal motion trajectory data within continuous time periods is extracted.

[0043] S42. Construct a motion evaluation tree based on biomechanical constraints. Each node of the motion evaluation tree corresponds to a combination of skeletal joints and a motion judgment rule. Perform state recognition and preliminary compliance judgment on the skeletal motion trajectory data of each motion segment to generate preliminary compliant skeletal motion trajectory data.

[0044] S43. Using a dynamic time warping algorithm, the preliminary compliant skeletal motion trajectory data is aligned with the key point temporal trajectory of the preset standardized action template, and the dynamic time warping matching degree is calculated. When the dynamic time warping matching degree is less than the threshold, it is determined to be a compliant action segment.

[0045] S45. Output the standardization judgment results of each action segment, including the skeletal key points of compliant action segments and the skeletal key points of non-compliant action segments.

[0046] Optionally, step S5 further includes:

[0047] S51. Combine the abnormal skeletal key points with the topological relationship of the human motion chain to determine the set of adjacent skeletal key points, and collect three-dimensional position information. Spatial correlation is completed by using the time motion vectors of adjacent points. The abnormal skeletal key points include skeletal key points that have been tracked and confirmed as abnormal and skeletal key points of non-compliant action segments.

[0048] S52. Using inverse kinematics reasoning, the optimal estimated position of the abnormal proximal bone point is calculated by using the spatial position of the distal joint directly connected to the abnormal bone key point. The position of the bone key point is completed by minimizing the objective function, and the reconstructed bone point trajectory is obtained. The objective function formula is:

[0049]

[0050] Among them, P j (t i () represents the joint point of the skeleton to be completed at time t. i Location, P k (t i ) represents the location of key points in adjacent bones, l jk Represents the static bone length vector between node j and its neighboring node k;

[0051] S53. The reconstructed skeletal point trajectory is smoothed temporally using a sliding weighted average filter to generate smooth skeletal key point temporal data.

[0052] S54. Integrate the smoothed skeletal keypoint timing data with the skeletal keypoint timing data of the compliant action segment to form complete skeletal keypoint timing data.

[0053] Optionally, step S6 further includes:

[0054] S61. The complete skeletal key point temporal data is processed through the input layer of the multimodal spatiotemporal pyramid network. The multimodal spatiotemporal pyramid network includes a four-level spatiotemporal pyramid convolutional structure, with each layer using a different dilation rate, and extracts short-term, long-term and multi-spatial scale features layer by layer.

[0055] The input layer generates joint angles, joint range of motion, and joint velocity based on the skeletal structure relationship, generates a motion energy heatmap based on the skeletal time trajectory, and judges the rationality and completeness of the generated joint angle sequence, joint range of motion, joint velocity, and motion energy heatmap. If anomalies, missing values, or physical parameters that do not conform to human physiological laws are detected, data correction processing is triggered.

[0056] S62. Use the corrected joint angle sequence, joint range of motion, joint velocity, and motion energy heatmap as multimodal data;

[0057] S63. Multimodal data is extracted and fused through a multimodal spatiotemporal pyramid network to obtain fused features. The fused features are then input into a four-level spatiotemporal pyramid convolutional structure, and multi-scale action features are extracted using spatiotemporal convolutions with different dilation rates.

[0058] S64. Based on multi-scale movement characteristics and integrating skeletal movement trajectory, joint angle sequence, joint range of motion and joint velocity, quantitatively output physical fitness assessment indicators, including movement standardization, movement range, movement speed and duration.

[0059] S65. Output various physical fitness assessment indicators and their corresponding quantitative scores.

[0060] Optionally, step S7 includes:

[0061] S71. Set reference intervals according to various evaluation indicators, standardize the indicator scores of the test subjects, and automatically generate a physical fitness evaluation report based on the test subjects' scores and movement trajectory analysis results. The report includes movement quality scores, movement compliance analysis, key movement suggestions, and risk warnings.

[0062] S72. For segments with non-standard movements or physical fitness indicators below the threshold, automatically label the corresponding time period and joint type, and output personalized movement optimization suggestions and training feedback.

[0063] S73. Save the evaluation data and reports, and support exporting them as PDF documents, data interfaces, or platform push methods.

[0064] A physical fitness assessment system based on human skeletal trajectory tracking according to an embodiment of the present invention includes:

[0065] The video acquisition module is used to acquire video data of the entire physical fitness test process and generate video frames according to the set acquisition frequency.

[0066] The skeleton parsing module is used to detect and extract key points of the human skeleton from video frames and construct a set of temporal motion trajectories of the key points of the skeleton.

[0067] The identity tracking module is used for tester identification and tracking based on skeletal time-series trajectories.

[0068] The action recognition and segmentation module is used to perform action recognition, action segmentation, and compliance determination on the temporal motion trajectory of the tracked and confirmed skeletal key points.

[0069] The skeleton completion module is used to complete the key points of abnormal skeletons and perform temporal smoothing.

[0070] The motion assessment module is used to output physical fitness assessment indicators and generate corresponding quantitative scores;

[0071] The Reporting and Feedback module is used to generate physical fitness assessment reports and output feedback and optimization suggestions.

[0072] The beneficial effects of this invention are:

[0073] First, this invention utilizes multi-view, full-process video acquisition and dynamic skeletal analysis technology to achieve temporal tracking and automated identification of key skeletal points in the human body. This effectively overcomes the limitations of traditional methods, which rely on expensive hardware and restricted venues, thus improving the convenience and accessibility of physical fitness testing. Second, by introducing a multimodal spatiotemporal pyramid network, the system integrates skeletal key point data with multi-source data such as biomechanical parameters and kinetic energy. This enables a comprehensive and detailed quantitative assessment of physical fitness indicators such as movement standardization, amplitude, speed, and duration, significantly improving the scientific rigor and accuracy of movement analysis. Furthermore, the system possesses the ability to intelligently complete abnormal skeletal points, automatically generate physical fitness assessment reports, and provide personalized feedback suggestions, significantly enhancing the reliability of the assessment results. In summary, this invention achieves automation, precision, and intelligence in the physical fitness assessment process. It boasts advantages such as low cost, high applicability, and comprehensive assessment, effectively promoting the widespread application of physical fitness testing in public health, sports training, and medical rehabilitation. Attached Figure Description

[0074] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0075] Fig. 1 This is a schematic diagram of a physical fitness assessment method and system based on human skeletal trajectory tracking proposed in this invention;

[0076] Fig. 2 This is a flowchart of the skeletal key point extraction and trajectory construction process in this invention;

[0077] Fig. 3 This is a schematic diagram of the core process of physical fitness assessment based on skeletal trajectory in this invention. Detailed Implementation

[0078] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0079] refer to Figs. 1-3 A physical fitness assessment method based on human skeletal trajectory tracking includes the following steps:

[0080] S1. Acquire video data of the entire physical fitness test process through video acquisition equipment, and generate video frames according to the acquisition frequency;

[0081] S2. Process video frames using a skeletal dynamic analysis algorithm to construct a set of temporal motion trajectories of skeletal key points;

[0082] S3. Based on the temporal motion trajectory set of skeletal key points, extract the identity features of the tester through the motion graph identity recognition network throughout the test process, and perform identity tracking and confirmation.

[0083] S4. Perform preliminary motion recognition on the time-series motion trajectory set of the tracked skeletal key points, segment and extract motion segments, and determine the compliance of each motion segment.

[0084] S5. For abnormal skeletal key points that do not pass identity tracking confirmation and compliance judgment, the inverse kinematics reasoning method is used to complete the position, the completion result is processed by temporal smoothing, and then synthesized with the temporal data of skeletal key points that pass the compliance judgment to form complete skeletal key point temporal data.

[0085] S6. Input the complete skeletal key point temporal data into the human motion assessment model based on the multimodal spatiotemporal pyramid network for processing, and quantify and output physical fitness assessment indicators.

[0086] S7. Automatically generate a physical fitness assessment report based on the assessment results, and output feedback on movement performance and optimization suggestions.

[0087] In this embodiment, step S1 further includes:

[0088] S11. Set up one or more video acquisition devices in a preset natural environment and initialize the parameters of the video acquisition devices. The video acquisition devices include fixed-point cameras, mobile cameras or depth cameras. The parameters of the video acquisition devices include camera angle θ, resolution R and exposure parameter E.

[0089] S12. The video acquisition device continuously acquires video data of the test subject throughout the entire process of the test preparation stage, the formal test stage and the test end stage according to the set time frequency f. The acquisition frequency f meets the requirements of the required action detail resolution and the acquisition frequency is not lower than the preset lower limit.

[0090] S13. Real-time recording of each video frame I(t) in the entire process video data. i The timestamp t i and all video frames The data is saved to a storage device in chronological order, where n represents the number of frames in the entire video data process. The storage device can be the built-in storage of the video acquisition device or an externally connected storage medium.

[0091] S14. For scenarios where multiple video acquisition devices work collaboratively, multiple video frames acquired at the same time are sorted according to their timestamps t. i Perform synchronization processing to obtain a time-aligned set of multi-view video frames {I} k (t iLet |k = 1, 2, ..., m}, where m is the number of cameras.

[0092] In this embodiment, step S2 further includes:

[0093] S21, transfer video frame I(t) i The algorithm detects key points of bones through a dynamic bone analysis algorithm, which is implemented by an improved spatiotemporal graph convolutional neural network. The improved spatiotemporal graph convolutional neural network uses a deformable three-dimensional convolutional neural network as the backbone and integrates a spatiotemporal graph convolutional module based on human kinematic chain topology as an anti-occlusion collaborative detection mechanism. Furthermore, the backbone network further incorporates a three-dimensional coordinate joint inference mechanism to generate a two-dimensional heat map, depth value, and confidence score of key points of bones.

[0094] S22. Process video frames using a deformable 3D convolutional neural network:

[0095] Deformable convolution kernels are introduced in the spatial dimension, and the sampling position of the convolution is dynamically adjusted according to the skeletal joint heatmap automatically generated by the network.

[0096] Introducing an inter-frame offset gating mechanism in the time dimension, specifically including using a Sigmoid gating function with a forget gate structure, dynamically constraining the sampling range of adjacent frames to within ±2 frames, and fusing real-time calculated optical flow features as motion sensing input;

[0097] A four-level feature pyramid structure is used to perform spatially sensitive region pooling and temporal difference feature extraction, generating an intermediate feature tensor that fuses spatiotemporal features.

[0098] S23. The intermediate feature tensor is processed by the spatiotemporal graph convolution module. When the skeletal joint is detected at t... i When occlusion occurs or the confidence level of the detection result is lower than the threshold, the anti-occlusion collaborative detection mechanism is activated according to the preset human motion chain topology. The spatial and temporal features of adjacent joints are communicated and fused, the spatial position of the occluded joint is automatically inferred, a virtual feature map is constructed and the features of the missing joints are filled in, and a corrected skeletal feature tensor is generated.

[0099] S24. The corrected skeletal feature tensor is processed through a three-dimensional coordinate joint calculation mechanism. Specifically, a heat map of each skeletal key point on the image plane is generated through a two-dimensional coordinate branch, and the two-dimensional coordinates (x, y, z) are output. j (t i ),y j (t i The depth value z of each skeletal keypoint relative to the torso reference plane is predicted using a depth regression branch. j (t iThe confidence score C for the 3D position of each skeletal keypoint in each frame is output through the confidence evaluation branch. j (t i );

[0100] S25. Obtain and output each time frame t. i The corresponding 3D coordinates P of the skeletal key points j (t i )=(x j (t i ),y j (t i ),z j (t i )) and confidence level C j (t i );

[0101] S26. Construct a set of temporal motion trajectories for skeletal key points. Where j = 1, 2, ..., J, J represents the number of skeletal key points, and n represents the number of frames captured in the entire video data process;

[0102] S27. During the tester registration or initialization process, extract basic identity features F based on the temporal motion trajectory of skeletal key points. base (u), and the basic identity features F base (u) is stored in the identity feature database corresponding to the tester's identity information, where u = 1, 2, ..., N, and N represents the number of registered users.

[0103] In this embodiment, step S3 further includes:

[0104] S31. Set of temporal motion trajectories based on skeletal key points Identity features are extracted and tracked using a motion graph identity recognition network. The motion graph identity recognition network includes a dual-channel feature extraction architecture and a biometric feature fusion module. The dual-channel feature extraction architecture includes a spatial topology analysis channel and a temporal dynamic modeling channel.

[0105] S32. A hierarchical graph attention convolution is used to construct a spatial topology analysis channel. The set of temporal motion trajectories of skeletal key points is input. The hierarchical graph attention convolution dynamically adjusts the connection weights between skeletal joints through a learnable adjacency matrix and introduces a local-global feature hierarchical aggregation mechanism. The first layer processes anatomical connections, and the second layer automatically learns functional action associations to generate spatial topology features characterized by spatial skeletal topology relationships.

[0106] S33. A temporal dynamic modeling channel is constructed through a causal dilation gating unit. The set of temporal motion trajectories of skeletal key points is input. The causal dilation gating unit embeds a dilated temporal convolutional layer in the gating unit network. During the dilation temporal convolution operation, a forward mask is used to retain current and historical frame information while masking future frame feature inputs, thereby generating temporal dynamic features.

[0107] S34. Through the biometric fusion module, the spatial topological features F topo (t i ) and temporal dynamic characteristics F seq (t i The features are then merged to generate a merged identity feature. The fusion method is

[0108]

[0109] Among them, f att (·) denotes the attention fusion function, f crs (·) denotes the feature cross function, which is performed element-wise by the Hadamard product. Indicates feature splicing;

[0110] S35, Calculate and fuse identity features F dyn (t i ) and basic identity features F base The cosine similarity between (u) and (u) is used to determine if the identities are consistent when the similarity exceeds a set threshold.

[0111] S36. If all similarities are below the set threshold within a consecutive preset number of frames, they are determined to be abnormal skeletal key points, and abnormal handling or re-initialization of the tracking confirmation process is initiated.

[0112] S37. Output the set of temporal motion trajectories of skeletal key points that are determined to be consistent with the identity and the set of skeletal key points that are determined to be abnormal.

[0113] In this embodiment, step S4 further includes:

[0114] S41. The set of skeletal key points with consistent identities is segmented into action segments, and skeletal motion trajectory data within continuous time periods is extracted.

[0115] S42. Construct a motion evaluation tree based on biomechanical constraints. Each node of the motion evaluation tree corresponds to a combination of skeletal joints and a motion judgment rule. Perform state recognition and preliminary compliance judgment on the skeletal motion trajectory data of each motion segment to generate preliminary compliant skeletal motion trajectory data.

[0116] S43. Using a dynamic time warping algorithm, the preliminary compliant skeletal motion trajectory data is aligned with the key point temporal trajectory of the preset standardized action template, and the dynamic time warping matching degree is calculated. When the dynamic time warping matching degree is less than the threshold, it is determined to be a compliant action segment.

[0117] S45. Output the standardization judgment results of each action segment, including the skeletal key points of compliant action segments and the skeletal key points of non-compliant action segments.

[0118] In this embodiment, step S5 further includes:

[0119] S51. Combine the abnormal skeletal key points with the topological relationship of the human motion chain to determine the set of adjacent skeletal key points, and collect three-dimensional position information. Spatial correlation is completed by using the time motion vectors of adjacent points. The abnormal skeletal key points include skeletal key points that have been tracked and confirmed as abnormal and skeletal key points of non-compliant action segments.

[0120] S52. Using inverse kinematics reasoning, the optimal estimated position of the abnormal proximal bone point is calculated by using the spatial position of the distal joint directly connected to the abnormal bone key point. The position of the bone key point is completed by minimizing the objective function, and the reconstructed bone point trajectory is obtained. The objective function formula is:

[0121]

[0122] Among them, P j (t i () represents the joint point of the skeleton to be completed at time t. i Location, P k (t i ) represents the location of key points in adjacent bones, l jk Represents the static bone length vector between node j and its neighboring node k;

[0123] S53. The reconstructed skeletal point trajectory is smoothed temporally using a sliding weighted average filter to generate smooth skeletal key point temporal data.

[0124] S54. Integrate the smoothed skeletal keypoint timing data with the skeletal keypoint timing data of the compliant action segment to form complete skeletal keypoint timing data.

[0125] In this embodiment, step S6 further includes:

[0126] S61. The complete skeletal key point temporal data is processed through the input layer of the multimodal spatiotemporal pyramid network. The multimodal spatiotemporal pyramid network includes a four-level spatiotemporal pyramid convolutional structure, with each layer using a different dilation rate, and extracts short-term, long-term and multi-spatial scale features layer by layer.

[0127] The input layer generates joint angles, joint range of motion, and joint velocity based on the skeletal structure relationship, generates a motion energy heatmap based on the skeletal time trajectory, and judges the rationality and completeness of the generated joint angle sequence, joint range of motion, joint velocity, and motion energy heatmap. If anomalies, missing values, or physical parameters that do not conform to human physiological laws are detected, data correction processing is triggered.

[0128] S62. Use the corrected joint angle sequence, joint range of motion, joint velocity, and motion energy heatmap as multimodal data;

[0129] S63. Multimodal data is extracted and fused through a multimodal spatiotemporal pyramid network to obtain fused features. The fused features are then input into a four-level spatiotemporal pyramid convolutional structure, and multi-scale action features are extracted using spatiotemporal convolutions with different dilation rates.

[0130] S64. Based on multi-scale movement characteristics and integrating skeletal movement trajectory, joint angle sequence, joint range of motion and joint velocity, quantitatively output physical fitness assessment indicators, including movement standardization, movement range, movement speed and duration.

[0131] S65. Output various physical fitness assessment indicators and their corresponding quantitative scores.

[0132] In this embodiment, step S7 includes:

[0133] S71. Set reference intervals according to various evaluation indicators, standardize the indicator scores of the test subjects, and automatically generate a physical fitness evaluation report based on the test subjects' scores and movement trajectory analysis results. The report includes movement quality scores, movement compliance analysis, key movement suggestions, and risk warnings.

[0134] S72. For segments with non-standard movements or physical fitness indicators below the threshold, automatically label the corresponding time period and joint type, and output personalized movement optimization suggestions and training feedback.

[0135] S73. Save the evaluation data and reports, and support exporting them as PDF documents, data interfaces, or platform push methods.

[0136] A physical fitness assessment system based on human skeletal trajectory tracking, comprising:

[0137] The video acquisition module is used to acquire video data of the entire physical fitness test process and generate video frames according to the set acquisition frequency.

[0138] The skeleton parsing module is used to detect and extract key points of the human skeleton from video frames and construct a set of temporal motion trajectories of the key points of the skeleton.

[0139] The identity tracking module is used for tester identification and tracking based on skeletal time-series trajectories.

[0140] The action recognition and segmentation module is used to perform action recognition, action segmentation, and compliance determination on the temporal motion trajectory of the tracked and confirmed skeletal key points.

[0141] The skeleton completion module is used to complete the key points of abnormal skeletons and perform temporal smoothing.

[0142] The motion assessment module is used to output physical fitness assessment indicators and generate corresponding quantitative scores;

[0143] The Reporting and Feedback module is used to generate physical fitness assessment reports and output feedback and optimization suggestions.

[0144] Example 1:

[0145] To verify the effectiveness of this invention in practical applications, it was applied to a physical fitness assessment activity for university students. A total of 120 students were recruited for the assessment, with a relatively balanced gender ratio, representing individuals with different exercise habits, including those who exercise regularly, those who exercise occasionally, and those who exercise almost never. The testing venue was a spacious and well-lit indoor sports space, equipped with multiple high-definition cameras to achieve multi-angle video data acquisition. Before acquisition, all devices underwent time synchronization processing to ensure accurate and continuous motion capture. The physical fitness assessment included squats, push-ups, vertical jumps, and high knees. The entire test was automatically recorded by the camera system and uploaded to a processing platform.

[0146] Each participant registered their identity and had their standard posture recorded before the test, then performed various physical fitness exercises in sequence. Real-time multi-view video data was first input into the skeletal analysis module of this invention. The system automatically detected and precisely tracked at least 17 key points of the human skeleton, while simultaneously performing skeletal point completion and data correction for situations such as short-term occlusion, ensuring the complete and accurate skeletal trajectory of each movement. Subsequently, the system extracted biomechanical parameters such as joint angles, range of motion, and velocity, and combined them with the skeletal movement trajectory to generate an energy heatmap as multimodal input. A movement evaluation model was then used to quantitatively analyze indicators such as movement standardization, amplitude, speed, and duration. After the evaluation, the system automatically generated a detailed physical fitness assessment report and provided the participant with targeted movement feedback and training suggestions.

[0147] In this test, the system demonstrated strong adaptability to situations such as partial occlusion and simultaneous movement of multiple people that may occur in real-world scenarios. Experimental results show that the system can automatically detect abnormal skeletal points and complete them through inverse kinematics and temporal filtering, ensuring the continuity and integrity of motion data. In contrast, when using traditional single-view two-dimensional recognition schemes for parallel data analysis,

[0148] In this test, complex environmental conditions such as partial occlusion and simultaneous movement of multiple people were set up to compare and analyze the processing capabilities of different technical solutions. Additionally, a traditional single-view 2D motion recognition method was simultaneously applied as a control. Through parallel analysis of data from the same scene and with the same actions, the performance of the present invention and the traditional method was compared in terms of motion accuracy scoring, skeletal point error (mm), and positioning error (seconds). This further reflects the advantages and disadvantages of the two methods in terms of motion capture completeness, skeletal keypoint tracking accuracy, and automatic data completion capabilities. Table 1 shows the performance comparison data of the present invention system and the traditional method in motion evaluation.

[0149] Table 1 Comparison of motion evaluation performance between the system of this invention and traditional methods.

[0150]

[0151] Data shows that in all test samples, the system of this invention generally scored higher than the traditional method in standardization scores for common physical fitness exercises such as squats, vertical jumps, push-ups, and high knees, with an average difference of 3 to 5 points. This indicates that the system of this invention can more accurately capture movement details and make scientific judgments. For example, in the standardization scores of movements numbered 01 and 03, the system of this invention achieved scores of 89 and 91 respectively, while the traditional method only achieved scores of 86 and 88, demonstrating that the movement quality judgment results of this invention are more stable and reliable.

[0152] Regarding skeletal point error, the skeletal point error of the system of this invention is generally less than 7.5 mm in different movements and in subjects of different genders, with a minimum of 6.6 mm, while the skeletal point error of traditional methods is generally between 21 mm and 24 mm. Taking samples 02 and 08 as examples, the skeletal point errors using traditional methods are 23.3 mm and 21.4 mm, respectively, while the corresponding errors of the system of this invention are only 7.2 mm and 6.6 mm, indicating that the present invention improves the ability to capture the dynamic features of the subject.

[0153] Regarding positioning error, the system of this invention automatically identifies the start and end times of actions, and the error is always controlled between 0.13 seconds and 0.15 seconds, while the positioning error of traditional methods mostly exceeds 0.60 seconds. This shows that the present invention effectively ensures the automation of the entire evaluation process and greatly improves the segmentation accuracy of each action unit.

[0154] Furthermore, the system of this invention can automatically quantify and output the amplitude score, speed score, and duration of the movement, while traditional methods, limited by their technical principles, struggle to obtain such quantitative data. Taking the vertical jump item (number 03) as an example, the system of this invention scores 90 points for amplitude, 89 points for speed, and 13.5 seconds for duration, comprehensively reflecting the subject's movement performance and physical fitness.

[0155] In summary, the system of this invention outperforms traditional solutions in terms of motion recognition, skeletal tracking, and automatic data completion. It can effectively reduce the data loss rate in complex environments, improve the completeness and scientific nature of physical fitness assessment, and has the advantages of accurate and efficient physical fitness assessment in diverse and complex scenarios.

[0156] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A physical ability evaluation method based on human skeleton trajectory tracking, characterized in that, The method comprises the following steps: S1, acquiring whole-process video data of physical fitness testing through a video acquisition device, and generating video frames according to a collection frequency; S2, processing the video frames through a skeletal dynamic analysis algorithm to construct a skeletal key point time sequence motion trajectory set, and the S2 specifically comprises: S21, the video frame The skeleton key points are detected through a skeleton dynamic analysis algorithm, and the skeleton dynamic analysis algorithm is implemented through an improved space-time graph convolutional neural network. S22, processing the video frames through a deformable three-dimensional convolutional neural network: Introducing a deformable convolution kernel in a spatial dimension, and dynamically adjusting the sampling position of the convolution according to the skeletal joint heat map automatically generated by the network; Introducing an inter-frame offset gating mechanism in a time dimension, specifically comprising adopting a sigmoid gating function with a forgetting gate structure, dynamically restricting the sampling range of adjacent frames within ±2 frames, and fusing the real-time calculated optical flow features as motion perception inputs; Respectively performing spatial sensitive region pooling and time sequence difference feature extraction through a four-level feature pyramid structure to generate an intermediate feature tensor fusing space-time features; S23. The intermediate feature tensor is processed by the spatiotemporal graph convolution module. When skeletal joints are detected... When occlusion occurs or the confidence level of the detection result is lower than the threshold, the anti-occlusion collaborative detection mechanism is activated according to the preset human motion chain topology. The spatial and temporal features of adjacent joints are communicated and fused, the spatial position of the occluded joint is automatically inferred, a virtual feature map is constructed and the features of the missing joints are filled in, and a corrected skeletal feature tensor is generated. S24, the corrected bone feature tensor is processed through a three-dimensional coordinate joint calculation mechanism, specifically: a heat distribution map of each bone key point in the image plane is generated through a two-dimensional coordinate branch, and a two-dimensional coordinate is output , a depth value of each bone key point relative to the trunk reference plane is predicted through a depth regression branch , a three-dimensional position confidence score of each bone key point in each frame is output through a reliability evaluation branch ; S25, obtaining and outputting each time frame three-dimensional coordinates of corresponding skeletal keypoints and a confidence ; S26, construct a skeletal key point timing motion trajectory set wherein, , represents the number of skeletal key points, represents the number of collected frames of the whole-process video data; S27、In the registration or initialization of the tester, the basic identity feature is extracted according to the skeletal key point time sequence motion trajectory , and the basic identity feature is stored in the identity feature library corresponding to the tester identity information, wherein, , represents the number of registered users; S3, based on the skeletal key point time sequence motion trajectory set, extracting the identity features of the tester through a motion atlas identity recognition network within the whole testing process, and performing identity tracking confirmation, and the S3 specifically comprises: S31, the set of timing motion trajectories based on the skeleton key points An identity feature is extracted by using a motion atlas identity recognition network, and tracking confirmation is performed, the motion atlas identity recognition network comprises a double-channel feature extraction architecture and a biological feature fusion module, the double-channel feature extraction architecture comprises a spatial topology analysis channel and a timing dynamic modeling channel. S32, adopt hierarchical graph attention convolution to construct a spatial topology analysis channel, input a time sequence motion trajectory set of the skeleton key points, and generate spatial topology features ; S33, constructing a time sequence dynamic modeling channel through a causal inflation gate unit, inputting a time sequence motion trajectory set of the skeleton key points, and generating a time sequence dynamic feature ; S34, fusing the spatial topological features by a biometric fusion module with the temporal dynamic features to generate a fused identity feature : ; wherein, denotes an attention fusion function, denotes a feature cross function, which takes an element-wise product with a Hadamard product, denotes a feature concatenation; S35、calculate the cosine similarity between the fusion identity feature and the base identity feature and the base identity feature When the similarity exceeds a set threshold, identity consistency is determined. S36, if all the similarities are lower than the set threshold within a continuous preset number of frames, it is determined that the skeletal key points are abnormal, and an abnormal processing or re-initialization tracking confirmation process is started; S37, outputting the skeletal key point time sequence motion trajectory set determined as having the same identity and the skeletal key points determined as abnormal; S4, performing preliminary action recognition on the tracked skeletal key point time sequence motion trajectory set, segmenting and extracting action segments, and performing compliance determination on each action segment; S5, for the abnormal skeletal key points that do not pass the identity tracking confirmation and compliance determination, a reverse kinematics reasoning method is used for position completion, time sequence smoothing processing is performed on the completion result, and the complete skeletal key point time sequence data is synthesized with the skeletal key point time sequence data passing the compliance determination, and the S5 specifically comprises: S51, combining the abnormal skeletal key points with the human motion chain topological relationship, determining a set of adjacent skeletal key points, and collecting three-dimensional position information, performing spatial correlation completion through the time motion vector of adjacent points, and the abnormal skeletal key points include the skeletal key points determined as abnormal by the tracking confirmation and the skeletal key points of the non-compliant action segment; S52, using a reverse kinematics reasoning method, using the spatial position of the end joint directly connected with the abnormal skeletal key point to reversely calculate the optimal estimated position of the abnormal proximal skeletal point, completing the skeletal key point position by minimizing the objective function, and obtaining the reconstructed skeletal point trajectory: ; wherein, is the position of the skeleton joint to be completed at time , is the position of the adjacent skeleton key point, denotes the static bone length vector of the joint to the adjacent joint . S53, performing time sequence smoothing processing on the reconstructed skeletal point trajectory through a sliding weighted average filter to generate smoothed skeletal key point time sequence data; S54, integrating the smoothed skeletal key point time sequence data with the skeletal key point time sequence data of the compliant action segment to form complete skeletal key point time sequence data; S6, inputting the complete skeletal key point time sequence data into a human action evaluation model based on a multi-modal space-time pyramid network for processing, and quantitatively outputting physical fitness evaluation indexes. S7, automatically generate a physical fitness evaluation report according to the evaluation results, and output action performance feedback and optimization suggestions.

2. The physical ability evaluation method based on human skeleton trajectory tracking according to claim 1, characterized in that, The S1 further comprises: S11, arranging one or more video acquisition devices in a preset natural environment, and initializing parameters of the video acquisition devices; S12, collect the video acquisition device according to the set time frequency Continuously collect the video data of the tester in the whole process of the test preparation stage, the formal test stage and the test end stage. S13, record each frame of the real-time video data timestamp , and save all video frames in chronological order to the storage device, wherein, indicates the number of frames of the whole-process video data; S14, for the scene of multiple video acquisition devices working cooperatively, multiple video frames collected at the same time are synchronized according to timestamps to obtain a time-aligned multi-view video frame set wherein, is the number of cameras. 3.The physical ability evaluation method based on human skeleton trajectory tracking according to claim 1, wherein, The S4 further comprises: S41, segmenting the set of skeletal key point time series motion trajectories determined to be consistent in identity into action segments, and extracting skeletal motion trajectory data in continuous time periods; S42, constructing an action evaluation tree based on biomechanical constraints, each node of the action evaluation tree corresponding to a combination of skeletal joints and an action determination rule, performing state recognition and preliminary compliance determination on the skeletal motion trajectory data of each action segment, and generating preliminary compliant skeletal motion trajectory data; S43, using a dynamic time warping algorithm to align the preliminary compliant skeletal motion trajectory data with the key point time series trajectories of the preset standard action templates, calculating a dynamic time warping matching degree, and determining a compliant action segment when the dynamic time warping matching degree is less than a threshold value; S45, outputting the standardization determination results of each action segment, including the skeletal key points of compliant action segments and the skeletal key points of non-compliant action segments.

4. The physical ability evaluation method based on human skeleton trajectory tracking according to claim 1, characterized in that, The S6 further comprises: S61, processing the complete skeletal key point time series data through a multi-modal spatio-temporal pyramid network, the input layer of the multi-modal spatio-temporal pyramid network generating joint angles, joint ranges of motion, and joint speeds according to skeletal structure relationships, generating a motion energy heat map according to skeletal time series trajectories, and performing rationality and integrity judgment on the generated joint angle sequences, joint ranges of motion, joint speeds, and motion energy heat maps, and if abnormalities, missing, or physical parameters are detected to be inconsistent with human physiological rules, triggering data correction processing; S62, taking the corrected joint angle sequences, joint ranges of motion, joint speeds, and motion energy heat maps as multi-modal data; S63, extracting and fusing features of the multi-modal data through the multi-modal spatio-temporal pyramid network to obtain fused features, and inputting the fused features into a four-level spatio-temporal pyramid convolution structure to extract multi-scale action features using spatio-temporal convolutions with different dilation rates; S64, based on the multi-scale action features, and integrating the skeletal motion trajectory, joint angle sequence, joint range of motion, and joint speed, quantitatively outputting physical fitness evaluation indicators, the physical fitness evaluation indicators including action standardization, action amplitude, action speed, and duration; S65, outputting each physical fitness evaluation indicator and the corresponding quantitative score. 5.The physical ability evaluation method based on human skeleton trajectory tracking according to claim 1, wherein, The S7 comprises: According to each evaluation indicator, a reference interval is set, the indicator scores of the tested person are standardized, and based on the scores and motion trajectory analysis results of the tested person, a physical fitness evaluation report is automatically generated, and for segments with non-standard actions or physical fitness indicators below a threshold value, the corresponding time period and joint type are automatically labeled, and individualized action optimization suggestions and training feedback are outputted.

6. A physical ability evaluation system based on human skeleton trajectory tracking, which executes the physical ability evaluation method based on human skeleton trajectory tracking according to any one of claims 1 to 5, characterized by, It comprises: A video acquisition module for acquiring full-process video data of physical fitness testing and generating video frames according to a set acquisition frequency; A skeletal analysis module for detecting and extracting human skeletal key points from video frames to construct a set of skeletal key point time series motion trajectories; An identity tracking module is configured to identify and track the identity of the tester based on the skeletal time-series trajectory; An action recognition and segment division module is configured to recognize actions, divide action segments, and determine compliance based on the time-series motion trajectory of the tracked and confirmed skeletal key points; A skeletal completion module is configured to complete abnormal skeletal key points and perform time-series smoothing processing; An action evaluation module is configured to output physical fitness evaluation indicators and generate corresponding quantitative scores; A report and feedback module is configured to generate a physical fitness evaluation report and output feedback and optimization suggestions.

Citation Information

Patent Citations

  • User identity recognition method and system in combination with user gait information

    CN112101176A

  • Physical fitness testing method and device, computer equipment and storage medium

    CN114708541A

  • Video action detection method based on convolutional neural network

    US20200057935A1