A sports action evaluation method based on a space-time diagram attention network
By constructing a human kinetic chain diagram and a spatiotemporal attention network, we can achieve accurate assessment of the quality of sports movements and trace the source of deviations. This solves the problem of inaccurate assessment in existing technologies and improves the fine-grainedness and multi-dimensional analysis capabilities of the assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RONGMENGYUESHI (SHANGHAI) SPORTS TECHNOLOGY CO LTD
- Filing Date
- 2026-05-09
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies are insufficient for objectively and multidimensionally evaluating the quality of sports movements, especially in handling the complex spatial relationships and dynamic changes over time between human joints. Furthermore, traditional methods are unable to reflect the structural characteristics of sports movements at different stages.
By constructing a human kinetic chain diagram and using a spatiotemporal graph attention network to segment movement stages and locate deviation sources, continuous motion capture data is acquired to establish a human kinetic chain diagram. Combined with the spatiotemporal graph attention network, movement stages are segmented and deviation sources are traced, thereby achieving accurate evaluation of the quality of sports movements.
It improves the accuracy and fine-grained analysis capabilities of sports movement assessment, can accurately characterize the synergistic force relationship between joints, and provides source analysis of movement deviations, solving the problem that traditional methods are insufficient in assessing the spatial correlation between joints and the characteristics of movement stages.
Smart Images

Figure CN122440172A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and sports data analysis technology, specifically to a sports motion assessment method based on spatiotemporal graph attention networks. Background Technology
[0002] With the development of computer vision, deep learning, and motion sensing technologies, the digital analysis and intelligent evaluation of sports movements have gradually become an important research direction in fields such as sports training, competitive sports, physical fitness testing, and rehabilitation training. Traditional sports movement evaluation usually relies on subjective judgment by coaches or experts through visual observation. This method not only depends heavily on the experience of the evaluators but also makes it difficult to accurately quantify the subtle differences in complex movements. Furthermore, manual evaluation is inefficient in scenarios with high movement frequency or multiple people training simultaneously, failing to meet the demands of modern sports training for objective, standardized, and real-time evaluation.
[0003] In existing technologies, sports motion capture typically acquires human joint motion data using devices such as optical motion capture systems, depth cameras, or inertial sensors, and reconstructs human posture using skeletal models. Building upon this, some studies utilize machine learning or deep learning models to identify or classify movements, such as motion recognition methods based on convolutional neural networks, recurrent neural networks, or graph convolutional networks. However, most existing methods focus on determining the movement category or identifying whether a movement has occurred, and their ability to evaluate movement execution quality, the degree of movement standardization, and the coordination of key joints remains limited. Furthermore, traditional algorithms still have shortcomings in handling the complex spatial relationships between human joints and the dynamic changes of movements over time, making it difficult to comprehensively reflect the structural characteristics of sports movements at different stages.
[0004] Furthermore, there are often reasonable differences in athletic movements among individuals. Factors such as height, limb proportions, and exercise habits can all affect the performance of movements, making it difficult to accurately assess movement quality using simple, fixed template comparison methods. Some existing movement evaluation systems typically determine deviations by matching the movement to a single standard movement model. However, this approach easily overlooks temporal rhythm changes during the movement and the synergistic relationships between key force chains, resulting in insufficient stability of the evaluation results. Therefore, how to construct an intelligent analysis model based on motion capture data that can simultaneously characterize the spatial relationships and temporal dynamics of human joints, and on this basis, achieve objective and multi-dimensional evaluation of the quality of athletic movement completion, remains a pressing technical problem to be solved in this field.
[0005] To address the above issues, this application proposes a sports motion assessment method based on a spatiotemporal graph attention network. Summary of the Invention
[0006] To address the problem that existing technologies struggle to objectively and multidimensionally evaluate the quality of sports movements, this application proposes a sports movement evaluation method based on a spatiotemporal graph attention network. This method constructs a human kinetic chain diagram and uses the spatiotemporal graph attention network to segment movement stages and locate deviation sources, thereby achieving accurate evaluation of sports movement quality and tracing of deviation sources.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A method for evaluating physical activity based on a spatiotemporal graph attention network, the method comprising: Acquire continuous motion capture data of the test subject during the execution of the target sports movement, wherein the continuous motion capture data includes at least the spatial pose information of multiple joints at consecutive moments; Based on the human kinematic connection relationship between the multiple joints and the joint synergistic force relationship corresponding to the target sports movement, a human kinetic chain diagram is constructed, and the continuous motion capture data is processed to obtain the spatiotemporal representation sequence of the movement to be tested. Based on the spatiotemporal representation sequence of the action to be tested, the action stages are divided to obtain at least two action stages corresponding to the target sports action; By combining the human kinetic chain diagram, the sources of motion deviation in each stage of the movement are identified, and the evaluation results of the test subject's performance of the target sports movement are determined based on the associated paths of the motion deviation sources in the human kinetic chain diagram.
[0008] The methods for acquiring the spatial pose information include: The process of the subject performing the target sports action is collected by at least one motion acquisition device to obtain corresponding motion observation data; Based on the motion observation data, the joint point observation results of the tested object at each sampling time are extracted; After completing the equipment calibration, the aligned joint point observation results are converted to a unified spatial reference coordinate system to obtain the spatial position data of multiple joint points at consecutive times. Based on the spatial location data, the human skeleton pose is restored to obtain the spatial pose information of the multiple joints at continuous time points.
[0009] The human skeleton pose recovery is achieved by combining the topological constraints of the human skeleton, the continuity constraints of joint motion at adjacent time points, and the observation confidence of joint points to jointly solve the spatial position and joint orientation of each joint.
[0010] The method for constructing the human kinetic chain diagram includes: Based on the human kinematic connection relationship between the multiple joints, a basic skeleton graph containing multiple joints and joint connection edges is established, wherein each joint is used as a graph node, and each joint connection edge is used to characterize the basic kinematic connection relationship between the corresponding joints. Based on the action type, force exertion part, and dominant coordination relationship in the execution process of the target sports action, at least one target kinetic chain path is determined from the basic skeleton diagram, and a coordination force exertion edge is established between the associated joints on the target kinetic chain path to characterize the linkage and transmission relationship between non-adjacent joints in the execution process of the target sports action. Based on the force exertion sequence and stability control requirements of the target sports movement in different movement stages, stage activation identifiers and link weight information are configured for the basic skeleton diagram and the cooperative force exertion edges, respectively, to characterize the effective participation degree of each joint and each edge in different movement stages, thus obtaining the human kinetic chain diagram.
[0011] The methods for determining the target kinetic chain path include: Determine the dominant force-generating part, power transmission part, and end-effector control part corresponding to the target sports movement; Based on the distribution of the main force-generating parts, power transmission parts, and end control parts at the joints in the basic skeleton diagram, the corresponding starting joints, intermediate joints, and end joints are determined. Based on the force transmission sequence during the execution of the target sports movement, candidate connection paths that satisfy the continuity of human kinematics are searched between the starting joint, intermediate joint, and terminal joint. Based on the joint coordination response intensity, action phase participation order, and attitude stability contribution of each candidate connection path, the candidate connection paths are screened to determine at least one target kinetic chain path.
[0012] The method for constructing the collaborative force edge includes: Determine at least one continuous joint segment along the target kinetic chain path whose contribution to the current action phase evaluation is lower than the preset condition; Fold the continuous joint segment into a virtual transfer segment; Establish a collaborative force-generating edge between the joints at both ends of the virtual transfer segment, and configure an equivalent transfer weight for the collaborative force-generating edge based on the response contribution of each joint in the continuous joint segment and the consistency of the state within the segment.
[0013] The calculation method for the current action phase evaluation contribution includes: Obtain the stage response information of each joint on the target kinetic chain path in the current action phase. The stage response information includes at least the joint pose change, joint motion stability, and stage activation degree corresponding to the current action phase. Based on the degree of influence of each joint on the coordinated changes among the dominant force-generating parts, power transmission parts, and end-effectors in the current action phase, determine the phase influence contribution value of each joint. Based on the position of each joint in the target kinetic chain path, determine the degree of relay necessity of each joint for the transmission of motion state between adjacent joints, and obtain the transmission necessity contribution value corresponding to each joint. Based on the stage response information, the stage impact contribution value, and the transmission necessity contribution value, determine the current action stage evaluation contribution corresponding to each key point; Adjacent joints whose current action phase evaluation contribution is continuously lower than the preset condition are identified as the continuous joint segment.
[0014] The generation methods of the spatiotemporal representation sequence of the action to be tested include: The continuous motion capture data is processed by time alignment, pose alignment and scale normalization to obtain joint temporal pose data under a unified reference frame. Based on the temporal pose data of the joints, the position change features, orientation change features, and relative motion features between joints at continuous time intervals are extracted. The position change features, orientation change features, and relative motion features between joints are then weighted and fused to obtain the spatiotemporal representation sequence of the action to be tested.
[0015] The action phase includes at least one of the following: preparation phase, force exertion phase, buffer phase, stabilization control phase, and recovery phase. The action phase is obtained by performing phase boundary detection and temporal segmentation on the spatiotemporal representation sequence of the action to be tested using a phase boundary recognition algorithm.
[0016] The method for determining the source of the motion deviation includes: By combining the human kinetic chain diagram, the target kinetic chain, the stage response degree of each joint, and the cooperative change relationship between the joints are determined within the corresponding action stage, and the kinetic chain response results corresponding to each action stage are obtained. Based on the kinetic chain response results and the preset spatiotemporal graph attention network, the stage correspondence is compared to determine the source of the action deviation in each stage of the action to be tested. Based on the associated path of the action deviation source in the human kinetic chain graph, the propagation direction, propagation intensity and propagation range of the action deviation between the associated joints are determined, and the corresponding action deviation propagation chain is generated.
[0017] The spatiotemporal graph attention network includes a spatiotemporal feature encoding unit, a spatial attention allocation unit, a temporal attention allocation unit, a stage alignment comparison unit, and a deviation source localization unit. The stage correspondence comparison includes: The power chain response results corresponding to each action stage are input into the spatiotemporal feature encoding unit to extract the stage feature representations corresponding to the joint response state, link transmission state and stage rhythm state. The spatial attention allocation unit determines the spatial attention weight of each joint and each link in the corresponding action stage based on the topological position, link weight information and cooperative change relationship of each joint in the human kinetic chain diagram. The time attention allocation unit determines the time attention weight of each time segment within the corresponding action stage based on the response change amplitude, rhythm transition position, and stage boundary proximity relationship within each action stage. The stage alignment and comparison unit matches the weighted stage feature representation with the pre-established standard action stage reference features to determine the difference response location, difference response amplitude, and difference response duration range within each action stage. The deviation source localization unit determines the motion deviation source based on the location, amplitude, and duration of the difference response, combined with the target kinetic chain distribution and link transmission relationship in the human kinetic chain diagram.
[0018] The methods for determining the evaluation results include: Based on the kinetic chain response results corresponding to each action stage, the completion degree, action stability, coordination consistency and rhythm matching degree of the target kinetic chain in each action stage are calculated to obtain the stage quality score corresponding to each action stage. Based on the propagation direction, propagation intensity, and propagation range in the propagation chain of the action deviation, determine the degree of impact of the action deviation on the key nodes and key links of the target power chain in each action stage, and generate the corresponding deviation deduction value. Based on the stage quality score corresponding to each action stage and the deviation deduction value, determine the sub-stage score corresponding to each action stage; The total score is obtained by weighting and merging the scores corresponding to each stage of the action and the stage weight of each stage in the target sports action.
[0019] Compared with the prior art, the beneficial effects of the present invention are: This invention constructs a human kinetic chain diagram, models the human body as a kinetic chain diagram, and calculates force chain consistency. This accurately depicts the coordinated force exertion relationships between joints, solving the problem of traditional methods neglecting the spatial relationships between joints and improving the accuracy of evaluation. By using a motion stage segmentation network to break down movements into stages such as preparation, force exertion, hitting, and recovery, and scoring within each stage, this invention achieves fine-grained analysis of the motion process, solving the problem that traditional methods cannot comprehensively reflect the structural characteristics of sports movements at different stages. Attached Figure Description
[0020] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 A flowchart illustrating a sports motion assessment method based on a spatiotemporal graph attention network, provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the spatiotemporal graph attention network provided in an embodiment of the present invention. Detailed Implementation
[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0022] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0023] like Figure 1 As shown, this application provides a sports movement evaluation method based on a spatiotemporal graph attention network. The method constructs a human kinetic chain graph and combines it with a spatiotemporal graph attention network to achieve fine-grained evaluation of sports movement quality and tracing of deviations. Specifically, the method includes the following steps: Step S100: Obtain continuous motion capture data of the test object during the execution of the target sports movement. The continuous motion capture data includes at least the spatial pose information of multiple joints at continuous moments.
[0024] Specifically, this step aims to provide a high-precision data foundation for subsequent analysis. The subjects can be athletes, fitness enthusiasts, or rehabilitation patients, and the target sports movements can be complex actions such as badminton smashes, basketball shots, and long jumps. Continuous motion capture data not only includes the three-dimensional spatial coordinates of joints but also implicitly contains dynamic changes over time. Data acquisition can be achieved in various ways, such as computer vision-based optical motion capture systems (e.g., multi-camera arrays), wearable devices based on inertial measurement units (IMUs), or depth cameras (e.g., ToF cameras). Spatial pose information can specifically include the three-dimensional coordinates of each joint, joint angles, velocities, and accelerations, among other kinematic parameters. It should be understood that to ensure data accuracy, the acquisition equipment typically needs to be calibrated during the acquisition process, and significant noise data needs to be removed.
[0025] Step S200: Based on the human kinematic connection relationship between the multiple joints and the joint synergistic force relationship corresponding to the target sports movement, construct a human kinetic chain diagram, and process the continuous motion capture data to obtain the spatiotemporal representation sequence of the movement to be tested.
[0026] Understandably, this step is one of the core steps of this application. Traditional methods often treat the human body as an isolated set of joints or a simple skeletal connection, ignoring the crucial kinetic chain characteristics in sports movements. In most sports movements, force is transmitted from the lower limbs to the trunk, and then through the upper limbs to the extremities, such as the hand or racket. This synergistic force relationship between non-adjacent joints is key to evaluating the quality of movement. This embodiment constructs a human kinetic chain diagram, modeling the human body as a graph structure containing nodes and edges. The edges include not only basic skeletal connection edges, such as the upper arm connecting to the forearm, but also synergistic force edges, such as the linkage between the hip and shoulder joints. This explicitly depicts the force transmission path during movement, thus solving the problem of existing technologies' difficulty in evaluating the quality of synergistic force. Simultaneously, continuous motion capture data is processed, including time alignment and posture normalization, to eliminate individual differences in height, limb proportions, and movement habits among different test subjects, extracting a universally applicable spatiotemporal representation sequence to provide input for subsequent standardized evaluation.
[0027] Step S300: Based on the spatiotemporal representation sequence of the action to be tested, the action stages are divided to obtain at least two action stages corresponding to the target sports action.
[0028] Specifically, a complete sports movement exhibits distinct phases over time. For example, a badminton smash can typically be divided into four phases: backswing, swing, hitting, and follow-through. Each phase demands significantly different joint force, body posture, and rhythm. Evaluating the entire movement as a whole often overlooks subtle deviations in local details. This step utilizes phase segmentation algorithms, such as those based on dynamic time warping or change point detection, to identify key turning points in the movement's temporal sequence, dividing the continuous spatiotemporal representation into several semantically defined movement phases. This fine-grained segmentation allows subsequent evaluations to differentiate based on the specific technical requirements of each phase, thereby improving the accuracy of the evaluation.
[0029] Step S400: Based on the human kinetic chain diagram, determine the sources of motion deviation in each motion stage, and determine the evaluation result of the test subject performing the target sports motion according to the associated paths of the motion deviation sources in the human kinetic chain diagram.
[0030] After obtaining the defined movement stages, this step utilizes a spatiotemporal graph attention network to perform in-depth analysis of the movement quality at each stage. Unlike traditional neural networks, this embodiment combines human kinetic chain diagrams for inference. The network not only focuses on the absolute positional deviations of joints but also on whether the transmission relationships on the kinetic chain diagram are abnormal. For example, if a low striking point is detected, the network will trace back using the kinetic chain diagram to determine whether the deviation originates from insufficient knee extension (e.g., insufficient lower limb power), delayed hip rotation (e.g., obstructed core power transmission), or incorrect wrist timing. In this way, not only can a quantitative score be provided, but a clear movement deviation propagation chain can also be generated, intuitively showing how the deviation is gradually transmitted from the source joint to the end and affects the final movement effect. This solves the pain point of only providing a score without offering specific corrective suggestions, achieving a leap from "what" to "why" in assessment.
[0031] Through the above steps, the method provided in this embodiment constructs a complete closed loop from data acquisition, structured modeling, stage segmentation to deviation tracing, effectively solving the problem that existing technologies are unable to objectively and multidimensionally evaluate the quality of sports movements.
[0032] In one example, this application embodiment further describes in detail the specific process of acquiring continuous motion capture data in step S100.
[0033] In one optional implementation, the spatial pose information acquisition method includes: acquiring the process of the test object performing a target sports movement through at least one motion acquisition device to obtain corresponding motion observation data; extracting the joint point observation results of the test object at each sampling time based on the motion observation data; after completing the device calibration, converting the aligned joint point observation results to a unified spatial reference coordinate system to obtain the spatial position data of multiple joint points at consecutive times; and performing human skeleton pose recovery based on the spatial position data to obtain the spatial pose information of multiple joint points at consecutive times.
[0034] It should be noted that in practical applications, motion capture devices can be multiple high-speed cameras distributed around the site, inertial measurement units worn on the subject, or RGB-D cameras capable of directly acquiring depth information. When using a multi-camera system, the intrinsic and extrinsic parameters of each camera must first be calibrated to determine the rotation matrix and translation vector between the coordinate systems of each camera. During the acquisition process, each camera independently captures a video stream. Using a pre-trained convolutional neural network model, such as OpenPose or HRNet, the pixel coordinates of key points on the human body in each frame are extracted as the joint point observation results. Since single-viewpoints may have occlusion or self-occlusion issues, it is necessary to transform the two-dimensional pixel coordinates from multiple viewpoints to a unified world coordinate system. This process typically utilizes epipolar geometry constraints or triangulation principles to fuse the two-dimensional observation data from multiple viewpoints into three-dimensional spatial coordinates, thereby obtaining the spatial position data of multiple joint points at consecutive moments.
[0035] Furthermore, in order to eliminate noise and jitter and obtain high-precision pose information, human skeleton pose recovery is achieved by combining human skeleton topological constraints, joint motion continuity constraints of adjacent time points, and joint observation confidence to jointly solve the spatial position and joint orientation of each joint.
[0036] This step is crucial for ensuring data quality. Due to occlusion, lighting variations, or sensor noise, spatial location data obtained through direct backprojection often suffers from jitter or incompleteness. This embodiment constructs a multi-constraint joint optimization model, solving the following objective function to obtain the optimal pose parameters: The objective function takes the following form: , in, This is the reprojection error term, used to measure the distance deviation between the obtained 3D coordinates of the joint points projected back to the image plane and the original observation points, ensuring that the solution results match the observed facts; To smooth the constraint, the continuity of joint motion at adjacent time points is used to penalize abrupt changes in joint position or velocity between adjacent frames, ensuring the smoothness of the motion trajectory. For example, the second-order difference of joint position between adjacent frames can be calculated as a smoothing cost. This is a confidence constraint term used to adjust and optimize weights based on the reliability of the observed data. , , These are the weighting coefficients for each item, which can be dynamically adjusted according to the actual scenario. They are mainly determined by considering the data reliability, motion speed characteristics, and occlusion noise level of the current acquisition scenario: When the multi-view imaging conditions are good, the joint point observations are clear, and the overall recognition confidence is high, the weight of the corresponding reprojection error term can be appropriately increased to enhance the fit of the solution results to the original observation data; when the measured motion is highly continuous, the sampling frame rate is high, and it is desirable to suppress trajectory jitter, the weight of the smoothing constraint term can be appropriately increased to ensure the continuity and stability of pose changes between adjacent time steps; when there are occlusions, missed detections, rapid rotations, or unstable local joint point recognition, the weight of the confidence constraint term can be appropriately increased to reduce the interference of low-reliability observations on the results.
[0037] Regarding the confidence level of joint point observations, this application further provides specific threshold settings and processing strategies. Typically, the confidence level of joint points output by image recognition algorithms ranges from 0 to 1. The specific determination of the confidence threshold can be achieved through extensive experimentation by those skilled in the art; this application exemplarily sets the confidence threshold to 0.7. When the confidence level of an observation at a certain joint point is lower than 0.7, the system determines that the observation data is unreliable and may contain severe occlusion or misidentification. In this case, the system will process the data using interpolation between adjacent frames or kinematic constraint completion. For example, if the confidence level of the right elbow joint in frame t is lower than 0.7, cubic spline interpolation is performed using the positions of the right elbow joint in frames t-1 and t+1 as the initial estimate for that point; or the position of the elbow joint is calculated using topological constraints of the human skeleton (such as the upper arm and forearm remaining constant) combined with the reliable positions of the shoulder and wrist joints. This joint solution strategy can effectively eliminate noise interference, fill in missing data, and output high-precision, robust spatial pose information, providing a reliable data foundation for the subsequent construction of accurate human kinetic chain diagrams.
[0038] In one example, this embodiment further elaborates on the specific algorithm and optimization mechanism for constructing the human kinetic chain diagram in step S200.
[0039] It should be emphasized that step S200 is the core of this application. By constructing this graph structure, the abstract human kinematic connections are transformed into a computable and inferable power transmission network.
[0040] Specifically, the methods for constructing human kinetic chain diagrams include: Step S201: Based on the human kinematic connection relationship between multiple joints, establish a basic skeleton graph containing multiple joints and joint connection edges, wherein each joint is used as a graph node, and each joint connection edge is used to represent the basic kinematic connection relationship between the corresponding joints.
[0041] The basic skeleton diagram is the underlying architecture of the human kinetic chain diagram, and its construction is based on the general human anatomical structure. For example, the node set V can be represented as {head, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, root of the spine, left hip, right hip...}, and the edge set E contains physical connections such as {(neck, left shoulder), (left shoulder, left elbow), (left elbow, left wrist)...}.
[0042] It should be understood that the basic skeleton diagram only reflects the rigid connections of the human skeleton and does not yet include the force logic under specific movements; it belongs to a static topological structure.
[0043] Step S202: Based on the action type, force exertion part, and dominant coordination relationship in the execution process of the target sports action, determine at least one target kinetic chain path from the basic skeleton diagram, and establish a coordinated force exertion edge between the associated joints on the target kinetic chain path to characterize the linkage and transmission relationship between non-adjacent joints in the execution process of the target sports action.
[0044] Specifically, determining the target kinetic chain path is a key step in constructing a kinetic chain diagram, which essentially involves searching for the main pathways for force transmission in the basic skeleton diagram.
[0045] In this embodiment, the determination of the target kinetic chain path includes: first, determining the dominant force-generating part, power transmission part, and end-effector corresponding to the target sports movement; based on the joint point distribution of the dominant force-generating part, power transmission part, and end-effector in the basic skeleton diagram, determining the corresponding starting joint, intermediate joint, and end-effector; according to the force transmission sequence during the execution of the target sports movement, searching for candidate connection paths that satisfy the continuity of human kinematics between the starting joint, intermediate joint, and end-effector; and filtering each candidate connection path based on the joint coordination response intensity, movement phase participation order, and posture stability contribution, to determine at least one target kinetic chain path.
[0046] During the search process, the joint coordination response strength is an important indicator for measuring whether a joint participates in the core force chain. This embodiment provides a specific calculation example: the joint coordination response strength R can be characterized by the normalized value of the product of joint velocity v and acceleration a, i.e. Its physical significance lies in the fact that the force-generating phase in sports movements is usually accompanied by a simultaneous surge in both velocity and acceleration, and this indicator can effectively identify the joints that are actively generating force. For example, in a badminton smash, the R-values of the hip, shoulder, elbow, and wrist joints will be significantly higher than those of the joints on the non-power-generating side. By calculating the cumulative or average R-values of all joints on each candidate path, the path with the highest response intensity is selected as the target kinetic chain path.
[0047] Specifically, in sports movement assessment scenarios, not all joints play an equal role throughout the entire movement. If all joints are treated with uniform importance at the initial stage of the movement, it is easy to include a large number of parts that only play a following, balancing, or auxiliary compensatory role in the subsequent kinetic chain search scope. This makes the path recognition results biased towards the static skeleton topology itself, and cannot reflect the key body segments that truly play a driving, transmission, and landing control role in the target sports movement. Especially in movements such as swinging, throwing, jumping, and supporting force exertion, the power output usually originates from the near-torso or lower limb support parts, completes the direction change and energy transfer through several transition joints, and finally completes the release or stabilization control at the upper limb end, foot end, or body posture control end. If the dominant force exertion part, power transmission part, and end control part are not distinguished before determining the path, the connection path obtained from the basic skeleton map, although satisfying the anatomical connection continuity, may only be a geometrically connected path, rather than a force transmission path with true movement significance. This will affect the accuracy of the movement deviation source localization and weaken the basis for identifying low-contribution continuous joint segments when constructing the subsequent collaborative force exertion side.
[0048] In this embodiment, the determination of the dominant force-generating parts, power transmission parts, and end-effector control parts is not simply based on pre-defined human body part names and is done manually. Instead, it is accomplished by combining the action category label of the target sports movement, the phased posture change trajectory in the standard action sample, and the response timing of each joint. For the input target sports movement, the candidate functional parts set corresponding to the type of movement is first extracted from the pre-defined action knowledge rules according to the action category. The action knowledge rules record the typical force-generating initiation area, typical relay transmission area, and typical action completion area of different sports movements. For example, for badminton smash movements, the candidate dominant force-generating parts may include pelvic rotation-related parts, trunk rotation-related parts, and shoulder initiation drive parts; the candidate power transmission parts may include upper arm, elbow, and forearm rotation-related parts; and the candidate end-effector control parts may include wrist, hand, and corresponding parts for equipment contact control. For vertical jump movements, the candidate dominant force-generating parts may include hip and knee; the candidate power transmission parts may include knee-ankle linkage parts; and the candidate end-effector control parts may include ankle-foot support parts. Subsequently, the qualifying samples corresponding to the target action are statistically analyzed in stages. The starting time of pose change, duration of angle change, consistency of displacement direction, and correlation with the final action result are extracted for each candidate part from the preparatory stage to the main action output stage. This identifies which part exhibits the earliest continuous active change, which part mainly undertakes the intermediate direction conversion, and which part corresponds to the final output or posture closure control of the action result. To avoid misleading results from a single sample, the determination of candidate parts adopts a consensus method based on multiple qualifying samples. That is, a part is officially included in the corresponding set only when it exhibits the same functional role in more than a preset proportion of qualifying samples. This preset proportion can be set according to the sample size. For example, when the sample size is greater than 50 groups, 70% can be used as the functional role confirmation threshold; when the sample size is small, 60% can be used, combined with manual rule verification, to ensure sufficient stability in the determination of functional parts under different action categories.
[0049] Furthermore, to ensure that test subjects of different body types and training levels remain within a unified functional framework during subsequent path searching, this embodiment also performs skeletal segmentation mapping on the human body region during the location determination process. Skeletal segmentation mapping does not use a single joint point as the location representative; instead, it combines several joint points with common biomechanical functions into local functional segments. For example, the pelvic center point and both hip joints are mapped as the lower limb initiation drive segment, the upper arm segment between the shoulder and elbow joints is considered the upper limb transmission segment, and the wrist joint and hand control point are mapped as the end control segment. This approach is because force application in real-world movements is often not strictly limited to a single joint point, but rather manifests as approximate local rigidity or coordinated local movement. Using only a single node can easily lead to unstable functional attribution due to sampling errors or momentary jitter. By mapping functional parts to local functional segments, and combining the activation sequence, duration of participation, and correlation with the action result of each local functional segment in the qualified samples, the dominant force-generating part, power transmission part, and end-effector of the target action can be determined more stably. This provides semantically relevant prior constraints for the subsequent localization of starting, intermediate, and ending joints. The subsequent path search space converges from the entire skeleton to a finite region related to the action mechanism, which not only helps reduce irrelevant paths from entering the candidate set but also makes the final determined target kinetic chain path closer to the actual force-generating process of the target athletic action.
[0050] Furthermore, even after identifying the functional parts, the entire functional part cannot be directly used as a path node. This is because the human kinetic chain path ultimately needs to be located at the specific node level of the basic skeleton diagram. Subsequent steps, such as candidate path searching, collaborative force-generating edge folding construction, and motion deviation source localization, all require clear keypoint correspondences between the starting point, transition point, and ending point of the path. If only regional descriptions such as the hip, torso, and hand are retained, it is impossible to determine the specific direction of the path along the graph structure, nor to distinguish which nodes within the same functional part play a major role in the current motion and which nodes are merely redundant observation nodes. Therefore, after the functional parts are identified, it is necessary to further refine the spatial distribution of the functional parts on the basic skeleton diagram into a representative set of keypoints, and on this basis, screen out the core nodes that can represent the initiation, transmission relay, and end-effector control of the motion.
[0051] In this embodiment, a mapping relationship between functional parts and joint point sets is first established based on a preset skeleton model. This mapping relationship can be pre-configured according to the skeleton definition file or automatically generated according to skeleton numbering rules. For example, if the three-dimensional skeleton includes nodes such as the pelvic center point, left and right hips, midpoint of the spine, left and right shoulders, left and right elbows, left and right wrists, left and right knees, and left and right ankles, and the dominant force-generating part is determined to be the hip-pelvic region, then the corresponding candidate starting node set may include the pelvic center point, left and right hip nodes, and the lower spine node connected to the pelvis; if the power transmission part is determined to be the trunk-upper limb transmission region, then the corresponding candidate relay node set may include the midpoint of the spine, shoulder joints, elbow joints, etc.; if the end-effector control part is determined to be the wrist-hand control region, then the corresponding candidate end-effector node set may include wrist nodes, hand reference points, and control nodes mapped to the device contact point. Subsequently, representativeness indicators are calculated for each candidate node set. These indicators consider at least three aspects: first, the average response precedence of a node relative to other nodes in its corresponding part within the qualifying action samples, used to determine whether the node enters the dominant action state earlier; second, the connectivity and continuity of a node's response to changes in adjacent parts, used to determine whether the node is more suitable as an interface for transmission between parts; and third, the positioning stability of a node in a continuous action sequence, used to avoid incorrectly selecting nodes that are easily occluded, prone to jitter, or have unstable estimations as critical path nodes. Taking the selection of the starting node as an example, the node that exhibits the earliest continuous pose change in the qualifying samples, and whose change can be propagated upwards to subsequent candidate relay nodes, can be prioritized as the starting joint. If multiple nodes simultaneously meet these criteria, the node with the smallest trajectory fluctuation and the highest stage activation consistency in the initial segment of the action is further selected as the final starting node. The determination of relay joints emphasizes their connectivity and transmission continuity, typically requiring stable motion associations with both preceding and following functional parts, and serving as a bridging role for direction conversion or force transmission during the action. The determination of the end joint points focuses more on their closed-loop control effect on the completion of the action. For example, in actions such as swinging, throwing, and kicking, the end joint points need to be highly correlated with the direction control, speed control, or contact accuracy at the moment of action output.
[0052] Furthermore, the force transmission in human movement is not an arbitrary, leapfrog process. A truly biomechanically significant kinetic chain must simultaneously satisfy the continuity of anatomical connections, the sequential response order within the movement phases, and the rationality of local transmission directions. Ignoring the continuity of human kinematic connections and establishing paths solely based on the statistical correlation between the starting and ending points may result in pseudo-connections traversing non-anatomically adjacent areas. Ignoring the force transmission order can easily lead to mistaking lag correlations resulting from compensation or balance after the movement's completion as the main transmission path. Ignoring the directional adjustment role of relay nodes during transmission may result in redundant paths that, while connected, do not actually perform the function of transmitting power. Therefore, in the path search phase, reachability constraints, temporal sequence constraints, and local transmission rationality constraints on the graph must be incorporated into the candidate path generation process.
[0053] In this embodiment, the search for candidate connection paths is based on the basic skeleton diagram, starting with the set of initial joints and ending with the set of final joints. The path must pass through at least one relay joint or its corresponding candidate region to ensure that the path has actual intermediate transmission segments. During the search, directional constraint annotations are first applied to the basic skeleton diagram. These directional constraint annotations do not change the physical connections of the skeleton edges, but rather assign a permissible transmission direction and a permissible transmission stage to each skeleton edge that may participate in the action, based on the phased response sequence of the target action in standard samples. For example, in an action released from the near-torso position to the upper limb, the permissible transmission direction is mainly from the pelvis or torso towards the shoulder, elbow, and wrist; while in the buffering or recovery phase, the permissible transmission direction may be opposite or exhibit a localized return transmission. Subsequently, during the candidate path search, only path segments whose edge direction matches the typical transmission direction of the current target stage are retained, thereby excluding connection paths that, although geometrically reachable, are inconsistent with the force generation mechanism. For each candidate path, it is also necessary to check whether the response timing of each adjacent node in the path satisfies a monotonically advancing or quasi-monotonically advancing relationship in the standard samples. That is, the start time of the main response of the leading node in the path should be earlier than or no later than the start time of the main response of the subsequent nodes. Small-scale synchronous overlap is allowed, but large-scale reverse ordering is not permitted. The reason for retaining small-scale synchronous overlap is that in real-world actions, some adjacent body segments may enter the main response state almost simultaneously. If a completely strict sequence is required, reasonable coordinated actions may be mistakenly excluded. For example, the allowed time offset window for quasi-monotonically advancing can be set to no more than two sampling frames or no more than 3% of the total action duration between the start times of the main responses of adjacent nodes. In high frame rate sampling scenarios, it can also be directly set to a range of 20 to 40 milliseconds based on the time length. After completing the above screening, the remaining paths constitute the candidate connection path set.
[0054] Furthermore, the candidate connection paths obtained through the aforementioned search are still only paths that may play a role in power transmission; this does not mean that all of these paths can directly participate in the subsequent construction of the human kinetic chain diagram as target kinetic chain paths. The same target athletic movement often corresponds to multiple candidate paths that meet the continuity of connection in the basic skeleton diagram. Some of these paths may primarily describe the main motion chain, while others reflect the accompanying motion chains formed by the body to maintain balance, compensate for errors, or complete auxiliary swings. If all candidate paths are included in the target kinetic chain path set without careful screening, it will not only over-inflate the human kinetic chain diagram, increasing the data processing burden for subsequent spatiotemporal representation and deviation propagation analysis, but it will also dilute the foundation for constructing synergistic force-generating edges with noisy paths, making it difficult for the main path that truly carries the quality information of the movement to stand out from the numerous paths. Therefore, at the candidate path level, a comprehensive screening mechanism needs to be established that can simultaneously reflect whether the path actually exists in the movement, whether the path's participation spans key stages, and whether the path has a constraining effect on the stable completion of the movement.
[0055] In this embodiment, for each candidate connection path, the joint coordination response strength is first calculated. This joint coordination response strength is not simply a comparison of whether the displacements of nodes on the path are synchronized, but rather a comprehensive description of the response consistency between adjacent nodes and cross-node pairs on the path. Specifically, the main response interval, response change direction, and response duration of each node on the path in each action phase are extracted first. Then, it is analyzed whether there is a continuous progression relationship between adjacent nodes, whether there is a stable coupling relationship between nodes before and after crossing relay nodes, and whether the response peaks of key nodes in the path appear within a reasonable sequence. Only when a path exhibits a continuous progression response from the starting node to the ending node in most qualifying samples, and there are no obvious breaks or reverse dominance between segments on the path, is the path considered to have a high joint coordination response strength. Then, the action phase participation order of the candidate path is analyzed. This analysis focuses not on whether the path is briefly active in a certain phase, but on whether the path participates in each phase in an order consistent with the target action. For example, in typical force-generating movements, the ideal path typically begins pre-activation in the latter part of the preparation phase, maintains high activity during the force-generating phase, and gradually decays during the buffering or recovery phases. If a path only exhibits a high response after the movement is completed, it is usually more likely a compensation chain or recovery chain than a primary force chain. Therefore, a phase participation record is established for each candidate path, statistically analyzing its activity level and order of occurrence in the preparation, power-building, force-generating, buffering, control, and recovery phases. Only when its participation order is consistent with or substantially consistent with the order of the main paths in the standard sample is a higher phase participation evaluation retained. Subsequently, the attitude stability contribution is determined. The attitude stability contribution is mainly used to screen out paths that, although related to force generation, would cause overall attitude instability or are essentially only used for local oscillation. Specifically, it can be observed whether the supporting end, trunk center, and non-moving side balance nodes of the tested object remain within a reasonable stable range when key nodes on the candidate path enter a high-response state. If a path exhibits high response accompanied by significant trunk drift, abnormal swaying of the supporting end, or overall center of gravity shift, then this path is generally unsuitable as a primary propulsion chain path and is more likely to be a case of erroneous force application or compensatory linkage. After this processing, each candidate path will correspond to a set of comprehensive descriptive results reflecting coordination, stages, and stability.
[0056] It is understood that the extraction of the main response interval, response change direction and response duration of each node on the path in each action stage in this application does not rely on a specific unique algorithm. The implementation process can refer to the existing technology of temporal action signal analysis, motion stage segmentation and joint response detection methods. Specifically, based on the joint pose sequence, angle sequence, velocity sequence, angular velocity sequence, or a combination thereof, the node response sequence can be used to perform a temporal scan of the response amplitude changes of each node in the corresponding action phase. Then, by combining sliding window statistics, local peak detection, change point detection, threshold crossing judgment, or temporal energy distribution analysis, the time interval in which each node's response changes from a low-activity state to a significantly active state and remains there can be determined as the main response interval. As for the direction of response change, it can be determined based on the node displacement vector, joint rotation direction, angle increase / decrease trend, or projection change along the target kinetic chain transmission direction to identify whether the node is in a positive force exertion, reverse retraction, or compensation swing state in the current phase. As for the duration of the response, it can be determined based on the continuous coverage length of the main response interval on the time axis, the number of continuous sampling frames, or the proportion of the duration to the total duration of the current action phase.
[0057] Step S203: Based on the force exertion sequence and stability control requirements of the target sports movement in different movement stages, configure stage activation identifiers and link weight information for the basic skeleton diagram and the collaborative force exertion edges respectively, so as to characterize the effective participation degree of each joint and each edge in different movement stages, and obtain the human kinetic chain diagram.
[0058] Understandably, during the preparation phase, the link weight of the lower limb joints is relatively high; while during the striking phase, the weight of the coordinated force exertion of the upper limbs and trunk increases significantly. This dynamic configuration allows the graph structure to adaptively match the movement process.
[0059] Furthermore, to improve computational efficiency and focus on key performance stages, this embodiment introduces a collaborative performance edge construction and graph structure optimization mechanism. The collaborative performance edge construction method includes: identifying at least one continuous joint segment along the target kinetic chain path whose contribution to the current action stage evaluation is lower than a preset condition; folding the continuous joint segment into a virtual transfer segment; establishing a collaborative performance edge between the joints at both ends of the virtual transfer segment; and configuring an equivalent transfer weight for the collaborative performance edge based on the response contribution of each joint within the continuous joint segment and the consistency of the segment's state.
[0060] It should be noted that the folding logic in this application is a graph structure simplification strategy. In long-distance kinetic chain transmission, there are often some intermediate joints, such as the middle segment of the spine. Although they physically participate in force transmission, their contribution to motion assessment is low, or their motion state is mainly controlled by adjacent joints, lacking active adjustment. Including all these joints in the graph attention network calculation would not only increase computational overhead but may also introduce noise. Therefore, this embodiment identifies these low-contribution segments by calculating the assessment contribution of the current motion stage.
[0061] In one optional implementation, the calculation method for the current action phase evaluation contribution includes: acquiring the phase response information of each joint on the target kinetic chain path within the current action phase, wherein the phase response information includes at least the joint pose change, joint motion stability, and the phase activation degree corresponding to the current action phase; determining the phase influence contribution value corresponding to each joint based on the degree of influence of each joint on the coordinated changes between the dominant force-generating part, the power transmission part, and the end-effector control part in the current action phase; determining the relay necessity of each joint for the motion state transmission between adjacent joints based on the position of each joint in the target kinetic chain path, thereby obtaining the transmission necessity contribution value corresponding to each joint; determining the current action phase evaluation contribution corresponding to each joint based on the phase response information, the phase influence contribution value, and the transmission necessity contribution value; and identifying adjacent joints whose current action phase evaluation contributions are continuously lower than a preset condition as continuous joint segments.
[0062] Specifically, the effective working mode of the same target kinetic chain path differs in different action stages. Some key points mainly serve the purpose of attitude building or support calibration in the preparation stage, and only enter a high-response state in the force exertion stage. Other key points, although located in the middle of the path, only passively follow in a certain stage. If their importance is judged directly based on their static topological position, it is easy to mistake relay nodes that should be compressed as effective nodes that must be retained, thus resulting in a large amount of segment-by-segment transmission redundancy when constructing subsequent collaborative force exertion edges. Based on this, in this embodiment, the time interval corresponding to the current action stage is first extracted from the complete action sequence, and then a stage response file is formed for each key point within that time interval. The stage response profile is not a single instantaneous observation, but rather a local temporal description formed by multiple consecutive sampling moments within the stage. It includes at least three aspects: First, the change in joint pose, characterizing whether the joint has undergone a sustained change in motion sufficient to affect the judgment of motion quality within the current stage. This can be obtained by comparing the spatial position offset, joint orientation offset, and local motion direction changes between the stage start and end times, and can be combined with the peak change range within the stage to determine whether the node has truly entered the active motion state. Second, the joint motion stability, used to distinguish between active control changes and noisy jitter or unbalanced swaying. This can be achieved by statistically analyzing the joint's response trajectory within the current stage. The fluctuation continuity, the consistency of changes between adjacent sampling times, and the degree of coordination with the changes of adjacent nodes are all important factors. If a node changes significantly but its trajectory shows frequent reverse jitters, local jumps, or disconnection from adjacent nodes, then this change should not be directly considered as high-quality stage participation. The third factor is the stage activation level, which is used to determine whether the joint is truly in the functional role corresponding to the current stage. Specifically, the response start time, duration, and response intensity of the node in the current stage can be aligned with the typical activation range of the corresponding node in the same stage in the standard action sample. If the node in the standard sample usually only enters a high activation state in the later stage of force exertion, while in the tested sample it has a low response throughout the current stage, then it indicates that its participation level in this stage is low.
[0063] In this embodiment, after obtaining the stage response file, it is necessary to further distinguish whether the joint point moves independently or whether its changes have a practical effect on the functional realization of the entire target kinetic chain in the current stage. Therefore, the determination of the stage impact contribution value is not directly sorted by the pose change amplitude, but rather a constraint analysis is performed around the cooperative change relationship between the dominant force-generating part, the power transmission part, and the end-effector control part. Specifically, in the current action stage, the stage main response trend of the node set corresponding to the dominant force-generating part, the stage relay response trend of the node set corresponding to the power transmission part, and the end-effector control trend of the node set corresponding to the end-effector control part are extracted respectively. Then, the position and role of each joint point on the target kinetic chain path among the above three types of trends are examined one by one. If the response change of a certain joint point can stably appear after the response of the dominant force-generating part and before the response of the end-effector, and the enhancement of the change at this joint point can synchronously improve the continuity of the response between the preceding and following parts, then this joint point can be considered to have a high degree of influence on the cross-part coordinated change within the current stage. Conversely, if a certain joint point changes significantly, but its change neither improves the continuity of transmission from the dominant force-generating part to the end-effector nor participates in the connection of the key action rhythm in the current stage, but only manifests as local compensatory swaying or accompanying adjustments, then its stage influence contribution value should be reduced. To avoid the stage influence contribution value only reflecting the strength of the influence and ignoring whether the transmission must be completed through this joint point, this embodiment further calculates the necessary transmission contribution value based on the positional relationship of the joint point in the target kinetic chain path. Specifically, for each target key node, the continuity of response between the adjacent path segments before and after the node in the current action phase is observed. If removing the node still maintains a high degree of response continuity and phase progression consistency between the preceding and following path segments, it indicates that the node plays more of a formal continuous connection role, and its necessity as a relay node is low. If, once the node is ignored, there is a misalignment in response timing, discontinuous direction switching, or a significant increase in end-point control fluctuations between the preceding and following path segments, it indicates that the node plays an irreplaceable relay role, and its necessary contribution value should be maintained at a high level. The reason for treating the degree of influence and the degree of relay necessity separately is that there is a type of node in actual actions that may not have a very strong dominant response, but plays a key role in direction transition or stabilization in the path. If only the magnitude of change is considered, it is easy to misjudge such nodes as low-contribution nodes. This avoids the crude approach of simply prioritizing nodes with large-amplitude actions and eliminating nodes with small-amplitude actions, and more accurately identifies which key nodes, although having limited response magnitudes, are still necessary relay nodes that cannot be collapsed.
[0064] Furthermore, after extracting the stage response information, stage impact contribution value, and necessary contribution value for transmission, this embodiment does not simply superimpose the three quantities. Instead, it first maps all three to the same contribution evaluation space and then performs continuous segment identification to ensure that the final selected continuous joint segment truly corresponds to a low-contribution relay segment suitable for folding and compression within the current stage. Specifically, the following processing logic can be adopted: First, the three types of evaluation results for each joint within the current stage are subjected to interval normalization and stability verification to make the results comparable between different types of actions, different sampling accuracies, and different body types of test subjects; then, a comprehensive evaluation contribution label is generated for each joint. This comprehensive evaluation contribution label reflects both whether the node actively participates in the target kinetic chain work within the current stage and whether the node is irreplaceable in the transmission of previous and subsequent states; finally, continuous scanning is performed along the target kinetic chain path according to the node sequence, and adjacent nodes whose comprehensive evaluation contribution is consistently lower than the preset condition are merged into continuous joint segments. The preset condition mentioned here is usually not a fixed constant, but is adaptively determined based on the node contribution distribution of the corresponding action stage in the standard action sample. For example, the quantile distribution of the comprehensive contribution of each path node in the current stage of the qualified samples can be statistically analyzed first. Then, segments with a certain proportion of contributions lower than the median of the sample contributions and at least two consecutive nodes can be identified as foldable candidate segments. The specific proportion can be determined by those skilled in the art through extensive experiments. In a set of exemplary embodiments, the low contribution judgment threshold can be set between 60% and 75% of the average comprehensive contribution of the standard sample nodes in the current stage, and the number of consecutive nodes can be at least two. When the total number of nodes in the path is long, the number of consecutive nodes can also be at least three to prevent a single occasional low response node from being mistakenly identified as a foldable segment. If an adjacent node has a slightly lower comprehensive contribution, but is immediately adjacent to an end control node, or its location corresponds to one of the few directional transition points in the path, it needs to be additionally retained and not directly included in the continuous joint segment.
[0065] It should be noted that the specific settings for the preset conditions in this embodiment are as follows: when the evaluation contribution of a certain joint point is less than 30% of the average contribution of all relevant nodes on the entire target kinetic chain path, the joint point is determined to be a low-contribution node. If there are multiple adjacent low-contribution nodes, they are merged into a continuous joint point segment. For example, when analyzing the lower limb-trunk kinetic chain, if the contribution of the three joint points in the middle of the spine is less than 30% of the average contribution of joint points, these three points are folded into a virtual transmission segment, and a collaborative force-generating edge is directly established between the hip and shoulder joints at both ends of the segment. The equivalent transmission weight of this edge can be calculated based on the average response contribution of each joint point within the folded segment. This processing method effectively reduces the complexity of the graph structure while ensuring that the key kinetic transmission path is not interrupted, allowing the subsequent spatiotemporal graph attention network to concentrate computational resources on key force transmission links such as "right hip-right shoulder", thereby improving the accuracy and efficiency of the evaluation.
[0066] In yet another example, this application embodiment further elaborates on the specific implementation process of generating the spatiotemporal representation sequence of the action to be tested in step S200 and dividing the action into stages in step S300.
[0067] Specifically, the generation method of the spatiotemporal representation sequence of the action to be tested includes: performing time alignment, pose alignment and scale normalization on continuous motion capture data to obtain joint temporal pose data under a unified reference frame; extracting position change features, orientation change features and relative motion features between joints at continuous time based on the joint temporal pose data; and performing weighted fusion of position change features, orientation change features and relative motion features between joints to obtain the spatiotemporal representation sequence of the action to be tested.
[0068] At the data processing level, due to differences in height and limb proportions among different test subjects, and the varying starting positions and orientations of each action, directly using the raw coordinate data can lead to difficulties in model convergence or inconsistent evaluation standards. Therefore, this embodiment first performs normalization processing. Specifically, time alignment uses a linear interpolation method to unify data from different sampling frequencies to a preset standard frequency. This standard frequency can be set by considering the speed characteristics of the target sports movement, the frequency of joint response changes, and the time resolution required for movement evaluation. For example, for relatively slow-paced movements such as walking, squatting, and balance control, the standard frequency can be set to 30 to 60 frames per second; for movements with rapid force generation and instantaneous rhythm changes, such as swinging, throwing, kicking, and jumping, the standard frequency can be set to 100 to 200 frames per second to ensure effective capture of key response peaks, stage boundary positions, and power transmission timing. Posture alignment uses the pelvic position of the test subject in the first frame as the origin and the spine orientation as the positive Y-axis to construct a local coordinate system, transforming all joint coordinates to this local coordinate system. Scale normalization uses the trunk length of the test subject (such as the distance from the center of the shoulder to the center of the hip) as a reference, scaling all joint coordinates to make the normalized trunk length 1, thereby eliminating the influence of individual height differences.
[0069] At the feature extraction level, to comprehensively characterize the dynamic properties of movements, this embodiment extracts three key features. First is the position change feature, mathematically expressed as the position difference vector of joints at adjacent moments. This feature represents the instantaneous velocity of the joint. Furthermore, second-order differences can be calculated to represent acceleration information. Second is the orientation change feature. Since human joints not only translate but also rotate, this embodiment uses quaternions to represent the orientation of the joints and calculates angular velocity features through quaternion differentiation. Compared to Euler angles, quaternion differentiation avoids gimbal lock-up problems and more stably describes joint rotational motion. Finally, there are the relative motion features between joints, including the relative distance and relative angle between adjacent joints. For example, for the elbow and wrist joints, the Euclidean distance between them and the angle between the upper arm and forearm are calculated. These features directly reflect the extension and flexion states of the limbs and are crucial for identifying force-generating movements. After extracting the above features, the system concatenates or weights and sums the feature vectors along the channel dimension to form a multidimensional spatiotemporal representation sequence that includes position, velocity, angular velocity, and relative geometric relationships. It should be understood that the weights of the weighted fusion can be adjusted according to the specific action type; for example, for actions emphasizing rotation, the weight of the direction change feature can be increased.
[0070] Furthermore, based on the spatiotemporal representation sequence of the action to be tested, the action stages are divided to obtain at least two action stages corresponding to the target sports action.
[0071] The action phase includes at least one of the following: preparation phase, force exertion phase, buffer phase, stabilization and control phase, and recovery phase. The action phase is obtained by performing phase boundary detection and temporal segmentation on the spatiotemporal representation sequence of the action to be tested using a phase boundary recognition algorithm.
[0072] Stage division is a prerequisite for achieving fine-grained evaluation. This embodiment employs a stage boundary identification algorithm based on a combination of energy mutation detection and dynamic time warping (DTW). The specific implementation process is as follows: The motion energy representation values at each sampling moment are extracted based on the spatiotemporal representation sequence of the action to be tested. These motion energy representation values are not limited to mechanical energy in a strictly physical sense, but rather are temporal indicators used to characterize the overall motion activity level at the current moment. They can be obtained by weighted fusion of the positional change amplitude, velocity change amplitude, angular velocity change amplitude, and relative motion change between joints at multiple joint points. In an optional implementation, the positional displacement, posture angle, and local bone segment direction changes of each joint point relative to the previous moment can be statistically analyzed, and given higher weights in conjunction with key nodes on the target kinetic chain to generate a whole-body motion energy curve for the corresponding moment. This motion energy curve can compress the original multidimensional joint temporal data into a single-dimensional response trajectory that facilitates observation of transitions, allowing the process of motion from stillness, build-up, burst, buffering to convergence to form a continuous evolution of strength and weakness on the time axis.
[0073] After obtaining the motion energy curve, it is smoothed and enhanced with local variation processing to reduce sampling noise and interference from individual joint jitter on stage boundary identification. Subsequently, the trend of the motion energy curve on the time axis is calculated, extracting the rapid energy rise region, energy peak region, rapid energy fall region, and energy stabilization region. Specifically, the segment where energy gradually rises from a low level and continues to increase typically corresponds to the transition from the preparation stage to the exertion stage; the segment where energy reaches a local peak or maintains high activity near the peak typically corresponds to the exertion stage; the segment where energy rapidly falls from a high level typically corresponds to the buffer stage; and the segment where energy tends to stabilize and overall fluctuations weaken can correspond to the stable control stage or the closing stage. Based on these characteristics, a set of candidate stage boundary points can be determined. These candidate stage boundary points are the locations in the motion energy curve where the response intensity shows a significant inflection point, used to characterize the initial boundaries between different motion stages.
[0074] In some optional implementations, to improve the accuracy of the segmentation, this embodiment also incorporates a change point detection (CPD) algorithm to verify candidate boundary points. The CPD algorithm identifies structural abrupt changes in time-series data by calculating changes in the statistical characteristics (such as mean and variance) of the data within a sliding window. When the statistical characteristics of the kinetic energy curve change significantly (such as a sudden increase in the mean exceeding a preset multiple of the standard deviation), it is confirmed as a stage boundary. Each stage has typical characteristics: the preparation stage usually corresponds to a low kinetic energy and low acceleration state, where the subject adjusts its body posture to prepare for exertion; the exertion stage corresponds to a high acceleration abrupt change, with the kinetic energy curve showing a steep rising edge, and the speed of each key joint point increasing rapidly, for example, in a badminton smash, the wrist joint speed can reach its extreme value at the moment of swinging the racket; the buffering stage corresponds to kinetic energy decay, the acceleration direction reverses, and various parts of the body begin to brake; the stabilization and control stage corresponds to low kinetic energy and small-amplitude posture adjustments, where the subject attempts to maintain balance; and the recovery stage corresponds to a static or extremely low kinetic energy state. Using the algorithm described above, the system can accurately segment continuous action sequences into action stages with clear semantics, providing a basis for setting differentiated evaluation criteria for different stages.
[0075] In yet another example, this application embodiment further elaborates on the specific network architecture and evaluation logic for determining the source of motion deviation in step S400 by combining the human kinetic chain diagram.
[0076] Specifically, the method for determining the source of motion deviation includes: combining the human kinetic chain diagram to determine the target kinetic chain within the corresponding motion stage, the stage response degree of each joint, and the cooperative change relationship between joints, to obtain the kinetic chain response result corresponding to each motion stage; comparing the kinetic chain response result with the preset spatiotemporal graph attention network to determine the source of motion deviation in each motion stage of the motion to be tested; and determining the propagation direction, propagation intensity, and propagation range of the motion deviation between the associated joints based on the association path of the motion deviation source in the human kinetic chain diagram, thereby generating the corresponding motion deviation propagation chain.
[0077] Specifically, when locating the source of motion deviation, the joint point with the largest deviation at a certain moment cannot be directly identified as the source of deviation. This is because in real sports movements, the location where a large deviation ultimately occurs is often just the outward manifestation of the deviation after it has spread, and not necessarily the initial location where the anomaly occurred. Especially in movements with obvious power transmission characteristics, there is a continuous stage-based coupling relationship between the dominant force-generating part, the power transmission part, and the end-effector. A slight anomaly at an upstream node, after being converted in direction, amplified in amplitude, or lost in stability by a relay node, may manifest as a more obvious posture deviation, speed mismatch, or control imbalance at the downstream end-effector. Based on this, in this embodiment, instead of directly comparing the overall difference between the movement under test and the standard movement, we first structure and unfold the actual power transmission state within the current movement stage based on the already determined human kinetic chain diagram. In specific processing, based on the stage division results, we first extract the local temporal segment of the corresponding movement stage from the spatiotemporal representation sequence of the movement under test, and then call the stage activation identifier and link weight information corresponding to that stage to perform intra-stage validity screening of nodes and edges in the human kinetic chain diagram, retaining only the target kinetic chain and its associated links that actually participate in power transmission in the current stage. Subsequently, the stage response degree is extracted node by node on the target kinetic chain. This stage response degree is not a single pose change amplitude, but rather a node stage response description formed by combining the displacement persistence, angle change continuity, response start time, response peak occurrence interval, and local stable state of the joints within the current stage. Simultaneously, the cooperative change relationships between adjacent nodes, cross-relay nodes, and end-control preceding nodes are analyzed along the connecting edges and cooperative force edges in the kinetic chain diagram. This is used to determine whether there are continuous propulsion relationships, synchronous enhancement relationships, or stable transmission relationships between nodes within the current stage that conform to standard motion mechanisms. Through the above processing, a set of kinetic chain response results containing the target kinetic chain structure, node stage responses, and link cooperative relationships can be generated for each action stage. This result is no longer the discrete observations in the original skeleton sequence, but rather a functional response expression within the stage after being reorganized according to the action stage and power transmission semantics. The reason for forming this stage-specific kinetic chain response result first is that the subsequent deviation source identification does not focus on "which joint is moving incorrectly", but rather on "which node or link first deviates from the kinetic chain response pattern that should exist in this stage". Only by putting the response analysis back into the target kinetic chain and its cooperative relationship can the initial anomaly and the diffusion anomaly be distinguished.
[0078] In this embodiment, after obtaining the kinetic chain response results, the spatiotemporal graph attention network is not used as a simple action classifier, but rather as a tool for fine-grained alignment and comparison between the standard response pattern and the response pattern to be tested within a stage. Specifically, the kinetic chain response results corresponding to the current action stage are first input into the spatiotemporal feature encoding part of the spatiotemporal graph attention network. This allows the network to receive the stage response sequences of each key point at the node dimension and the cooperative change sequences of the corresponding links at the edge dimension. Furthermore, the functional role information of the nodes, the link weight information, and the stage boundary position are encoded into the spatiotemporal feature representation of the current stage. Subsequently, the spatial attention component allocates different levels of attention to different nodes and links based on their topological position, node functional roles, and link transmission priorities in the human kinetic chain diagram. This avoids treating nodes that only play a subordinate balancing or weak following role in the current stage as equivalent to nodes that truly play a leading role in force generation, key transmission, and end-effector control. The temporal attention component further identifies response mutation zones, rhythm transition zones, and stability maintenance zones within the local time sequence of the current stage. This prevents the comparison process from being influenced by local short-term jitter or sampling noise, allowing for a more focused attention on the time segments that have the most decisive impact on the quality of stage completion. After completing the dual attention weighting, the weighted stage feature representation is then aligned with the reference features of the corresponding stage in the standard action samples for positional and pattern comparison. The reference features here are not fixed templates for a single sample, but rather a set of stage reference features extracted from multiple qualified samples within the same stage. These features are formed through common interval extraction and reasonable tolerance fitting, thus accommodating normal differences in body shape, movement habits, and execution range among different test subjects. During the comparison process, the total difference value across all stages is not used directly as the basis for judgment. Instead, the order of occurrence, duration, and expansion sequence of differences within a stage are tracked. If a certain key node or critical link is the first to exhibit an abnormal response that continuously exceeds the reference tolerance range in the current stage, and this abnormality subsequently triggers secondary difference responses at downstream nodes, then the location that first deviates from the standard stage response pattern is more likely to correspond to the true source of the motion deviation. Conversely, if a node has a large deviation value, but its abnormality only appears after the abnormalities of other nodes, and its location is in a typical downstream control or compensation section, it is more likely to be a result of deviation propagation rather than the source location. By comparing stages according to the principle of "first to appear, persistent, and interpretable to spread downstream," explicit deviations within and outside a stage can be distinguished from primary deviations, thereby improving the reliability and consistency of motion deviation source localization.
[0079] Furthermore, after the source of the motion deviation is initially identified, this embodiment does not simply provide an abnormal node number, but continues to use the associated paths in the human kinetic chain diagram to recover the deviation propagation process, in order to obtain a motion deviation propagation chain that can be used for evaluation and correction suggestion generation. Specifically, starting from the node or link where the deviation source is located, the response changes of its adjacent nodes and the nodes at both ends of the cooperating force edge are checked along the target kinetic chain direction activated in the current stage. This determines whether the anomaly propagates upstream, downstream, or in a branching manner at a local directional transition point. The determination of the propagation direction mainly combines three criteria: first, the temporal relationship of the abnormal responses, i.e., whether subsequent nodes exhibit the same type or explainable secondary anomalies after the deviation source; second, the link transmission relationship, i.e., whether there are effectively activated connecting edges or cooperating force edges between the deviation source and subsequent abnormal nodes; and third, the consistency of the transmission logic within the stage, i.e., whether the direction of anomaly propagation conforms to the allowed main direction of power transmission in that motion stage. The propagation intensity can be determined by classifying the degree of maintenance, amplification, or attenuation of the abnormal amplitude between the deviation source and subsequent nodes. If the deviation continuously increases at multiple relay nodes during propagation along the path, it can be considered that amplification propagation exists on that path. If the deviation only appears briefly at adjacent nodes and attenuates rapidly, it can be considered as local absorption or limited diffusion. The propagation range is determined by comprehensively considering the duration of the abnormality, the number of covered nodes, and whether it crosses key directional transition points. Only when the abnormality meets the continuous diffusion condition at adjacent or cross-segment nodes is it included in the same propagation chain. To avoid misjudging localized sporadic jitter as propagation, this embodiment also performs a stage consistency check on the propagation chain. That is, it requires that the order of occurrence of each abnormal node on the chain is basically consistent with the response and advancement order of the target power chain in that stage, and that there is an interpretable correlation between the abnormality types. For example, insufficient upstream force can reasonably lead to downstream end-control delay, directional deviation, or decreased stability, rather than discrete abnormalities unrelated to the action mechanism. The motion deviation propagation chain obtained through the above processing can completely represent "where the deviation starts, through which nodes it spreads, at which locations it is amplified, and ultimately affecting which control terminals." In this way, the key deviation locations, descriptions of the causes of deviations, and motion correction suggestions in subsequent evaluation results can all be based on a propagation chain with structural and stage-based evidence, rather than relying solely on static judgments based on a single high-deviation node. This allows the entire motion evaluation process to reflect both the final state of motion quality and the path leading to that final state.
[0080] refer to Figure 2 , Figure 2 This is a schematic diagram of the spatiotemporal graph attention network provided in an embodiment of this application.
[0081] It is understandable that the spatiotemporal graph attention network in this application is a pre-trained network model. During actual evaluation, the pre-trained model parameters are directly called to perform stage correspondence comparison and action deviation source localization for the test action, without having to retrain for each test. The reason for this setting is that the target sports actions in sports action evaluation scenarios usually have relatively clear action categories and stage structures. By pre-training offline using qualified and unqualified samples, the network can learn in advance the human kinetic chain response laws, key node collaboration relationships, and deviation propagation patterns under different action stages. Thus, in actual applications, it can directly output stable stage feature representations, differential response locations, and deviation source localization results, improving evaluation efficiency and ensuring evaluation consistency among different test subjects.
[0082] Furthermore, training data for the network can be obtained through standard motion capture and sample annotation. Specifically, multiple subjects can be pre-organized for motion capture of the target sports movement. These subjects include at least qualified individuals who can perform the target sports movement correctly, and may also include trainees with typical motion deviations. During the capture process, continuous motion capture data is acquired using optical motion capture equipment, inertial sensing equipment, depth cameras, or combinations thereof, and the spatial pose information of multiple joints at continuous moments is recovered according to the aforementioned method of this application. Based on this, a corresponding human kinetic chain diagram is constructed for each sample, and a spatiotemporal representation sequence of the movement to be tested is generated. Subsequently, coaches, professional evaluators, or a rule system based on the statistical results of qualified samples are used to annotate the stage boundaries, stage quality, and deviation locations of each training sample. The stage boundary annotation is used to indicate the start and end intervals of the preparatory stage, the exertion stage, the buffer stage, the stabilization and control stage, and the recovery stage. The stage quality annotation is used to indicate the completion level of each stage, and the deviation location annotation is used to indicate the joint, link, or propagation path where the source of the motion deviation is located. Through the above methods, a sample set for training the spatiotemporal graph attention network can be formed.
[0083] In one optional implementation, the qualified samples in the training data can be derived from demonstrations by professional athletes and coaches with the corresponding technical level, or from high-quality action samples selected after repeated collection. The deviation samples in the training data can be derived from the actual training movements of ordinary trainees, or they can be constructed based on the qualified samples by injecting perturbations into the timing of local joints, delaying the response sequence of critical links, or applying amplitude offsets to the end-effector control nodes. The purpose of this processing is to ensure that the training samples not only cover standard movement patterns but also various common movement deviation types such as insufficient force, coordination lag, rhythm mismatch, and end-effector control offset, thereby enhancing the network's ability to identify and generalize movement deviation sources.
[0084] like Figure 2 As shown, the spatiotemporal graph attention network includes a spatiotemporal feature encoding unit, a spatial attention allocation unit, a temporal attention allocation unit, a stage alignment and comparison unit, and a deviation source localization unit.
[0085] First, the kinetic chain response results corresponding to each action stage are input into the spatiotemporal feature encoding unit to extract the stage feature representations corresponding to the joint response state, link transmission state, and stage rhythm state.
[0086] Specifically, the spatiotemporal feature encoding unit adopts an architecture combining Graph Convolutional Network (GCN) and Temporal Convolutional Network (TCN). For the spatial dimension, a 1xK kernel size is used, where K is the number of neighboring nodes. The number of neighboring nodes can be determined by the number of nodes in the human kinetic chain graph that have direct connections to the current joint. Feature information of the joint and its neighboring nodes is aggregated through graph convolution operations. For the temporal dimension, a 1xT kernel size is used, where T represents the length of the temporal kernel in the temporal dimension, for example, T=9, to capture the local dependencies of joint motion over time. Through the stacking of multiple spatiotemporal convolutional layers, the network can extract high-dimensional stage feature representations. These representations not only contain the motion state of a single joint but also implicitly contain topological relationship information within the kinetic chain graph structure.
[0087] Specifically, the spatiotemporal feature encoding unit can be implemented using a multi-layer spatiotemporal convolution stacked structure. In an exemplary embodiment, the spatiotemporal feature encoding unit includes at least an input mapping layer, a first spatiotemporal feature extraction layer, a second spatiotemporal feature extraction layer, a third spatiotemporal feature extraction layer, and an output encoding layer connected in sequence. The input mapping layer maps the original input features of the joints to a unified feature dimension, facilitating joint processing by subsequent graph convolution and temporal convolution. The first spatiotemporal feature extraction layer extracts basic spatial correlation features and short-term change features of the joints within their local neighborhood. The second spatiotemporal feature extraction layer further expands the perception range of cross-node transmission relationships and intra-stage temporal evolution relationships in the kinetic chain graph. The third spatiotemporal feature extraction layer extracts high-level semantic features related to action stage semantics, kinetic chain transmission patterns, and local anomaly diffusion trends. The output encoding layer compresses or maps the features extracted from the multiple layers into stage feature representations for use by subsequent spatial attention allocation units, temporal attention allocation units, and stage alignment comparison units.
[0088] In one optional implementation, the first, second, and third spatiotemporal feature extraction layers all include a graph convolution branch, a temporal convolution branch, a normalization processing unit, and a nonlinear activation unit. The graph convolution branch aggregates features of the current node and its neighboring nodes based on the topological connections and collaborative force edge information of the human kinetic chain graph. The temporal convolution branch performs local temporal modeling of the response sequence of the same node at consecutive sampling times. The normalization processing unit reduces distribution fluctuations between different batches of samples. The nonlinear activation unit enhances the network's ability to express complex action patterns. For example, the normalization processing unit can be implemented using batch normalization, and the nonlinear activation unit can be implemented using a linear rectified function, a linear rectified function with leakage slope, or a Gaussian error linear unit. To balance computational efficiency and training stability, a linear rectified function can be used as the activation function, effectively alleviating the gradient vanishing problem in deep networks and improving convergence efficiency during spatiotemporal feature extraction.
[0089] Furthermore, regarding the number of layers, the spatiotemporal feature encoding unit is not limited to a fixed number of layers. However, to balance the ability to extract spatiotemporal patterns of movements with network complexity, it can be set to three spatiotemporal feature extraction layers in one example; in another example, it can be set to four or five spatiotemporal feature extraction layers depending on the complexity of the target sports movement. When the kinetic chain of the target sports movement is short and the stage structure is relatively clear, three spatiotemporal feature extraction layers are usually sufficient to model the local joint coordination relationships and stage rhythm changes. When the target sports movement has a longer power transmission path, more relay nodes, or a more complex multi-stage rhythm, the number of spatiotemporal feature extraction layers can be appropriately increased to enhance the representation ability of cross-segment coordination relationships and long-term dependencies. Correspondingly, the number of output channels in each layer can be set in a progressively increasing manner. For example, the number of output channels in the input mapping layer can be set to 64, the number of output channels in the first spatiotemporal feature extraction layer can be set to 64 or 128, the number of output channels in the second spatiotemporal feature extraction layer can be set to 128, and the number of output channels in the third spatiotemporal feature extraction layer can be set to 256. The output encoding layer then maps the features into a stage representation vector of a preset dimension. This method of progressively expanding the channel dimension helps the network retain local motion details in shallow layers and aggregate more abstract stage semantic information in deeper layers.
[0090] In a further implementation, residual connection structures can be set between each spatiotemporal feature extraction layer to directly pass the input features of the previous layer to the output of the next layer, thereby reducing the risk of information decay and gradient degradation during deep stacking. For the temporal convolution branch, the temporal convolution kernel size can be set to 3, 5, 7, or 9, and can be adjusted according to the duration of the target action and the sampling frequency; for example, when the target sports action is a high-speed burst action, a smaller convolution kernel can be preferentially selected to improve the ability to capture short-term abrupt responses; when the target sports action is an action with a long rhythm duration, the temporal convolution kernel size can be appropriately increased to enhance the ability to perceive longer-term dependencies. For the graph convolution branch, in addition to the basic skeleton connection edges, the collaborative force edges and their link weights can also be included in the adjacency relation matrix, so that the graph convolution not only perceives anatomical adjacent connections, but also perceives the dynamic transmission associations in the target sports action, thereby improving the sensitivity of the stage feature representation to the target kinetic chain structure.
[0091] Subsequently, the spatial attention allocation unit determines the spatial attention weight of each joint and link in the corresponding action stage based on the topological position, link weight information and cooperative change relationship of each joint in the human kinetic chain diagram.
[0092] In this embodiment, the spatial attention allocation unit utilizes a kinetic chain graph to guide attention allocation, avoiding the problem that traditional attention mechanisms may focus on irrelevant key points. Specifically, this embodiment introduces the link weights of the kinetic chain graph as a bias term when calculating the attention coefficient. If there is a collaborative force edge between two key points, their corresponding attention weights will receive an additional bonus, thereby making the network pay more attention to the key links in the force transmission path.
[0093] Meanwhile, through the time attention allocation unit, the time attention weight of each temporal segment in the corresponding action stage is determined based on the response change amplitude, rhythm transition position and stage boundary proximity relationship within each action stage.
[0094] Specifically, the temporal attention allocation unit assigns temporal attention weights by calculating the similarity between the feature vectors of each time segment and the global feature vector of the stage. For turning points where the rhythm of the action changes drastically (such as the moment of hitting the ball), the temporal attention weights are significantly increased, thereby ensuring that the features of key action frames dominate in subsequent comparisons.
[0095] Furthermore, through the stage alignment comparison unit, the weighted stage feature representation is matched with the pre-established standard action stage reference features to determine the position, amplitude, and duration of the difference response within each action stage.
[0096] The standard action stage reference features are not a single template, but rather a feature distribution range generated statistically from a large number of standard action samples. The stage alignment comparison unit uses the Dynamic Time Warping (DTW) algorithm or a Transformer-based cross-attention mechanism to calculate the distance between the action feature to be tested and the standard features. The difference response location is the keypoint or time frame where the distance exceeds a preset threshold; the difference response amplitude reflects the degree of deviation; and the difference response duration records the duration of the deviation.
[0097] It should be noted that the standard action stage reference features in this application are stage reference expressions formed by statistical analysis of the kinetic chain response results of multiple compliant action samples under the same action category and the same action stage.
[0098] Specifically, the methods for establishing reference features for the standard action phase include the following aspects: In the first aspect, for a specific target sports movement, multiple qualifying movement samples are selected from a pre-set sample library. These qualifying movement samples can be samples performed by coaches, professional athletes, or pre-assessed qualified trainees, and each sample carries a movement category identifier. To ensure the stability of the reference features, samples that meet preset standards for completion quality, have a complete movement process, and whose acquisition quality satisfies preset requirements are selected as standard samples. For example, samples with severe occlusion, continuous missing key joints, abnormal sampling frame rates, or obvious missing segments in the movement phase can be removed, and only samples with high joint recognition stability, clear phase boundaries, and high movement completion are retained in the standard sample set.
[0099] Secondly, uniform preprocessing is performed on all qualified action samples. This uniform preprocessing includes at least time alignment, posture alignment, and scale normalization. Time alignment ensures comparability between different samples at key moments such as the start of the action, the transition of force, the output of the action, and the end of the movement; posture alignment eliminates differences in the orientation of the acquisition coordinate system between different samples; and scale normalization reduces apparent amplitude differences caused by variations in height, limb length, and equipment size.
[0100] Thirdly, the action phases are segmented for the keypoint temporal pose data of each standard sample. This segmentation can be achieved using the stage boundary recognition algorithm described earlier. Specifically, based on the temporal response changes of the action to be analyzed, the complete action sequence is divided into at least two stages: a preparation stage, a force exertion stage, a buffer stage, a stabilization and control stage, and a recovery stage. For standard samples, to improve the statistical consistency of subsequent stage reference features, the initial stage segmentation results for each standard sample can be obtained using the stage boundary recognition algorithm. Then, the start and end positions of each stage are corrected based on the proximity of the stage boundaries between multiple standard samples, ensuring high consistency of the same action phase across different samples.
[0101] Fourthly, after completing the phase division, the kinetic chain response results for each standard sample are extracted within each action phase, based on the human kinetic chain diagram. These phase kinetic chain response results can include: the pose change, angle change, velocity change, angular velocity change, response start time, response peak time, response duration, local trajectory stability, and the coordinated change amplitude of each link within the current action phase, the temporal connection relationship between preceding and following nodes, the phase rhythm transition position, and the continuity of link transmission. For node-level features, they are mainly used to characterize the participation mode and intensity of a single joint within the current phase; for link-level features, they are mainly used to characterize whether the action achieves smooth coordinated transmission along the target kinetic chain within the current phase; and for phase rhythm features, they are mainly used to characterize whether the temporal progression of the action is consistent with the target action.
[0102] Fifthly, after obtaining the stage power chain response results of multiple standard samples under the same action stage, statistical modeling is performed on multiple sets of response results for the same action stage to form standard action stage reference features for the corresponding action stage. In specific implementation, node-level features, link-level features, and stage rhythm features can be aggregated separately. For node-level features, the typical response start interval, typical response peak interval, typical amplitude range, and typical stability range of each joint in the stage can be statistically analyzed; for link-level features, the typical response time difference range, typical cooperative change intensity range, and typical transmission continuity interval between adjacent nodes or nodes at both ends of the cooperative force-generating edge can be statistically analyzed; for stage rhythm features, the rhythm turning point, the length of the main active period, and the response change patterns near the front and back boundaries of the stage can be statistically analyzed.
[0103] It should be noted that the reference features for the standard action stage can be represented as either a reference feature center or a reference feature distribution range. Using a reference feature center represents the central tendency of multiple standard samples within that action stage; using a reference feature distribution range represents the allowable fluctuation range of multiple standard samples within that action stage. In specific implementations, the choice can be made based on the sample size and action complexity. When the action category is fixed, the number of samples is large, and the performance of the qualifying actions is relatively concentrated, a combination of reference feature centers and discrete ranges can be used. When the action itself allows for a certain degree of style variation, a reference feature distribution range can be used to improve the tolerance for reasonable individual differences.
[0104] In a further specific embodiment, the preset spatiotemporal graph attention network is used to perform a stage correspondence comparison between the stage kinetic chain response results of the action under test and the stage reference features of the standard action.
[0105] In one example, for the action to be tested, after the action stages are divided, the stage dynamic chain response results corresponding to each action stage are extracted and input into the preset spatiotemporal graph attention network to obtain the weighted stage feature representation of the action to be tested under each action stage. Then, this weighted stage feature representation is matched with the standard action stage reference features of the corresponding action stage. This stage matching does not simply compare whether a single feature value is equal, but comprehensively compares the deviations between node-level response features, link-level coordination features, and stage rhythm features to determine the location, amplitude, and duration of the differential response within the current action stage.
[0106] Specifically, the difference response location is used to characterize the location of the object that deviates abnormally during the current action phase. This object location can be a key point, a link, or a local response area corresponding to a time segment. The difference response amplitude is used to characterize the degree of deviation of the abnormal deviation relative to the reference features of the standard action phase. The difference response duration interval is used to characterize the length of time that the abnormal deviation persists during the current action phase.
[0107] For example, when the response start time of a certain key point during the exertion phase is significantly later than the typical start interval of the standard sample, and this lag continuously covers the main active period of the current phase, this key point can be marked as one of the differential response locations. The deviation of the start time and the deviation of the peak amplitude can be used as part of the differential response amplitude, and the continuous time period covered by the lag can be recorded as the differential response duration interval. As another example, when the response amplitudes of the preceding and following nodes in a power chain are both within a reasonable range, but the typical temporal connection between them is disrupted, leading to a decrease in transmission continuity, this link location can also be marked as a differential response location, and the degree of decrease in coordinated change can be used as the corresponding differential response amplitude.
[0108] Furthermore, the deviation source localization unit determines the motion deviation source based on the location, amplitude, and duration of the difference response, combined with the target kinetic chain distribution and link transmission relationship in the human kinetic chain diagram.
[0109] Understandably, traditional evaluation methods often only point out where the error is, but cannot explain why it is wrong. This embodiment utilizes the topological structure of a kinetic chain graph for reverse tracing. The specific process is as follows: First, the location of the difference response is mapped onto the kinetic chain graph; then, a reverse search is performed along the edges of the kinetic chain graph to find the joint with the largest difference response amplitude that is upstream in the kinetic chain, and this joint is identified as the initial source of the deviation. For example, if a difference of "low striking point" is detected, the system will trace back along the "wrist-elbow-shoulder-hip" kinetic chain. If it finds that the rotation angle difference of the hip joint is significantly greater than that of the wrist joint, and the deviation of the hip joint occurs earlier than that of the wrist joint in time, then "hip joint rotation lag" is identified as the initial source of the deviation, while the deviation of the wrist joint is merely a propagated result. Based on this, the system generates a motion deviation propagation chain, visually demonstrating how the deviation propagates from the initial source to the end point.
[0110] Furthermore, in this embodiment, after obtaining the location of the difference response, the amplitude of the difference response, and the duration of the difference response, the deviation source localization unit does not directly identify the node with the largest difference as the action deviation source, but instead performs the following determination process: First, all differential response locations detected during the current action phase are mapped to the target kinetic chain and its related links in the human kinetic chain diagram, forming a candidate abnormal object set. This candidate abnormal object set may include abnormal joints, abnormal basic skeletal edges, and abnormal collaborative force edges.
[0111] Secondly, for each candidate object in the candidate anomaly object set, its anomaly occurrence start time, anomaly duration, and anomaly amplitude level are determined to construct a corresponding anomaly response profile. The anomaly occurrence start time can be determined based on the starting position of the differential response duration interval, the anomaly duration can be determined based on the continuous length of the differential response duration interval, and the anomaly amplitude level can be determined based on the degree of deviation of the differential response amplitude from the reference characteristics of the standard action phase.
[0112] Furthermore, considering the upstream and downstream link transmission relationships in the human body kinetic chain diagram, upstream priority screening is performed on each candidate anomaly object. Specifically, if a candidate anomaly object is located upstream of another candidate anomaly object, and its anomaly occurrence time is earlier than or no later than that of the downstream candidate anomaly object, and its anomaly can provide a reasonable explanation for the downstream anomaly in the direction of kinetic chain transmission, then the upstream candidate anomaly object is preferentially retained as a candidate deviation source. Conversely, if a candidate anomaly object has a large deviation amplitude, but it is located at the end of the kinetic chain, and its preceding upstream node or link has an anomaly that can explain the end anomaly, then the end candidate anomaly object is more likely to be a result deviation after propagation, rather than the initial action deviation source. In other words, the deviation source determination in this application follows the joint principle of "earliest occurrence in time, upstream in structure, and logically able to explain subsequent anomalies," rather than simply relying on the principle of maximum numerical deviation at a certain moment.
[0113] After upstream priority screening, the remaining candidate anomalies are comprehensively ranked to determine the final source of action deviation. This comprehensive ranking can combine the anomaly's onset time, duration, amplitude level, position in the kinetic chain, and propagation interpretability score. For example, candidate anomalies with earlier onset times, longer durations, key relay positions in the target kinetic chain, and strong propagation interpretability for downstream anomalies can be assigned higher priority. Ultimately, the candidate anomaly with the highest priority is determined as the source of action deviation in the current action phase.
[0114] After identifying the source of the action deviation, this embodiment further performs propagation analysis along the target kinetic chain activated in the current action phase, targeting adjacent joints and associated links associated with the action deviation source to generate an action deviation propagation chain. Specifically, starting from the action deviation source, subsequent joints and links are sequentially checked along the kinetic chain transmission direction for any abnormal responses that are temporally connected to the deviation source, consistent in the direction of change, and matched in the cooperative disruption mode. If such responses exist, the corresponding objects are included in the action deviation propagation chain, and the deviation propagation intensity and range are determined based on the coverage, transmission level, and link impact of the abnormal spread.
[0115] In another example, the method for determining the evaluation results includes: calculating the completion degree, movement stability, coordination consistency, and rhythm matching degree of the target kinetic chain in each movement stage based on the kinetic chain response results corresponding to each movement stage, and obtaining the stage quality score corresponding to each movement stage; determining the degree of influence of movement deviation on key nodes and key links of the target kinetic chain in each movement stage based on the propagation direction, propagation intensity, and propagation range in the movement deviation propagation chain, and generating corresponding deviation deduction values; determining the sub-stage score corresponding to each movement stage based on the stage quality score and deviation deduction value corresponding to each movement stage; and performing weighted fusion based on the sub-stage score corresponding to each movement stage and the stage weight of each movement stage in the target sports movement to obtain the total score.
[0116] Specifically, when determining the evaluation results, instead of directly outputting a single score for the entire movement as a whole, we first quantify the completion status of the target kinetic chain at each stage of the movement, focusing on the kinetic chain response results. This approach is because the requirements for the same sporting movement differ at different stages. The preparatory stage focuses on the adequacy of posture establishment and force preparation; the force exertion stage focuses on the continuity of power output and link transmission; and the buffering or recovery stage focuses on the coordination of end-effector control and overall stability. If a uniform standard were used to score the entire movement all at once, the differences between stages would easily cancel each other out, resulting in a score that only reflects the overall quality of the movement, failing to indicate whether the problem occurred at the start of force exertion, during transmission, or in the end-effector control stage. In this example, we first extract the target kinetic chain coverage, key node activation, link transmission continuity, and stage rhythm progression for each movement stage, and then construct stage quality evaluation items based on these. The completion rate is used to characterize whether the target power chain in this stage has completed the required node activation and link connection according to the reference characteristics of the standard action stage. If the standard requires that the dominant force-generating part sequentially drives multiple relay nodes to enter a continuous response in a certain stage, but only local nodes are activated for a short time or the link is interrupted in the test action, the completion rate of this stage will be reduced accordingly. The action stability is used to characterize whether the key support nodes, trunk stability nodes and end control nodes in this stage maintain a reasonable fluctuation range during the action execution. If there is obvious attitude swing, support offset or control end jitter, the stability will be reduced. The coordination consistency is used to reflect whether the coordination change relationship between the dominant force-generating part, power transmission part and end control part in the current stage is continuous and consistent. If a certain path segment responds early, late or has a local break that is inconsistent with the standard timing, the coordination consistency will be reduced. The rhythm matching is used to characterize whether the order of the main response of the nodes, the duration and the timing of the stage transition in the current stage match the rhythm of the standard stage, so as to avoid the action being completed but the rhythm being unbalanced and still being misjudged as a high-quality action. Based on the above sub-evaluation results, a corresponding stage quality score can be generated for each action stage. This stage quality score is not just a simple average, but has comprehensively reflected the completion of multiple aspects such as the establishment, transmission, control and rhythm maintenance of the kinetic chain within the stage. Therefore, it can provide a basic scoring foundation with stage semantics for the subsequent superposition of deviation propagation effects.
[0117] In this example, after obtaining the stage quality score, it is necessary to further incorporate the motion deviation propagation chain into the scoring process to avoid the scoring merely focusing on the surface motion results while ignoring the formation method and depth of impact of the deviation. This is because even if two tested objects ultimately exhibit similar end-point control deviations, their causes may be completely different: one scenario might be insufficient upstream driving force, gradually amplified through relay links, ultimately leading to end-point control imbalance; another scenario might be that the overall power chain is basically normal, with only localized operational deviations occurring at the end-point control. Simply deducting points uniformly based on the magnitude of the end-point deviation would lead to insufficient identification of the former's structural problems and weaken the logical connection between the scoring results and subsequent corrective recommendations. Therefore, in this example, the impact of the deviation on key nodes and key links of the target power chain is further stratified based on the propagation direction, intensity, and range within the motion deviation propagation chain. In specific processing, first determine whether the deviation originates from the dominant force-generating node, key relay node, direction-transfer node, or end-control node; then, based on the propagation direction, determine whether the deviation is spreading downstream along the main transmission direction of the current stage, or forming a short-range disturbance within a local link. If the deviation has already spread to multiple key nodes along the main transmission direction, it indicates that the deviation has a higher degree of damage to the entire target power chain; next, combine the propagation intensity and propagation range to identify whether the deviation gradually attenuates, locally amplifies, or continuously spreads during the propagation process. If the same deviation is amplified on multiple key links and crosses the core nodes within the stage, the corresponding deviation deduction value should be higher than that of a short-term disturbance limited to a single local node. To make the deduction process feasible, deviation impact classification rules can be established in advance according to different action types. For example, the case where the propagation is limited to a single non-key relay node and does not affect end-control is defined as low impact level; the case where the propagation covers key relay links and causes end-control deviation is defined as medium impact level; and the case where it starts from the dominant force-generating part and spreads to multiple key nodes along the main path is defined as high impact level. Deduction intervals adapted to the stage quality score are configured for each. The resulting deviation deduction value is not a uniform penalty independent of the motion mechanism, but rather a phased correction amount formed by combining the structural location of the deviation propagation path, the diffusion process, and the final impact on the result. Subsequently, the phase quality score of each motion phase is combined with the corresponding deviation deduction value to obtain the phase score of each motion phase. The phase score can simultaneously reflect two levels of information: "how well the phase itself was completed" and "how much damage the deviations that occurred within the phase caused to the overall motion quality".
[0118] In one example, the process of constructing the deviation impact grading rule includes: classifying samples according to the differences in the movement mechanism of the target sports movement to form a set of movement types such as explosive output, swing and throw, balance control, continuous movement, or jump landing; and determining the corresponding dominant force-generating part, key relay part, end control part, and standard kinetic chain path for each movement type; secondly, based on the qualified and unqualified samples under each movement type, statistically analyzing the occurrence location, propagation direction, propagation length, propagation consequences, and degree of influence on the final movement result of the movement deviation in the human kinetic chain diagram, and extracting grading indicators that can characterize the degree of deviation impact. The grading indicators may include at least whether the deviation starting position is located at the dominant force-generating node or key relay node, whether the deviation propagates along the main kinetic chain direction, whether it crosses multiple key nodes during propagation, whether it causes end control deviation, and whether... This can lead to a mismatch in rhythm or a decrease in posture stability. Then, based on the aforementioned grading indicators, deviation cases in historical samples are categorized into at least three levels of deviation impact: low impact, medium impact, and high impact. Low impact corresponds to deviations limited to non-critical local nodes, not spreading along the main force chain, and not significantly affecting the movement result. Medium impact corresponds to deviations that have propagated to critical relay nodes or end-point control nodes and have a visible impact on the quality of local movements. High impact corresponds to deviations that originate at the dominant force-generating part or critical transmission link, spread along the main force chain, and significantly disrupt the overall quality of movement completion. Finally, based on the actual score decrease in the sample statistics for each deviation category, coach evaluation results, or standard movement comparison results, corresponding deduction intervals or deduction coefficients are assigned to each deviation category, thus forming a deviation impact grading rule applicable to the corresponding movement type.
[0119] Furthermore, after completing the phase-by-phase scoring of each action stage, it is necessary to determine the total score by combining the phase weights of the target sports action. It is not advisable to simply treat all stages with equal weight, because the contribution of each stage to the overall action quality varies across different action categories. For example, for explosive output actions, the power generation and power transmission phases typically have a higher evaluation weight, while for balance control actions, the evaluation weights of the stability control and recovery phases are often no lower than those of the output phase alone. Therefore, in this example, the phase weights are not fixed values, but are configured by combining the action category of the target sports action, the phase contribution distribution in the standard sample, and the criticality of the target kinetic chain in each stage. Specifically, the activation ratio of key nodes, the participation length of key links, the proportion of stage duration, and the impact of stage deviations on the final action result can be statistically analyzed in the standard action sample for each stage. Based on this, a phase weight template for that action category can be generated. During actual scoring, the phase weight template matching the current target sports action is called to weight and fuse the phase-by-phase scores of each action stage to obtain the total score. To avoid the distortion of the overall score caused by extreme high or low scores in individual stages, consistency correction can be performed on the stage scores before fusion. For example, if the score of a certain stage is abnormally low and the deviation propagation chain shows that the deviation has significantly spread to subsequent stages, the actual influence ratio of that stage in the overall score can be appropriately increased to make the overall score more consistent with the actual quality of the movement. Conversely, if the score of a certain stage is slightly low but the deviation has not spread, its normal weight is maintained to avoid over-penalizing the overall result. The overall score obtained after the above processing is not a rough average of the entire movement, but a comprehensive evaluation result formed based on the stage quality score, with deviation propagation chain correction as the core, and the stage weight of the movement category as a constraint. Therefore, it can not only reflect the overall completion level of the test subject in performing the target sports movement, but also maintain the same technical logic chain with the aforementioned stage scoring, key deviation identification, and movement correction suggestions, so that the subsequent output results have a consistent interpretive basis between scoring, positioning, and correction.
[0120] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for evaluating physical activity based on a spatiotemporal graph attention network, characterized in that, The method includes: Acquire continuous motion capture data of the test subject during the execution of the target sports movement, wherein the continuous motion capture data includes at least the spatial pose information of multiple joints at consecutive moments; Based on the human kinematic connection relationship between the multiple joints and the joint synergistic force relationship corresponding to the target sports movement, a human kinetic chain diagram is constructed, and the continuous motion capture data is processed to obtain the spatiotemporal representation sequence of the movement to be tested. Based on the spatiotemporal representation sequence of the action to be tested, the action stages are divided to obtain at least two action stages corresponding to the target sports action; By combining the human kinetic chain diagram, the sources of motion deviation in each stage of the movement are identified, and the evaluation results of the test subject's performance of the target sports movement are determined based on the associated paths of the motion deviation sources in the human kinetic chain diagram.
2. The sports movement evaluation method based on spatiotemporal graph attention network according to claim 1, characterized in that, The methods for acquiring the spatial pose information include: The process of the subject performing the target sports action is collected by at least one motion acquisition device to obtain corresponding motion observation data; Based on the motion observation data, the joint point observation results of the tested object at each sampling time are extracted; After completing the equipment calibration, the aligned joint point observation results are converted to a unified spatial reference coordinate system to obtain the spatial position data of multiple joint points at consecutive times. Based on the spatial location data, the human skeleton pose is restored to obtain the spatial pose information of the multiple joints at continuous time points.
3. The sports movement evaluation method based on spatiotemporal graph attention network according to claim 2, characterized in that, The human skeleton pose recovery is achieved by combining the topological constraints of the human skeleton, the continuity constraints of joint motion at adjacent time points, and the observation confidence of joint points to jointly solve the spatial position and joint orientation of each joint.
4. The sports movement evaluation method based on spatiotemporal graph attention network according to claim 1, characterized in that, The method for constructing the human kinetic chain diagram includes: Based on the human kinematic connection relationship between the multiple joints, a basic skeleton graph containing multiple joints and joint connection edges is established, wherein each joint is used as a graph node, and each joint connection edge is used to characterize the basic kinematic connection relationship between the corresponding joints. Based on the action type, force exertion part, and dominant coordination relationship in the execution process of the target sports action, at least one target kinetic chain path is determined from the basic skeleton diagram, and a coordination force exertion edge is established between the associated joints on the target kinetic chain path to characterize the linkage and transmission relationship between non-adjacent joints in the execution process of the target sports action. Based on the force exertion sequence and stability control requirements of the target sports movement in different movement stages, stage activation identifiers and link weight information are configured for the basic skeleton diagram and the cooperative force exertion edges, respectively, to characterize the effective participation degree of each joint and each edge in different movement stages, thus obtaining the human kinetic chain diagram.
5. The sports movement evaluation method based on spatiotemporal graph attention network according to claim 4, characterized in that, The methods for determining the target kinetic chain path include: Determine the dominant force-generating part, power transmission part, and end-effector control part corresponding to the target sports movement; Based on the distribution of the main force-generating parts, power transmission parts, and end control parts at the joints in the basic skeleton diagram, the corresponding starting joints, intermediate joints, and end joints are determined. Based on the force transmission sequence during the execution of the target sports movement, candidate connection paths that satisfy the continuity of human kinematics are searched between the starting joint, intermediate joint, and terminal joint. Based on the joint coordination response intensity, action phase participation order, and attitude stability contribution of each candidate connection path, the candidate connection paths are screened to determine at least one target kinetic chain path.
6. The sports movement evaluation method based on spatiotemporal graph attention network according to claim 4, characterized in that, The method for constructing the collaborative force edge includes: Determine at least one continuous joint segment along the target kinetic chain path whose contribution to the current action phase evaluation is lower than the preset condition; Fold the continuous joint segment into a virtual transfer segment; Establish a collaborative force-generating edge between the joints at both ends of the virtual transfer segment, and configure an equivalent transfer weight for the collaborative force-generating edge based on the response contribution of each joint in the continuous joint segment and the consistency of the state within the segment.
7. The sports movement evaluation method based on spatiotemporal graph attention network according to claim 6, characterized in that, The calculation method for the current action phase evaluation contribution includes: Obtain the stage response information of each joint on the target kinetic chain path in the current action phase. The stage response information includes at least the joint pose change, joint motion stability, and stage activation degree corresponding to the current action phase. Based on the degree of influence of each joint on the coordinated changes among the dominant force-generating parts, power transmission parts, and end-effectors in the current action phase, determine the phase influence contribution value of each joint. Based on the position of each joint in the target kinetic chain path, determine the degree of relay necessity of each joint for the transmission of motion state between adjacent joints, and obtain the transmission necessity contribution value corresponding to each joint. Based on the stage response information, the stage impact contribution value, and the transmission necessity contribution value, determine the current action stage evaluation contribution corresponding to each key point; Adjacent joints whose current action phase evaluation contribution is continuously lower than the preset condition are identified as the continuous joint segment.
8. The sports movement evaluation method based on spatiotemporal graph attention network according to claim 4, characterized in that, The generation methods of the spatiotemporal representation sequence of the action to be tested include: The continuous motion capture data is processed by time alignment, pose alignment and scale normalization to obtain joint temporal pose data under a unified reference frame. Based on the temporal pose data of the joints, the position change features, orientation change features, and relative motion features between joints at continuous time intervals are extracted. The position change features, orientation change features, and relative motion features between joints are then weighted and fused to obtain the spatiotemporal representation sequence of the action to be tested.
9. The sports movement evaluation method based on spatiotemporal graph attention network according to claim 8, characterized in that, The action phase includes at least one of the following: preparation phase, force exertion phase, buffer phase, stabilization control phase, and recovery phase. The action phase is obtained by performing phase boundary detection and temporal segmentation on the spatiotemporal representation sequence of the action to be tested using a phase boundary recognition algorithm.
10. The sports movement evaluation method based on spatiotemporal graph attention network according to claim 8, characterized in that, The method for determining the source of the motion deviation includes: By combining the human kinetic chain diagram, the target kinetic chain, the stage response degree of each joint, and the cooperative change relationship between the joints are determined within the corresponding action stage, and the kinetic chain response results corresponding to each action stage are obtained. Based on the kinetic chain response results and the preset spatiotemporal graph attention network, the stage correspondence is compared to determine the source of the action deviation in each stage of the action to be tested. Based on the associated path of the action deviation source in the human kinetic chain graph, the propagation direction, propagation intensity and propagation range of the action deviation between the associated joints are determined, and the corresponding action deviation propagation chain is generated.
11. The sports movement evaluation method based on spatiotemporal graph attention network according to claim 10, characterized in that, The spatiotemporal graph attention network includes a spatiotemporal feature encoding unit, a spatial attention allocation unit, a temporal attention allocation unit, a stage alignment comparison unit, and a deviation source localization unit. The stage correspondence comparison includes: The power chain response results corresponding to each action stage are input into the spatiotemporal feature encoding unit to extract the stage feature representations corresponding to the joint response state, link transmission state and stage rhythm state. The spatial attention allocation unit determines the spatial attention weight of each joint and each link in the corresponding action stage based on the topological position, link weight information and cooperative change relationship of each joint in the human kinetic chain diagram. The time attention allocation unit determines the time attention weight of each time segment within the corresponding action stage based on the response change amplitude, rhythm transition position, and stage boundary proximity relationship within each action stage. The stage alignment and comparison unit matches the weighted stage feature representation with the pre-established standard action stage reference features to determine the difference response location, difference response amplitude, and difference response duration range within each action stage. The deviation source localization unit determines the motion deviation source based on the location, amplitude, and duration of the difference response, combined with the target kinetic chain distribution and link transmission relationship in the human kinetic chain diagram.
12. The sports movement evaluation method based on spatiotemporal graph attention network according to claim 10, characterized in that, The methods for determining the evaluation results include: Based on the kinetic chain response results corresponding to each action stage, the completion degree, action stability, coordination consistency and rhythm matching degree of the target kinetic chain in each action stage are calculated to obtain the stage quality score corresponding to each action stage. Based on the propagation direction, propagation intensity, and propagation range in the propagation chain of the action deviation, determine the degree of impact of the action deviation on the key nodes and key links of the target power chain in each action stage, and generate the corresponding deviation deduction value. Based on the stage quality score corresponding to each action stage and the deviation deduction value, determine the sub-stage score corresponding to each action stage; The total score is obtained by weighting and merging the scores corresponding to each stage of the action and the stage weight of each stage in the target sports action.