A progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution

CN122552035APending Publication Date: 2026-08-11CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0006]针对现有技术存在的不足,本发明提出基于分解式时空归因的渐进式康复训练动作评估方法,以解决现有技术中存在的无法反映动作表现中的细粒度差异的技术问题

Benefits of technology

1.本发明通过面向康复任务的选择性状态空间建模机制,包括因果时间建模、双向时间建模、带有解剖先验的组感知时间建模以及时间到空间细化,以更好地捕获运动进程、身体部位协调性和关节级空间一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122552035A_ABST
    Figure CN122552035A_ABST
Patent Text Reader

Abstract

This invention provides a progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution. First, the original human skeletal sequence data is embedded and mapped to generate C-dimensional spatial embedding features. A bidirectional time-dependent model is constructed based on the state-space model, extracting anterior and posterior temporal features joint by joint. Prior information from anatomical groupings is introduced, and group-aware temporal models are used to obtain the specific features of each group. A spatial refinement scanning strategy is employed to complete the refined modeling of spatiotemporal features, and residual fusion yields the comprehensive modeling features. A decompositional spatiotemporal attribution FSTA head is added, decomposing global features into temporal and spatial attribution components. A visualization feedback mechanism is built based on the two types of attribution results, outputting temporal attribution curves, multi-class spatial attribution 3D graphs, attribution bar charts, and edge importance 3D graphs, intuitively presenting the basis for rehabilitation movement scoring. This invention solves the technical problem in existing technologies that fail to reflect fine-grained differences in movement performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rehabilitation training movement assessment technology, specifically to a progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution. Background Technology

[0002] Rehabilitation training plays a vital role in restoring motor function and improving the physical capabilities of patients with neurological or musculoskeletal disorders. Because rehabilitation typically requires repetitive and long-term training, home-based rehabilitation has become an important supplement to traditional therapist-guided treatment. However, home-based training is often conducted with limited or no professional supervision, making it difficult to provide timely feedback on the quality of movement execution. Therefore, patients may complete training with incorrect posture, insufficient range of motion, poor coordination, or unstable movement rhythms. These deviations can reduce rehabilitation effectiveness, and automated monitoring and assessment can help identify deviations from correct execution and support timely feedback.

[0003] Therefore, automated rehabilitation training movement assessment has received increasing attention, with the goal of providing objective, quantitative, and timely feedback for the evaluation of exercise quality.

[0004] Traditional assessments of skeletal-based rehabilitation exercises primarily treat the task as a classification problem, aiming to determine whether a particular exercise is performed correctly, rather than evaluating its continuous performance quality. While this classification paradigm provides a simple way to distinguish between normal and abnormal movements, it fails to reflect the fine-grained differences in movement performance.

[0005] Therefore, there is an urgent need to develop a progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution to address the aforementioned problems. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention proposes a progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution, in order to solve the technical problem that existing technologies cannot reflect fine-grained differences in movement performance.

[0007] The technical solution adopted in this invention is a progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution; including: The original human skeleton sequence data is embedded and mapped to obtain a C-dimensional distribution space. Embedded features ; Where T is the time length of the human skeleton sequence data, V is the number of joints in the human skeleton sequence data, and C is the dimension of the data features after embedding mapping modeling; Establish a bidirectional time-dependent model within the state-space model, and embed the features. Each joint trajectory is independently analyzed using a two-way time-dependent model to obtain forward features. With backward features Then, the features are fused to obtain the output features. ; Input the prior data of anatomical grouping into the state-space model Establish a group-perception time model, based on the prior data of each anatomical group. Embedding mapping modeling yields the corresponding embedding features. , and embed the corresponding features Prior data for each anatomical group were obtained using a group-perception time model. Corresponding output features Then the output of each group Aggregation yields output features 2; In the state-space model, a time-to-space refined model is established using a spatial refinement scan method. This is achieved by analyzing the feature outputs of the bidirectional time-dependent model and the group-aware time model. and 2. Perform spatial refinement scanning to obtain spatial refinement features. Then refine the spatial features. Residual fusion with temporal modeling features yields global output features. ; A spatiotemporal attribution model is established using the decompositional spatiotemporal attribution FSTA head, and global output features are analyzed. Mapped as spatial and temporal attribution components and The model then performs a weighted fusion of spatial and temporal attributions to obtain a quality prediction score for rehabilitation training exercises represented by human skeletal sequence data. ; Based on the output of the spatiotemporal attribution model and A visualization feedback mechanism based on attribution is established to output a visualization of rehabilitation training execution scores, including time attribution curves, spatial attribution 3D graphs, anatomical group perception spatial attribution 3D graphs, anatomical group perception spatial attribution bar charts, and edge importance 3D graphs.

[0008] Furthermore, the original human skeleton sequence data is as follows: The embedded mapping modeling is expressed as follows:

[0009] in, Original human skeleton sequence data The The first frame Data for each joint; For embedding features The The first frame Data for each joint; For learning parameters, ; It is a linear projection, and ; To learn the joint embedding position, The final embedded features are denoted as , The input feature dimension associated with each joint in the human skeleton sequence data.

[0010] Furthermore, the state-space model includes embedding feature sequences. Each joint As input to the selective state-space model, it is processed through a selective state-space scanning mechanism. The state-space model at time 1 The state is updated by the previous state and the current input, generating the potential state. Then the obtained state is mapped to output features. ;in,

[0011]

[0012] in, , and Represent and The discretized matrix; This represents the discretization step size, which depends on the input. The state transition matrix is ​​a learnable matrix. , For a parameter matrix that depends on the input, This represents the state dimension of the state-space model.

[0013] Furthermore, the bidirectional time-dependent model includes mapping the input sequence to the output sequence by independently applying state-space recursion to each joint trajectory along the time dimension, as expressed below:

[0014]

[0015] in These are forward and backward time representations, respectively. It is a time reversal operator applied along the time dimension; To perform a selective state-space scan for each joint along the time dimension, For C-dimensional distribution space Embedded features in For output features.

[0016] Furthermore, the group-aware temporal model includes dividing each joint. Groups This includes the trunk, upper limbs, or lower limbs, among which For the first Grouping by key points. For each group... Extract the corresponding feature subsequences :

[0017] in, To perform a selective state-space scan for each joint along the time dimension, To obtain the group-level feature representation, the outputs of each group are aggregated using normalized scatter-add to output the features. :

[0018] in, Representatives and Groups The corresponding all-one cover mask, Scatter() is the normalized aggregation.

[0019] Furthermore, the aforementioned time-to-space refinement model, at each time step ,right 1 and 2. Perform a selective state-space scan along the joint dimension to obtain the spatial scan feature sequence; the expression is as follows:

[0020]

[0021] in Selective state-space scanning applied along the joint dimension. and The spatial scanning feature sequence is used; then, the output features of the temporal processing model are fused with the spatial scanning features through residual fusion to obtain a refined representation, as shown in the following expression:

[0022] in, Representation layer normalization, This is the refined representation.

[0023] Furthermore, a rehabilitation movement model is constructed based on a progressive backbone network built from a bidirectional time-dependent model, a group-aware time model, and a time-to-space refinement model. This rehabilitation movement model comprises two sequential modules: CSTEB and SSTRB. These two modules have the same network architecture, and their network structure expression is as follows:

[0024] Among them, the input features of the rehabilitation movement model are: CSTEB module For the SSTRB module, This is the output of CSTEB; the output representation expression is: ; First, parallel complementary feature modeling is performed. The CSTEB and SSTRB modules in the backbone network execute two complementary transformations in parallel. The first branch applies graph-based relational modeling, with the expression:

[0025] in The second branch is a stacked AE-GC adaptive graph convolutional module; it applies multiple selective state-space modeling techniques, expressed as:

[0026] in, For and The parallel selective state-space modeling module, depending on the module configuration, includes a bidirectional time-dependent model, a group-aware time model, and a time-to-space refinement model, through... Emphasizing topology-aware relational structures. Capture context-adaptive sequence dynamics; Then, on as well as The output is adaptively fused using an input-conditional adaptive fusion mechanism. First, the input of the rehabilitation action model is... Perform global average pooling on the time and joint dimensions:

[0027] in, Represents the pooling result, Represents pooling operations, express The Frame number Feature data of each joint point represent A real vector space.

[0028] Lightweight gating network Generate normalized branch weights:

[0029] in, Represents a soft maximum function. It represents a 2-dimensional real vector space.

[0030] The fusion characterization is calculated as follows:

[0031] and For the above formula The first and second components in.

[0032] Following adaptive fusion, the selective state-space modeling module is further applied sequentially. The fusion is refined along the time dimension to obtain the fusion features. The expression is as follows:

[0033] Application module-level residual connections and layer normalization generate output:

[0034] in For layer normalization; Through progressive hierarchical modeling, a two-stage modeling process is formed, gradually refining the motion representation. Given embedded features... Hierarchical modeling is performed using the CSTEB module. embed features Transformed into first-layer global features SSTRB module take over As input, it is refined into a second-layer global feature. The expression for this process is as follows: .

[0035] Furthermore, the spatiotemporal attribution model first predicts a dense score map frame-by-frame and joint-by-joint. The expression is as follows:

[0036] in It is a two-layer fully connected network; Then, the space and time logits are calculated using global pooling, as shown in the following expression:

[0037] in, and and They employ similar structures, but their parameters are independent. Second layer global features The Frame data, for The The data for each joint is then normalized using a temperature-controlled softmax function to obtain the logits.

[0038] in Spatial attribution for joints, For temporal attribution on the frame, τ>0 is the temperature parameter in softmax normalization, used to control the concentration of normalized logits; Then, spatiotemporal attribution weights for skeleton joints and time steps are constructed through outer product factorization:

[0039] in Spatiotemporal attribution weights; The attribution at each spatiotemporal location is decomposed into separable temporal and spatial contributions, given the input skeleton sequence. Quality score of rehabilitation exercise execution The global predicted value is obtained through weighted aggregation, as shown in the following expression: .

[0040] Furthermore, the attribution-based visualization feedback mechanism includes generating a temporal attribution vector using the FSTA head for each input skeleton sequence. and spatial attribution vector ,Will Visualize the time attribution curve, label representative keyframes and high response intervals, and identify the representative keyframes as those with the highest time attribution, as shown in the following expression:

[0041] Highlight the high response time range and determine the performance of a specific frame in the visualization. If the attribution satisfies the following condition, it is considered a high-response frame, as expressed below:

[0042] in, Indicates the first The temporal attribution value of the frame. In the experiment, [the value was taken as...]. , and , This indicates a maximum value operation, where consecutive high-response frames are merged into time segments, and only segments with a length of at least 2 frames are retained for highlighting; Select keyframe Above, the pose is normalized using reference joints and bone lengths, and spatial attribution values ​​are encoded by marker size and color intensity; The spatial attribution value is further defined in the predefined anatomical group. In the aggregation process, when there is overlap between groups, the joint contribution is redistributed proportionally to avoid duplicate counting, as shown in the following expression:

[0043] in For including joints The number of groups, This represents the spatial attribution of the k-th group, and the aggregated group-level attribution. Highlight body groups that contribute to motion quality in the 3D visualization; Next, construct a dynamic adjacency matrix, expressed as follows:

[0044] in, It is a dynamic adjacency matrix. Represents the skeleton adjacency matrix. This is the scaled edge importance matrix; This is a learnable edge importance matrix in the human skeleton model; since skeleton connections are treated as undirected connections in visualization, the edges... Importance The learned edge importance is calculated using a symmetric method, as shown in the following expression:

[0045] in, Representing an edge The corresponding importance value after scaling, Representing an edge The corresponding scaled edge importance values ​​are then mapped to line widths and color intensities in the 3D skeleton visualization, making connections with higher values ​​visually more prominent.

[0046] Furthermore, a computer storage medium stores a computer program that can run on a processor, characterized in that, when the processor executes the computer program, it implements the progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution as described above.

[0047] As can be seen from the above technical solution, the beneficial technical effects of the present invention are as follows: 1. This invention utilizes a selective state-space modeling mechanism for rehabilitation tasks, including causal time modeling, bidirectional time modeling, group-aware time modeling with anatomical priors, and time-to-space refinement, to better capture movement processes, body part coordination, and joint-level spatial consistency.

[0048] 2. By learning coarse-to-fine spatiotemporal representations and combining adaptive edge-weighted graph convolution with a Mamba-based selective state-space modeling mechanism, topology-aware inter-joint dependencies and long-range motion dynamics are captured. The selective state-space modeling design for rehabilitation tasks further enhances the modeling of motion processes, body part coordination, and joint-level spatial consistency, improving the accuracy of joint movement assessment during rehabilitation training. Attached Figure Description

[0049] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0050] Figure 1 This is a flowchart of the method of the present invention; Figure 2 A sample image of a squat training movement performed by an expert of this invention; Figure 3 This is a time attribution curve diagram of the present invention; Figure 4 This is a three-dimensional spatial attribution diagram of the present invention; Figure 5 This is a three-dimensional diagram of anatomical grouping perception spatial attribution according to the present invention; Figure 6 This is a bar chart illustrating the spatial attribution of anatomical groupings according to the present invention. Figure 7 This is a three-dimensional diagram illustrating the edge importance of the present invention. Figure 8 This invention uses an expert sample, with a score chart of subjects whose quality score is 50.0; Figure 9 For the purpose of this invention, a group-level attribution visualization of pelvic rotation and squat with a quality score of 50.0 using expert samples; Figure 10 The score chart for subjects with a quality score of 44.3, using a non-expert sample, is shown in the present invention. Figure 11 For the purpose of this invention, a group-level attribution visualization of pelvic rotation and squat with a quality score of 44.3 is provided for the use of non-home samples. Figure 12 This is a score chart of subjects with a quality score of 26.0, using a non-expert sample, for the present invention. Figure 13 For the purpose of this invention, a group-level attribution visualization of pelvic rotation and squat with a quality score of 26.0 using non-expert samples; Figure 14 The score chart for the subject with a quality score of 48.3, who is a sample of experts using this invention; Figure 15 For the expert sample used in this invention, a group-level attribution visualization of pelvic rotation and squat with a quality score of 48.3; Figure 16 The score chart for subjects with a quality score of 47.7, using a non-expert sample, is shown in the present invention. Figure 17 For the purpose of this invention, a group-level attribution visualization of pelvic rotation and squat with a quality score of 47.7 using non-expert samples; Figure 18 This is a rating chart of subjects with a quality score of 17.0, using a non-expert sample, for the present invention. Figure 19 For the purpose of this invention, a group-level attribution visualization of pelvic rotation and squat with a quality score of 17.0 using non-expert samples; Detailed Implementation

[0051] The embodiments of the technical solution of the present invention will now be described in detail with reference to the accompanying drawings. These embodiments are merely illustrative of the technical solution of the present invention and are therefore intended to limit the scope of protection of the present invention.

[0052] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application should have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0053] Example 1: like Figure 1 As shown, this embodiment provides a progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution; including: The original human skeleton sequence data is embedded and mapped to obtain a C-dimensional distribution space. Embedded features ; Where T is the time length of the human skeleton sequence data, V is the number of joints in the human skeleton sequence data, and C is the dimension of the data features after embedding mapping modeling; Establish a bidirectional time-dependent model within the state-space model, and embed the features. Each joint trajectory is independently analyzed using a two-way time-dependent model to obtain forward features. With backward features Then, the features are fused to obtain the output features. ; Input the prior data of anatomical grouping into the state-space model Establish a group-perception time model, based on the prior data of each anatomical group. Embedding mapping modeling yields the corresponding embedding features. , and embed the corresponding features Prior data for each anatomical group were obtained using a group-perception time model. Corresponding output features Then the output of each group Aggregation yields output features 2; In the state-space model, a time-to-space refined model is established using a spatial refinement scan method. Spatial refinement features are obtained by performing spatial refinement scans on the feature outputs of the bidirectional time-dependent model and the group-aware time model. Then refine the spatial features. Residual fusion with temporal modeling yields global output features. ; A spatiotemporal attribution model is established using the decompositional spatiotemporal attribution FSTA head, and global output features are analyzed. Mapped as spatial and temporal attribution components and The model then performs a weighted fusion of spatial and temporal attributions to obtain a quality prediction score for rehabilitation training exercises represented by human skeletal sequence data. ; Based on the output of the spatiotemporal attribution model and A visualization feedback mechanism based on attribution is established to output a visualization of rehabilitation training execution scores, including time attribution curves, spatial attribution 3D graphs, anatomical group perception spatial attribution 3D graphs, anatomical group perception spatial attribution bar charts, and edge importance 3D graphs.

[0054] In this embodiment, the original human skeleton sequence data is: The embedded mapping modeling is expressed as follows:

[0055] in, Original human skeleton sequence data The The first frame Data for each joint; For embedding features The The first frame Data for each joint; For learning parameters, ; It is a linear projection, and ; To learn the joint embedding position, The final embedded features are denoted as , For human skeleton sequence data, the input feature dimensions related to each joint are such as 3D joint coordinates, joint velocities, or other kinematic features. The dimension of the data features after embedding mapping modeling should be 128 or higher.

[0056] In this embodiment, the state-space model includes the embedding feature sequence. Each joint As input to the selective state-space model, it is processed through a selective state-space scanning mechanism. The state-space model at time 1 The state is updated by the previous state and the current input, generating the potential state. Then the obtained state is mapped to output features. ;in,

[0057]

[0058] in, , and Represent and The discretized matrix; This represents the discretization step size, which depends on the input. The state transition matrix is ​​a learnable matrix. , The input parameter matrix, For the state-space model, The state is mapped to the output feature.

[0059] In this embodiment, the bidirectional time-dependent model includes mapping the input sequence to the output sequence by independently applying state-space recursion to each joint trajectory along the time dimension, as shown in the following expression:

[0060]

[0061] in These are forward and backward time representations, respectively. It is a time reversal operator applied along the time dimension; To perform a selective state-space scan for each joint along the time dimension, For C-dimensional distribution space Embedded features in For output features.

[0062] In this embodiment, the group-aware time model includes dividing each joint. Groups This includes the trunk, upper limbs, or lower limbs, among which For the first Grouping of key points. For each group Extract the corresponding feature subsequences :

[0063] in, To perform a selective state-space scan for each joint along the time dimension, To obtain the group-level feature representation, The outputs from each group are aggregated using normalized scatter-add to output features. :

[0064] in, Representatives and Groups The corresponding all-one cover mask, Scatter() is the normalized aggregation.

[0065] In this embodiment, the time-to-space refinement model, at each time step ,right 1 and 2. Perform a selective state-space scan along the joint dimension to obtain the spatial scan feature sequence; the expression is as follows:

[0066]

[0067] in Selective state-space scanning applied along the joint dimension. and The spatial scanning feature sequence is used as the basis for the refinement. Then, the output features of the temporal processing model are fused with the spatial scanning feature sequence using residual fusion to obtain the refined representation, as shown in the following expression:

[0068] in, Representation layer normalization, This is the refined representation.

[0069] In this embodiment, a rehabilitation movement model is constructed based on a progressive backbone network built from a bidirectional time-dependent model, a group-aware time model, and a time-to-space refinement model. The rehabilitation movement model includes two sequential modules: CSTEB and SSTRB. These two modules have the same network architecture, and their network structure expression is as follows:

[0070] Among them, the input features of the rehabilitation movement model are: CSTEB module For the SSTRB module, This is the output of CSTEB; the output representation expression is: ; First, parallel complementary feature modeling is performed. The CSTEB and SSTRB modules in the backbone network execute two complementary transformations in parallel. The first branch applies graph-based relational modeling. The expression is:

[0071] in For stacked AE-GC adaptive graph convolutional modules; For the output of CSTEB, the second branch applies multiple selective state-space modeling. The expression is:

[0072] in, For and The parallel selective state-space modeling module, depending on the module configuration, includes a bidirectional time-dependent model, a group-aware time model, and a time-to-space refinement model, through... Emphasizing topology-aware relational structures. Capture context-adaptive sequence dynamics; Then, on as well as Adaptive fusion employs an input-conditional adaptive fusion mechanism. First, the input of the rehabilitation movement model... Perform global average pooling on the time and joint dimensions:

[0073] in, Represents the pooling result, Represents pooling operations, express The Frame number Feature data of each joint point represent A real vector space.

[0074] Lightweight gating network Generate normalized branch weights:

[0075] in, Represents a soft maximum function. It represents a 2-dimensional real vector space.

[0076] The fusion characterization is calculated as follows:

[0077] and For the above formula The first and second components in; Following adaptive fusion, the selective state-space modeling module is further applied sequentially. The fusion is refined along the time dimension to obtain the fusion features. The expression is as follows:

[0078] Application module-level residual connections and layer normalization generate output:

[0079] in For layer normalization; Through progressive hierarchical modeling, a two-stage modeling process is formed that gradually refines the motion representation; given embedded features Hierarchical modeling is performed using the CSTEB module. embed features Transformed into first-layer global features SSTRB module take over As input, it is refined into a second-layer global feature. The expression for this process is as follows: .

[0080] In this embodiment, the spatiotemporal attribution model first predicts the dense score map frame by frame and joint by joint. The expression is as follows:

[0081] in It is a two-layer fully connected network; Then, the space and time logits are calculated using global pooling, as shown in the following expression:

[0082] in, and and They employ similar structures, but their parameters are independent. Second layer global features The Frame data, for The The data for each joint is then normalized using a temperature-controlled softmax function to obtain the logits.

[0083] in Spatial attribution for joints, For temporal attribution on the frame, τ>0 is the temperature parameter in softmax normalization, used to control the concentration of normalized logits; Then, spatiotemporal attribution weights for skeleton joints and time steps are constructed through outer product factorization:

[0084] in Spatiotemporal attribution weights; The attribution at each spatiotemporal location is decomposed into separable temporal and spatial contributions, given the input skeleton sequence. Quality score of rehabilitation exercise execution The global predicted value is obtained through weighted aggregation, as shown in the following expression: .

[0085] In this embodiment, the attribution-based visualization feedback mechanism includes generating a temporal attribution vector using the FSTA head for each input skeleton sequence. and spatial attribution vector ,Will Visualize the time attribution curve, label representative keyframes and high response intervals, and identify the representative keyframes as those with the highest time attribution, as shown in the following expression:

[0086] Highlight the high response time range and determine the performance of a specific frame in the visualization. If the attribution satisfies the following condition, it is considered a high-response frame, as expressed below:

[0087] in, Indicates the first The temporal attribution value of the frame. In the experiment, [the value was taken as...]. , and , This indicates a maximum value operation, where consecutive high-response frames are merged into time segments, and only segments with a length of at least 2 frames are retained for highlighting; Select keyframe Above, the pose is normalized using reference joints and bone lengths, and spatial attribution values ​​are encoded by marker size and color intensity; The spatial attribution value is further defined in the predefined anatomical group. In the aggregation process, when there is overlap between groups, the joint contribution is redistributed proportionally to avoid duplicate counting, as shown in the following expression:

[0088] in For including joints The number of groups, This represents the spatial attribution of the k-th group, and the aggregated group-level attribution. Highlight body groups that contribute to motion quality in the 3D visualization; Next, construct a dynamic adjacency matrix, expressed as follows:

[0089] in, It is a dynamic adjacency matrix. It is a skeleton adjacency matrix. This is the scaled edge importance matrix; This is a learnable edge importance matrix in the human skeleton model; since skeleton connections are treated as undirected connections in visualization, the edges... Importance The learned edge importance is calculated using a symmetric method, as shown in the following expression:

[0090] in, Representing an edge The corresponding importance value after scaling, Representing an edge The corresponding scaled edge importance values ​​are then mapped to line widths and color intensities in the 3D skeleton visualization, making connections with higher values ​​visually more prominent.

[0091] Example 2 In this embodiment, a computer storage medium stores a computer program that can run on a processor. The processor executes the computer program to implement the progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution as described above.

[0092] Example 3 In this embodiment, to evaluate the overall effectiveness of the method of the present invention, it is compared with state-of-the-art methods on two widely used rehabilitation training movement assessment benchmarks, UI-PRMD and KIMORE. The present invention uses MAD as the evaluation metric on the UI-PRMD dataset, and reports MAD, RMSE, and MAPE to provide a more comprehensive assessment. For all metrics, lower values ​​indicate better assessment performance.

[0093] In this embodiment, Table 1 presents the comparison results on the UI-PRMD dataset, with lower values ​​indicating better performance. This invention achieves a best average MAD of 0.007, outperforming recent strong competing methods such as SSL-Rehab (Kourbaneetal.2025b), hierarchical contrastive representation methods, and optimized spatiotemporal graph convolutional networks. Specifically, this method achieves the lowest MAD on 6 out of 10 actions, including Ex1, Ex2, Ex3, Ex5, Ex7, and Ex10, and obtains competitive results on the remaining actions. Compared to earlier graph-based and deep learning-based methods, this method significantly reduces the average prediction error. The improvement on UI-PRMD demonstrates that the proposed framework can effectively model fine-grained quality differences in skeleton sequences. Based on a progressive GCN-Mamba architecture, hybrid GCN-Mamba modules jointly capture topology-aware inter-joint dependencies and long-term temporal motion dynamics. Furthermore, the progressive selective state-space modeling scheme further enhances representation capabilities by combining causal temporal modeling, bidirectional temporal modeling, group-aware temporal modeling with anatomical priors, and temporal-to-spatial refinement.

[0094] Table 2 presents the comparison results on the KIMORE dataset using MAD, RMSE, and MAPE as metrics, with lower values ​​indicating better performance. Our method achieves the best average performance across all three metrics: average MAD of 0.137, RMSE of 0.257, and MAPE of 0.405. Compared to the strongest comparison method, Kuang et al., our method further reduces the average MAD from 0.162 to 0.137, the average RMSE from 0.284 to 0.257, and the average MAPE from 0.463 to 0.405. Furthermore, our method achieves the best MAD on Ex1, Ex3, and Ex5, the best RMSE on Ex1 and Ex4, and the best MAPE on Ex1, Ex2, Ex3, and Ex5. These consistent improvements demonstrate that the proposed method can generalize well to different rehabilitation training movements. The results on KIMORE further validate the effectiveness of the proposed model under more diverse training conditions. Compared to Transformer-based spatiotemporal modeling methods, contrastive representation learning methods, and self-supervised transfer learning methods, this method explicitly introduces rehabilitation task-oriented modeling priors in the selective state-space model design. Bidirectional temporal modeling helps utilize past and future motion contexts, group-aware temporal modeling emphasizes anatomically significant body part dynamics, and temporal-to-spatial refinement improves joint-level spatial consistency. These mechanisms are particularly suitable for rehabilitation training movement assessment because the final quality score is closely related to movement progression, limb coordination, and local postural consistency.

[0095] Specifically, our method consistently outperforms or reaches state-of-the-art levels on both datasets. Its superior performance is primarily attributed to three aspects: (1) a progressive GCN-Mamba framework for learning coarse-to-fine spatiotemporal representations; (2) a hybrid GCN-Mamba module for complementary modeling of topologically aware inter-joint dependencies and long-term temporal motion dynamics; and (3) a selective state-space modeling mechanism for robust motion representation oriented towards rehabilitation tasks. Furthermore, the FSTA head provides discriminative spatiotemporal aggregation for final quality prediction while supporting interpretable feedback. These results demonstrate the effectiveness of the proposed method in regression-based rehabilitation training motion assessment.

[0096] Table 1 Performance comparison of the UI-PRMD dataset using MAD as the metric

[0097] Table 2 Performance comparison of MAD, RMSE, and MAPE on the KIMORE dataset.

[0098] Table 3 Ablation studies of different component combinations of the overall framework on the UI-PRMD and KIMORE datasets.

[0099] In this embodiment, the effectiveness of each key component in the proposed method is further investigated. The present invention conducted a series of ablation experiments on the UI-PRMD and KIMORE datasets. All ablation results are reported as the average MAD, RMSE and MAPE of all action categories in each dataset. The lower the value, the better the performance.

[0100] In this embodiment, Table 3 evaluates the effects of different component combinations in the overall framework, including CSTEB, SSTRB, and the FSTA head. Comparing F1 and F2, it can be observed that using SSTRB alone is superior to using CSTEB alone on both datasets. This indicates that the structural enhancement module, which introduces richer temporal context and anatomical priors, has a stronger expressive power for high-level rehabilitation motor representation. The complete framework in F4 consistently achieves the best results, reducing MAD, RMSE, and MAPE to 0.007, 0.020, and 1.003 on UI-PRMD, and to 0.137, 0.257, and 0.405 on KIMORE, respectively.

[0101] Compared to F3 with the FSTA head removed, F4 achieves significant improvements on both datasets. Particularly on KIMORE, MAD decreases from 0.226 to 0.137, RMSE from 0.369 to 0.257, and MAPE from 0.665 to 0.405. The results demonstrate that the decompositional spatiotemporal attribution head not only improves interpretability but also contributes to more accurate regression results by adaptively emphasizing informative frames and joints. Overall comparisons validate the effectiveness of the progressive framework design. Specifically, CSTEB and SSTRB jointly construct a coarse-to-fine representation hierarchy, while FSTA provides discriminative spatiotemporal aggregation for the final quality score prediction.

[0102] Specifically, Table 4 examines the contributions of different components in the GCN-Mamba module, including the GCN branch, the parallel Mamba branch, and the serial Mamba module. The full configuration B4 achieves best performance on both datasets, confirming the necessity of integrating these three components.

[0103] As shown in B1, when using only the GCN branch and serial Mamba, the model performs well on UI-PRMD but degrades on KIMORE. This indicates that in more complex and diverse rehabilitation scenarios, relying solely on graph-based spatial modeling is insufficient, as larger subject populations, heterogeneous health conditions, and stronger inter-subject movement variability make quality assessment more challenging. In contrast, B2 removes the GCN branch, relying only on parallel and serial Mamba modules. Although this setting improves performance on KIMORE compared to B1, it performs worse on UI-PRMD, demonstrating that explicit skeletal topology remains important for modeling inter-joint dependencies. B3 further shows that removing the serial Mamba module weakens the final representation, particularly in RMSE and MAPE metrics.

[0104] The superior performance of B4 demonstrates that the proposed hybrid GCN-Mamba module benefits from the complementarity between topology-aware graph convolution and Mamba-based state-space-time modeling. Specifically, the adaptive edge-weighted graph convolution branch captures discriminative inter-joint relationships, the parallel Mamba branch models long-range motion dynamics, and the serial Mamba module further refines the fused representation by improving temporal consistency.

[0105] Specifically, Table 5 evaluates different selective state-space modeling strategies in SSTRB. Starting from the causal temporal modeling baseline M1, adding bidirectional temporal modeling to M2 slightly reduces RMSE and MAPE on both datasets. This indicates that introducing future context helps generate more robust temporal representations. Further introducing group-aware temporal modeling with anatomical priors, M3 reduces RMSE on both UI-PRMD and KIMORE. This suggests that time modeling based on anatomical body parts is beneficial for capturing motion patterns at the body part level. However, the performance improvement is relatively limited, indicating that relying solely on anatomical grouping may not be sufficient to adequately enhance inter-joint spatial coordination.

[0106] Table 4 Ablation studies of different component combinations of the GCN-Mamba module on the UI-PRMD and KIMORE datasets.

[0107] Table 5 Ablation data of different modeling strategies in SSTRB on the UI-PRMD and KIMORE datasets.

[0108] Table 6 Ablation data of loss functions on the UI-PRMD and KIMORE datasets.

[0109] The full SSTRB configuration M4, which further incorporates temporal-to-spatial refinement, achieves the best results with a significant advantage. On UI-PRMD, M4 reduces MAD, RMSE, and MAPE from 0.012, 0.030, and 1.709 in M3 to 0.007, 0.020, and 1.003, respectively. On KIMORE, the corresponding metrics decrease from 0.287, 0.520, and 0.843 to 0.137, 0.257, and 0.405. These results demonstrate that temporal-to-spatial refinement plays a crucial role in enhancing joint-level spatial consistency after temporal modeling. Overall, the ablation results validate the effectiveness of the proposed progressive selective state-space modeling scheme, where causal temporal modeling, bidirectional temporal modeling, group-aware temporal modeling, and temporal-to-spatial refinement collectively improve the robustness and discriminative power of skeleton-based motion representations.

[0110] Table 6 examines the impact of different loss components, including regression loss. Spatial attribution regularization and time attribution regularization Use only The baseline L1 performance is weaker than the configuration using attribution regularization, indicating that simple regression supervision is insufficient to fully guide the attribution head.

[0111] After adding spatial attribution regularization to L2, the performance on both datasets is better than L1. This shows that encouraging more focused spatial attribution helps the model identify important joints in motion quality assessment. Similarly, L3 introduces temporal attribution regularization and achieves better results than L1, indicating that penalizing abrupt changes in temporal attribution helps to obtain more stable temporal importance estimates.

[0112] The complete loss function L4 consistently achieves the best performance across all metrics and datasets. Compared to L1, L4 reduces MAD, RMSE, and MAPE from 0.011, 0.030, and 1.709 to 0.007, 0.020, and 1.003 on UI-PRMD, and from 0.206, 0.455, and 0.581 to 0.137, 0.257, and 0.405 on KIMORE. These results demonstrate the complementary role of spatial and temporal attribution regularization terms. Specifically, Promote spatial attribution to focus on joints that contribute more, while Encouraging smoother temporal attribution across consecutive frames. Combining these two approaches can lead to more accurate and interpretable regressions.

[0113] In this embodiment, to evaluate the computational efficiency of the proposed model, it is compared with a Transformer-based variant, in which the Mamba-based selective state space module is replaced with a Transformer module, while the rest of the framework remains unchanged. All experiments were performed on a single NVIDIA A800 GPU.

[0114] As shown in Table 7, the two models have similar parameter counts. M, h, and s represent millions of parameters, hours, and seconds, respectively. The slight differences on different datasets are due to their different skeleton layouts, as UI-PRMD and KIMORE contain 39 and 25 joints, respectively. Compared to the Transformer-based variants, the proposed model consistently achieves lower computational cost and better regression performance. On UI-PRMD, the proposed model saves 2.638 hours of training time and 0.339 seconds of testing time, while reducing MAD, RMSE, and MAPE by 0.002, 0.008, and 0.379, respectively. On KIMORE, the proposed model saves 1.192 hours of training time and 0.213 seconds of testing time, while reducing MAD, RMSE, and MAPE by 0.121, 0.155, and 0.244, respectively.

[0115] Table 7 Comparison of computational cost and test performance of the proposed model and its Transformer-based variants on UI-PRMD and KIMORE.

[0116] The above results demonstrate that the Mamba-based selective state-space model of this invention provides an efficient alternative to Transformer-based sequence modeling, suitable for skeleton-based rehabilitation assessment. By using state-space recursion to model long-range motion dynamics, rather than explicitly calculating dense dependencies between skeleton frames, the proposed GCN-Mamba framework achieves a better trade-off between computational efficiency and assessment accuracy.

[0117] Specifically, this invention visualizes attribution patterns and learned edge importance on the KIMORE dataset. These visualizations aim to provide clinically meaningful insights into the prediction process by answering four key questions: which time phases are emphasized, which joints contribute more to the prediction, which anatomical groups are more relevant to motion quality, and which skeletal connections are highlighted by the model.

[0118] Figures 2 to 7 A detailed visualization of expert-performed squat samples is presented. For example... Figure 3As shown, the time attribution curve assigns a high response to a continuous time interval, and the selected keyframe lies within this high-response region. This indicates that the model focuses on key motion phases rather than isolated frames, which is consistent with the fact that motion quality in rehabilitation assessments is typically determined by the key phases of motion execution.

[0119] Figure 4 Spatial attribution is projected onto the 3D skeleton at the selected keyframe, where larger and more yellow markers indicate joints that contribute more. Figure 5 and Figure 6 Further spatial attribution was aggregated at the anatomical group level. For expert-performed squat samples, the model assigned higher group-level attributions to the spine and lower limbs. This aligns with the standard execution of squats, where trunk stability and lower limb control are crucial for movement quality. Furthermore, Figure 7 The learned edge importance is demonstrated, showing that the adaptive edge-weighted graph convolutional branch highlights discriminative skeletal connections during evaluation. These results illustrate that the proposed framework can provide interpretation at multiple levels, including temporal phases, individual joints, anatomical groups, and skeletal connections.

[0120] Figures 8-19 Further comparisons were made of group-level attribution patterns for pelvic rotation and squats under different quality scores. For pelvic rotation, expert samples showed dominant attribution to the lower limbs, consistent with the standard execution pattern of this movement. During pelvic rotation, the feet typically act as the fulcrum, while the lower limbs serve as the primary support structure for body rotation. High-scoring non-expert samples exhibited a similar group-level attribution pattern, indicating that their movement was close to standard execution. Conversely, low-scoring non-expert samples shifted dominant attribution to the left arm, deviating from the expected movement pattern and providing a possible explanation for their lower quality scores. A similar phenomenon was observed for squats. Expert samples primarily activated the spine and lower limbs, reflecting the importance of trunk posture and lower limb control in standard squats. High-scoring non-expert samples had a similar attribution distribution, indicating relatively standard execution. However, for low-scoring non-expert samples, dominant attribution was concentrated on the right arm, which is inconsistent with the key biomechanical requirements of squats. This anomalous attribution pattern may have contributed to their lower quality scores.

[0121] Overall, the visualization results demonstrate that the proposed FSTA head can provide interpretable temporal and spatial explanations for regression-based rehabilitation assessments. Temporal attribution identifies key motor phases, spatial attribution highlights important joints, and group-level attribution reveals contribution patterns at the body part level. Furthermore, the learned edge importance complements the attribution-based explanation by showcasing skeletal connectivity emphasized by adaptive edge-weighted graph convolutional branches. These results support the interpretability of the proposed framework and its potential to provide clinically relevant feedback in rehabilitation training movement assessments.

[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.

Claims

1. A progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution, characterized in that, include: The original human skeleton sequence data is embedded and mapped to obtain a C-dimensional distribution space. Embedded features Where T is the time length of the human skeleton sequence data, V is the number of joints in the human skeleton sequence data, and C is the dimension of the data features after embedding mapping modeling. Establish a bidirectional time-dependent model within the state-space model, and embed the features. Each joint trajectory is independently analyzed using a two-way time-dependent model to obtain forward features. With backward features Then, the features are fused to obtain the output features. ; Input the prior data of anatomical grouping into the state-space model Establish a group-perception time model, based on the prior data of each anatomical group. Embedding mapping modeling yields the corresponding embedding features. , and embed the corresponding features Prior data for each anatomical group were obtained using a group-perception time model. Corresponding output features Then the output of each group Aggregation yields output features 2; In the state-space model, a time-to-space refined model is established using a spatial refinement scan method. This is achieved by analyzing the feature outputs of the bidirectional time-dependent model and the group-aware time model. and 2. Perform spatial refinement scanning to obtain spatial refinement features. Then refine the spatial features. Residual fusion with temporal modeling features yields global output features. ; A spatiotemporal attribution model is established using the decompositional spatiotemporal attribution FSTA head, and global output features are analyzed. Mapped as spatial and temporal attribution components and The model then performs a weighted fusion of spatial and temporal attributions to obtain a quality prediction score for rehabilitation training exercises represented by human skeletal sequence data. ; Based on the output of the spatiotemporal attribution model and A visualization feedback mechanism based on attribution is established to output a visualization of rehabilitation training execution scores, including time attribution curves, spatial attribution 3D graphs, anatomical group perception spatial attribution 3D graphs, anatomical group perception spatial attribution bar charts, and edge importance 3D graphs.

2. The progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution as described in claim 1, characterized in that, in, The original human skeleton sequence data is The embedded mapping modeling is expressed as follows: in, Original human skeleton sequence data The The first frame Data for each joint; For embedding features The The first frame Data for each joint; For learning parameters, ; It is a linear projection, and ; To learn the joint embedding position, The final embedded features are denoted as ; The input feature dimension associated with each joint in the human skeleton sequence data.

3. The progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution as described in claim 1, characterized in that, The state-space model includes the embedding feature sequence. Each joint As input to the selective state-space model, it is processed through a selective state-space scanning mechanism. The state-space model at time 1 The state is updated by the previous state and the current input, generating the potential state. Then the obtained state is mapped to output features. ; in, , and Represent and The discretized matrix; This represents the discretization step size, which depends on the input. The state transition matrix is ​​a learnable matrix. , For a parameter matrix that depends on the input, This represents the state dimension of the state-space model.

4. The progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution according to claim 1, characterized in that, The bidirectional time-dependent model includes mapping the input sequence to the output sequence by independently applying state-space recursion to each joint trajectory along the time dimension, as shown in the following expression: in These are forward and backward time representations, respectively. It is a time reversal operator applied along the time dimension. To perform a selective state-space scan for each joint along the time dimension, For C-dimensional distribution space Embedded features in For output features.

5. The progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution according to claim 1, characterized in that, The group-aware time model includes dividing each joint. Groups This includes the trunk, upper limbs, or lower limbs, among which For the first Grouping key points into groups, and for each group Extract the corresponding feature subsequences : in, To perform a selective state-space scan for each joint along the time dimension, To obtain the group-level feature representation, the outputs of each group are aggregated using normalized scatter-add to output the features. : in, For grouping The corresponding all-one cover mask, Scatter() is the normalized aggregation.

6. The progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution according to claim 1, characterized in that, The aforementioned time-to-space refinement model, at each time step ,right 1 and 2. Perform a selective state-space scan along the joint dimension to obtain the spatial scan feature sequence; the expression is as follows: in Selective state-space scanning applied along the joint dimension. and The spatial scanning feature sequence is then used. The output features of the temporal processing model are then fused with the spatial scanning features using residuals to obtain a refined representation, as shown in the following expression: in, Representation layer normalization, This is the refined representation.

7. The progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution according to claim 1, characterized in that, A progressive backbone network based on a bidirectional time-dependent model, a group-aware time model, and a time-to-space refinement model is used to construct a rehabilitation movement model. This model comprises two sequential modules: CSTEB and SSTRB. These two modules have the same network architecture, and their network structure expressions are as follows: Among them, the input features of the rehabilitation movement model are: CSTEB module For the SSTRB module, This is the output of CSTEB; the output expression is: ; First, parallel complementary feature modeling is performed. The CSTEB and SSTRB modules in the backbone network execute two complementary transformations in parallel. The first branch applies graph-based relational modeling. The expression is: in For stacked AE-GC adaptive graph convolutional modules; For the output of CSTEB, the second branch applies multiple selective state-space modeling. The expression is: in, For and The parallel selective state-space modeling module, depending on the module configuration, includes a bidirectional time-dependent model, a group-aware time model, and a time-to-space refinement model, through... Modeling topology-aware relational structures. Capture context-adaptive sequence dynamics; Then, on as well as The output is adaptively fused using an input-conditional adaptive fusion mechanism. First, the input of the rehabilitation action model is... Perform global average pooling on the time and joint dimensions: in, Represents the pooling result, Represents pooling operations, express The Frame number Feature data of each joint point represent 3D real vector space; Lightweight gating network Generate normalized branch weights: in, Represents a soft maximum function. Represents a 2-dimensional real vector space; Fusion characterization The calculation is as follows: and For the above formula The first and second components in; Following adaptive fusion, the selective state-space modeling module is further applied sequentially. The fusion is refined along the time dimension to obtain the fusion features. The expression is as follows: Application module-level residual connections and layer normalization generate output: in For layer normalization; Through progressive hierarchical modeling, a two-stage modeling process is formed that gradually refines the motion representation; Given embedding features Hierarchical modeling is performed using the CSTEB module. embed features Transformed into first-layer global features SSTRB module take over As input, it is refined into a second-layer global feature. The expression for this process is as follows: 。 8. The progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution according to claim 1, characterized in that, The spatiotemporal attribution model first predicts a dense score map frame by frame and joint by joint. The expression is as follows: in It is a two-layer fully connected network; Then, the space and time logits are calculated using global pooling, as shown in the following expression: in, and and They employ similar structures, but their parameters are independent. Second layer global features The Frame data, for The The data for each joint is then normalized using a temperature-controlled softmax function to obtain the logits. in Spatial attribution for joints, For temporal attribution on the frame, τ>0 is the temperature parameter in softmax normalization, used to control the concentration of normalized logits; Then, spatiotemporal attribution weights for skeleton joints and time steps are constructed through outer product factorization: in Spatiotemporal attribution weights; The attribution at each spatiotemporal location is decomposed into separable temporal and spatial contributions, given the input skeleton sequence. Quality score of rehabilitation exercise execution The global predicted value is obtained through weighted aggregation, as shown in the following expression: 。 9. The progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution according to claim 1, characterized in that, The attribution-based visualization feedback mechanism includes generating a temporal attribution vector using the FSTA head for each input skeleton sequence. and spatial attribution vector ,Will Visualize the time attribution curve, label representative keyframes and high response intervals, and identify the representative keyframes as those with the highest time attribution, as shown in the following expression: Highlight the high response time range and determine the performance of a specific frame in the visualization. If the attribution satisfies the following condition, it is considered a high-response frame, as expressed below: in, Indicates the first The temporal attribution value of the frame; in the experiment, it was taken as... , and , This indicates a maximum value operation, where consecutive high-response frames are merged into time segments, and only segments with a length of at least 2 frames are retained for highlighting; Select keyframe Above, the pose is normalized using reference joints and bone lengths, and spatial attribution values ​​are encoded by marker size and color intensity; The spatial attribution value is further defined in the predefined anatomical group. In the aggregation process, when there is overlap between groups, the joint contribution is redistributed proportionally to avoid duplicate counting, as shown in the following expression: in For including joints The number of groups, This represents the spatial attribution of the k-th group, and the aggregated group-level attribution. Highlight body groups that contribute to motion quality in the 3D visualization; Next, construct a dynamic adjacency matrix, expressed as follows: in, It is a dynamic adjacency matrix. Represents the skeleton adjacency matrix. This is the scaled edge importance matrix; This is a learnable edge importance matrix in the human skeleton model. Since skeleton connections are treated as undirected connections in visualization, the edges... Importance The learned edge importance is calculated using a symmetric method, as shown in the following expression: in, Representing an edge The corresponding importance value after scaling, Representing an edge The corresponding scaled edge importance value is then mapped to the line width and color intensity in the 3D skeleton visualization, making connections with larger values ​​visually stand out.

10. A computer storage medium storing a computer program executable on a processor, characterized in that, When the processor executes the computer program, it implements the progressive rehabilitation training movement assessment method based on decompositional spatiotemporal attribution as described in any one of claims 1-9.