Rehabilitation monitoring system based on multi-modal sensing and AI evaluation

The rehabilitation monitoring system, which utilizes multimodal sensing and AI assessment, solves the problems of data interruption and misidentification in complex environments of existing systems. It achieves stable monitoring and accurate rehabilitation assessment around the clock, supports personalized treatment plans, and shortens the rehabilitation cycle.

CN121237387AActive Publication Date: 2025-12-30LEDETANG (SHANGHAI) DIGITAL MEDICAL TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511796876.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2025-12-30
Estimated Expiration
2045-12-02

AI Technical Summary

Technical Problem

Existing rehabilitation monitoring systems are prone to data interruption or misidentification in complex environments, leading to distorted rehabilitation assessments. They are unable to adapt to the unique movement patterns of different patient groups and lack effective data quality assessment and correction mechanisms, which affects treatment decisions and efficacy.

Method used

The rehabilitation monitoring system employs multimodal sensing and AI assessment. Through comprehensive analysis of visual skeletal tracking, inertial measurement units, and pressure distribution data, it establishes a multimodal fusion correction model to identify and compensate for data interruptions and errors, generate continuous and complete skeletal motion sequences, and provide accurate rehabilitation assessments.

Benefits of technology

It improves the reliability and continuity of rehabilitation monitoring, ensures stable monitoring around the clock, provides objective assessment of rehabilitation progress, supports personalized treatment plans, shortens the rehabilitation cycle, and improves the quality of functional recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237387A_ABST
    Figure CN121237387A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical rehabilitation, and discloses a rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation. A multi-dimensional skeleton key point confidence evaluation mechanism is established by integrating three types of sensing data of visual skeleton tracking, an inertial measurement unit and pressure distribution, and potential tracking interruption risk frames and error recognition candidate frames are recognized in real time. And a multi-modal fusion correction model is adopted to carry out accurate compensation on error recognition, missing data are reconstructed in combination with a skeleton motion mode memory bank and a bidirectional time sequence prediction algorithm, and a continuous and complete all-day skeleton motion sequence is generated. On the basis of the calculation, the time-phased activity level index and the rehabilitation action completion degree score of the patient are calculated, and an objective and quantitative rehabilitation evaluation basis is provided. According to the invention, the stability and continuity of skeleton tracking in a complex rehabilitation environment are improved, so that a medical team can obtain a real and complete all-day activity portrait of a patient, accurately position a rehabilitation bottleneck, formulate a personalized treatment scheme and promote functional recovery of the patient.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical rehabilitation, more particularly, the present application relates to a rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation. BACKGROUND

[0002] The existing rehabilitation treatment monitoring system faces some technical challenges in actual clinical application. In a typical rehabilitation environment, the visual monitoring system frequently encounters various interference factors: partial obstruction of rehabilitation equipment, visual obstruction caused by the movement of medical staff, self-obstruction of the patient's body, and complex situations such as changes in indoor light over time, resulting in frequent interruption or incorrect identification of skeletal key point tracking. For example, when the patient is doing balance training, the support equipment often obstructs the lower limb key points, causing the system to incorrectly infer the joint position; when a stroke patient is doing upper limb rehabilitation training, due to slow or shaking movements, the system often misjudges normal but slow movements as tracking errors and discards the data. These problems are particularly prominent in long-term monitoring. The traditional visual tracking system has data interruption or serious distortion for some time in all-day monitoring, and lacks effective warning and repair mechanisms. When the interruption occurs, the system usually uses simple linear interpolation or directly discards abnormal data, which completely ignores the biomechanical characteristics and continuity rules of human movement. Especially when the patient is doing complex actions such as turning around and converting from sitting to lying down, the system has a high false alarm rate, resulting in a serious underestimate of the patient's activity ability. More importantly, these data loss and errors will be amplified in the all-day activity volume statistics, making the rehabilitation evaluation distorted: the system cannot distinguish between the patient's true rest and data loss, cannot accurately capture short but important activity peaks (such as independent standing attempts), and lacks a reliable correction mechanism for incorrect identification, ultimately causing the medical team to make treatment decisions based on incomplete or inaccurate data, directly affecting the rehabilitation efficacy evaluation and the development of individualized treatment plans. The existing system also generally lacks real-time quantitative evaluation of data quality, cannot take differentiated processing strategies for data of different confidence levels, and is difficult to adapt to the unique movement patterns and rehabilitation needs of different patient groups (such as the elderly, children, and different disease types), making the system less practical in complex and variable clinical environments.

[0003] In view of this, the present application proposes a rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation to solve the above problems. SUMMARY

[0004] In order to overcome the above-mentioned defects of the prior art, in order to achieve the above-mentioned purposes, the present application provides the following technical scheme: a rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation, comprising: a data acquisition module for acquiring visual skeletal tracking data, inertial measurement unit data, and pressure distribution data during the patient's rehabilitation process; The confidence evaluation module is configured to perform multi-dimensional confidence evaluation on image features of each skeletal key point in the visual skeletal tracking data, and generate a skeletal key point confidence matrix. The interruption risk identification module is configured to identify potential tracking interruption risk frames and error identification candidate frames based on a time sequence change rate of the skeletal key point confidence matrix. The multi-modal data fusion module is configured to timestamp-align inertial measurement unit data and pressure distribution data, and construct an auxiliary sensing feature vector; and establish a multi-modal fusion correction model according to kinematic consistency constraints of the auxiliary sensing feature vector and the visual skeletal tracking data. The reconstruction module is configured to perform error compensation on skeletal key point coordinates in the error identification candidate frames based on the multi-modal fusion correction model; extract joint angle time sequence features and speed time sequence features of the visual skeletal tracking data, and construct a skeletal motion pattern memory bank. For the potential tracking interruption risk frames, the skeletal key point data missing is reconstructed in combination with the skeletal motion pattern memory bank and a bidirectional time sequence prediction algorithm. The compensated skeletal key point coordinates and the reconstructed skeletal key point data are spliced to generate a continuous and complete all-day skeletal motion sequence. The rehabilitation evaluation module is configured to calculate time-periodic activity indicators and rehabilitation action completion degree scores of the patient based on the all-day skeletal motion sequence. The above various modules are connected through wired and / or wireless modes to realize data transmission between the modules.

[0005] The rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation has the following technical effects and advantages: The present application improves the reliability and continuity of rehabilitation treatment monitoring. Through the innovative multi-dimensional monitoring mechanism, stable all-weather monitoring capability is maintained even in the case of partial obstruction of the patient by the rehabilitation device, interference by the medical staff walking around, or changes in light conditions, ensuring the integrity and consistency of rehabilitation evaluation data. This high robustness enables the clinical medical team to obtain a true and complete all-day activity profile of the patient, accurately distinguish between rest and activity periods, capture short but clinically significant functional motion attempts, and fully reflect the dynamic changes in rehabilitation progress. In terms of rehabilitation evaluation dimensions, the present application provides a multi-level evaluation system including activity intensity, quality, distribution, and rehabilitation special motion completion, enabling the medical team to accurately locate rehabilitation bottlenecks and adjust treatment plans accordingly. For patients, the intuitive progress feedback provided by the present application significantly improves rehabilitation compliance and confidence, replacing subjective feelings with objective and quantifiable results, and inspiring the motivation to continue participating in rehabilitation training. In clinical practice, the present application reduces the observation and recording burden of medical staff, freeing up more time for treatment itself, while providing detailed data support to enable the development and adjustment of rehabilitation programs based on objective data rather than traditional experience-based judgments. The present application optimizes the allocation of rehabilitation resources, shortens the patient rehabilitation period, and improves the quality of functional recovery. BRIEF DESCRIPTION OF DRAWINGS

[0006] Figure 1 A rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation according to the present application. DETAILED DESCRIPTION

[0007] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0008] The present application provides a rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation. The execution subject of the system includes but is not limited to the following: rehabilitation medical devices, intelligent wearable terminals, rehabilitation training monitoring platforms, patient rehabilitation data analysis centers, and other general computing nodes of the present application. The AI evaluation system includes but is not limited to the following: cloud-based rehabilitation evaluation engines, distributed multi-modal data processing systems, and intelligent rehabilitation motion analyzers.

[0009] Please refer to Figure 1 In the embodiments of the present application, the rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation includes: The data acquisition module is used to acquire visual skeletal tracking data, inertial measurement unit (IMU) data, and pressure distribution data during the patient's rehabilitation process. Visual skeletal tracking data is obtained by capturing the coordinate sequence of key skeletal points during patient movement using a depth camera or a regular camera. IMU data records the acceleration, angular velocity, and posture parameters of the patient's joint movements using wearable sensors. Pressure distribution data records the plantar pressure distribution during standing and walking using smart mats or insoles. These three types of data complement each other, forming a multimodal comprehensive monitoring system that provides rich raw data for subsequent analysis.

[0010] The confidence assessment module performs multi-dimensional confidence evaluation of the image features of each skeletal keypoint in the visual skeleton tracking data, generating a skeletal keypoint confidence matrix. This module comprehensively evaluates the reliability of each keypoint by analyzing its image stability features, skeletal structure consistency, temporal stability, and visibility, forming a quantified confidence matrix. This confidence assessment provides crucial information for subsequent error identification and compensation, ensuring the system can accurately determine which keypoint data requires optimization.

[0011] The interruption risk identification module identifies frames with potential tracking interruption risks and erroneous candidate frames based on the temporal change rate of the skeletal keypoint confidence matrix. This module accurately identifies frames that may cause problems in the visual tracking system by analyzing the temporal decay rate of confidence, the proportion of low-confidence keypoints, anomalous positional jumps, and anatomical constraint verification, providing targets for subsequent compensation and reconstruction. This risk identification mechanism proactively detects data quality issues, preventing erroneous data from being passed to subsequent analysis stages.

[0012] The multimodal data fusion module timestamps and aligns inertial measurement unit (IMU) data with pressure distribution data to construct an auxiliary sensing feature vector. Based on the kinematic consistency constraints between the auxiliary sensing feature vector and the visual skeleton tracking data, a multimodal fusion correction model is established. This module achieves precise alignment of multi-source data through a time synchronization algorithm and establishes a mapping relationship between different modalities based on kinematic principles, providing a correction mechanism for errors in visual skeleton tracking. This fusion strategy effectively utilizes the complementarity of various sensor data, improving the system's robustness and accuracy.

[0013] The reconstruction module, based on a multimodal fusion correction model, compensates for errors in the coordinates of skeletal key points in incorrectly identified candidate frames. It extracts temporal features of joint angles and velocity from visual skeletal tracking data to construct a skeletal motion pattern memory. For frames with potential tracking interruption risks, it reconstructs missing skeletal key point data by combining the skeletal motion pattern memory with a bidirectional temporal prediction algorithm. Finally, it concatenates the compensated skeletal key point coordinates with the reconstructed skeletal key point data to generate a continuous and complete 24-hour skeletal motion sequence. This module uses various intelligent algorithms to repair identified problematic data, ensuring the continuity and integrity of the skeletal motion sequence and providing high-quality foundational data for rehabilitation assessment.

[0014] The rehabilitation assessment module calculates patients' activity levels and rehabilitation movement completion scores based on a full-day skeletal movement sequence. By analyzing patients' daily movement data, this module quantifies the amount of activity and the quality of movements during the rehabilitation process, providing medical personnel with an objective and quantitative assessment of rehabilitation progress and supporting the optimization and adjustment of rehabilitation plans.

[0015] The modules are connected via wired and / or wireless means to enable data transmission between them.

[0016] In this embodiment of the invention, the detailed implementation steps for performing multi-dimensional confidence assessment of the image features of each skeletal key point in visual skeleton tracking data and generating a skeletal key point confidence matrix include: The local gradient intensity, edge response value, and texture contrast of each skeletal keypoint in the image are extracted as image stability features. Image stability features are fundamental visual parameters for evaluating the reliability of keypoints, reflecting their saliency and stability in the image. The extraction process first defines a fixed-size pixel region (typically 16×16 pixels) around the keypoint, and then calculates the image features of this region: the local gradient intensity is calculated using the Sobel operator to calculate the average pixel gradient magnitude, reflecting edge strength; the edge response value is calculated using the Harris corner response function, reflecting corner characteristics; and the texture contrast is calculated using the contrast feature of the local region's gray-level co-occurrence matrix, reflecting texture richness. These three features form the image stability vector of the keypoint; higher values ​​indicate that the keypoint is more stable and reliable in the image, and easier to accurately identify.

[0017] The bone length ratio between skeletal key points and adjacent key points is calculated, and this ratio is compared with the preset human skeletal anatomical proportions to obtain a skeletal structure consistency score. The skeletal structure consistency score, based on human anatomy knowledge, assesses whether the key point positions conform to normal human structure. The calculation process first establishes the connection relationships between skeletal key points, defining adjacent key point pairs; then, the distance between each pair of adjacent key points is calculated to obtain the bone length; the ratio of each pair of bone lengths is calculated to obtain the bone length ratio; finally, the actual ratio is compared with the standard human skeletal proportions to calculate the degree of deviation. The skeletal structure consistency score is calculated using the following formula: ; in, Scoring for skeletal structure consistency This is a proportion of the actual bone length. To pre-determine the anatomical proportions of the human skeleton, For proportion to quantity, This is the scaling factor (usually 3-5).

[0018] The score ranges from [0, 1], with higher values ​​indicating that the skeletal structure conforms more closely to human anatomy and that the key point locations are more reliable. This scoring mechanism effectively identifies misidentifications that violate human anatomy, such as abnormal limb elongation or compression.

[0019] The temporal stability score of keypoints is calculated by statistically analyzing the positional jitter of skeletal keypoints across consecutive frames and then using Kalman filtering to determine the deviation between the predicted and actual detected positions. This temporal stability score assesses the stability of keypoints over time and is particularly important for identifying jumps and jitter. The evaluation process first establishes a Kalman filter model for each keypoint, predicting its theoretical position in the current frame based on historical positions and velocities. Then, the Euclidean distance between the predicted and actual detected positions is calculated to obtain the prediction deviation. Finally, the temporal stability score is calculated based on the prediction deviation. The calculation formula is as follows: ; Where TS is the temporal stability score, d is the Euclidean distance between the predicted position and the actual position, and σ is the standard deviation parameter (usually dynamically set according to the type of skeletal keypoints, with core keypoints such as the torso being smaller and the extremities of the limbs being larger).

[0020] The score ranges from (0, 1), with values ​​closer to 1 indicating higher temporal stability and less jitter and jumps. This scoring mechanism is particularly sensitive to detecting sudden abnormal movements and can effectively filter out rapid jumps that are irrationally unreasonable in exercise physiology.

[0021] Occlusion probability is estimated for the pixel region surrounding skeletal keypoints. Visibility scores are obtained by analyzing the continuity of the color histogram and abrupt changes in the depth gradient. The visibility score assesses whether keypoints are occluded by other objects and is a crucial dimension for judging data reliability. Different strategies are employed for depth cameras and conventional cameras: for depth cameras, the distribution of depth values ​​around keypoints is analyzed to detect abrupt changes in the depth gradient and identify potential occlusion boundaries; for conventional cameras, changes in the color histogram of the region surrounding keypoints in consecutive frames are analyzed to detect sudden histogram discontinuities and identify potential occlusion events. The two methods are combined to calculate the final visibility score, reflecting the complete visibility of the keypoints. The score ranges from [0, 1], with higher values ​​indicating better visibility and lower occlusion. This scoring mechanism is particularly effective in detecting partial and complete occlusion, providing an important basis for subsequent data compensation.

[0022] Image stability features, skeletal structure consistency scores, temporal stability scores, and visibility scores are weighted and fused to generate a skeletal keypoint confidence matrix. The confidence matrix is ​​a comprehensive representation of the multi-dimensional evaluation results, intuitively reflecting the reliability level of each keypoint in each frame. The fusion process uses a weighted average method, with weights dynamically adjusted according to keypoint type and application scenario: core joints (such as the hip and shoulder joints) are typically assigned higher weights to skeletal structure consistency; distal joints with large range of motion (such as the wrist and ankle) are assigned higher weights to temporal stability; and easily occluded joints are assigned higher weights to visibility scores. The weighted fusion formula is: ; in, These are the element values ​​of the confidence matrix. This represents the normalized value of the image stability feature. Scoring for skeletal structure consistency The time series stability score is given. For visibility score, , , , These are the weighting coefficients, and .

[0023] Each element of the confidence matrix ranges from [0, 1], with higher values ​​indicating greater reliability of the critical point. This matrix provides a precise quantitative assessment basis for subsequent interruption risk identification and error compensation, and is the core quality control mechanism of the system.

[0024] In this embodiment of the invention, the detailed implementation steps for identifying potential tracking interruption risk frames and erroneous identification candidate frames based on the temporal change rate of the skeletal keypoint confidence matrix include: The confidence decay rate is denoted as the confidence decay gradient of the skeletal keypoint confidence matrix between adjacent frames. The confidence decay rate is a key indicator of changes in detection and tracking quality, quantifying the temporal trend of confidence change. The calculation process first obtains the confidence values ​​of corresponding keypoints in two adjacent frames, then calculates the difference and divides it by the inter-frame time interval to obtain the confidence change rate per unit time. For systems with stable frame rates, the confidence difference can be calculated directly. A negative confidence decay rate indicates a decrease in confidence; a larger absolute value indicates a more drastic decrease and a higher potential risk. The system calculates the confidence decay rate for each keypoint separately and can set different warning thresholds based on the importance of the keypoint. Core keypoints typically have lower warning thresholds and are more sensitive to quality changes.

[0025] The percentage of keypoints with a confidence level below a preset confidence threshold in a single frame is recorded as the intra-frame low confidence ratio. This ratio is a comprehensive indicator of overall frame quality, reflecting the overall tracking quality level of the current frame. The calculation process first sets a preset confidence threshold (typically 0.4-0.6, adjusted according to the application scenario), then counts the number of keypoints in the current frame with a confidence level below this threshold, and finally divides this number by the total number of keypoints to obtain the ratio. A higher ratio indicates poorer frame quality and a greater risk of tracking interruption. The system can dynamically adjust the threshold and evaluation strategy based on different skeletal models (such as a 25-point full-body model, a 15-point upper-body model, etc.) to ensure the targeted and accurate identification of risks.

[0026] Frames with a confidence decay rate greater than a preset decay threshold and a low-confidence proportion exceeding a preset proportion threshold are marked as frames at risk of potential tracking interruption. This step comprehensively considers both temporal changes and current state to accurately identify risk frames that may soon experience tracking interruption. The marking process employs a dual-threshold judgment strategy: the preset decay threshold is typically set to -0.1 to -0.2, representing a 10%-20% decrease in single-frame confidence; the preset proportion threshold is typically set to 0.3-0.4, representing 30%-40% of key points in a low-confidence state. When both conditions are met simultaneously, the system determines that the current frame has a high risk of tracking interruption and initiates a defense mechanism in advance to prepare for subsequent data reconstruction. This early warning mechanism significantly improves the system's ability to respond to tracking interruptions and reduces the impact of data loss.

[0027] The system detects the position jump distance of skeletal keypoints between consecutive frames. When the position jump distance exceeds the reasonable motion range calculated based on historical motion velocities, it is marked as an abnormal jump frame. Abnormal jump detection is a crucial step in identifying sudden errors, based on the principle of continuous human motion, to identify physically unreasonable position changes. The detection process first calculates the displacement distance of each keypoint between adjacent frames, and then establishes a reasonable motion range model for that keypoint based on the average velocity and standard deviation of N historical frames (usually 5-10 frames). Reasonable range of motion ; in, The historical average speed For the speed standard deviation, This is the inter-frame time interval. This is the expansion factor (usually taken as 2.5-3.5).

[0028] When the actual displacement exceeds this range, the system determines it as an abnormal jump. This determination method based on adaptive historical velocity can adapt to the different movement characteristics of different patients and the speed differences of different rehabilitation movements, thus improving the accuracy and adaptability of abnormality detection.

[0029] Anatomical constraint verification is performed on key skeletal points in abnormal jump frames to calculate whether joint angles exceed the range of human physiological activity. Frames violating the range of physiological activity are marked as candidate frames for incorrect identification. Anatomical constraint verification is the final step in confirming whether an abnormal jump is a misidentification. Based on medical knowledge of the range of human joint activity, it identifies physiologically impossible postures. The verification process first calculates the three-dimensional angles of major joints (shoulder, elbow, wrist, hip, knee, ankle, etc.) based on the coordinates of key points. Then, it queries a preset database of human joint range of activity to determine whether the angles of each joint are within a reasonable range. The range of activity in the database is personalized according to the patient's age, gender, and rehabilitation stage to ensure the adaptability of the judgment criteria. When the number of joints exceeding the reasonable range is greater than a preset threshold, or when there is a significant change in bone length, the system confirms it as a candidate frame for incorrect identification and marks it as a target requiring error compensation. This biomechanical and anatomical-based verification mechanism greatly improves the reliability of misidentification judgment and avoids erroneous intervention in normal but large-amplitude movements.

[0030] In this embodiment of the invention, the detailed implementation steps for aligning the inertial measurement unit data and pressure distribution data with timestamps to construct the auxiliary sensing feature vector include: Cross-correlation analysis was performed on the sampling timestamps of the inertial measurement unit (IMU) data and the frame timestamps of the visual skeleton tracking data to identify the system time delay offset. The time delay offset is a key parameter for achieving accurate alignment of multimodal data, quantifying the time difference between different sensing systems. The analysis process is based on cross-correlation techniques in signal processing. By calculating the change in similarity between two time-series signals over time, the time offset value corresponding to the maximum similarity is identified. Specifically, clearly defined patient motion events (such as standing, turning, and raising an arm) are selected as feature points. The corresponding feature responses are extracted from each modality of data, and the time difference between the responses is calculated. A stable system time delay offset is obtained by averaging multiple samples. This offset is typically in the millisecond range, but it is crucial for accurate motion analysis, especially in high-speed motion and fine motor skill assessment scenarios.

[0031] Based on the system time delay offset, timestamp correction is performed on the inertial measurement unit (IMU) data and pressure distribution data to achieve synchronization with the visual skeleton tracking data. Timestamp correction is a preprocessing step in data fusion, ensuring that data from different sources are aligned in the time dimension, providing a foundation for subsequent feature extraction. The correction process uses a linear mapping method to adjust the timestamps of the IMU and pressure distribution data according to the system time delay offset, unifying them to the time reference of the visual skeleton tracking data. For data sources with different sampling frequencies, interpolation or downsampling techniques are used to match them to the reference frequency, ensuring a one-to-one correspondence between data points. The corrected data meets the time synchronization requirements, accurately reflecting the multidimensional state information of the patient at the same moment, providing reliable data for kinematic consistency analysis.

[0032] The triaxial peak acceleration, rate of change of angular velocity, and attitude quaternions are extracted from inertial measurement unit (IMU) data. Plantar pressure center trajectory calculation and gait phase identification are then performed on pressure distribution data. This step is crucial for extracting high-value features from raw sensor data, transforming low-level sensor signals into biomechanically meaningful feature parameters. For IMU data, the extracted triaxial peak acceleration reflects motion intensity, the rate of change of angular velocity reflects motion transition characteristics, and the attitude quaternions reflect spatial orientation. For pressure distribution data, the calculated plantar pressure center trajectory reflects changes in the center of gravity, and the identified gait phases (such as ground contact and swing phases) reflect gait cycle characteristics. These extracted feature parameters have clear biomechanical interpretations and can be directly used for subsequent kinematic consistency analysis and data correction, improving the utilization efficiency of auxiliary sensor data.

[0033] The peak triaxial acceleration, rate of change of angular velocity, attitude quaternions, plantar pressure center trajectory, and gait phase are concatenated according to time windows to generate an auxiliary sensing feature vector. Feature concatenation is the final step in forming a comprehensive auxiliary sensing feature vector, combining features with different dimensions and physical meanings into a unified feature vector, which facilitates subsequent model processing. The concatenation process first sets a fixed-size time window (usually 0.5-2 seconds, adjusted according to motion characteristics), then calculates the statistics of each feature (such as mean, standard deviation, peak value, etc.) within each window, and finally combines these statistics in a predetermined order to form a feature vector with fixed dimensions. For different types of features, appropriate normalization is used to ensure dimensional consistency and prevent a single feature from dominating model behavior due to an excessively large numerical range. The generated auxiliary sensing feature vector contains rich kinematic information and has structured and standardized characteristics, providing ideal input data for subsequent multimodal fusion correction models.

[0034] In this embodiment of the invention, the detailed implementation steps for establishing a multimodal fusion correction model based on the kinematic consistency constraints between the auxiliary sensing feature vector and the visual skeleton tracking data include: Joint angular velocity vectors and center-of-mass acceleration vectors are calculated from visual skeletal tracking data to serve as visual kinematic features. Visual kinematic features are a high-level representation of skeletal motion, reflecting the core dynamic characteristics of the movement. The calculation process first obtains velocity information based on the difference in keypoint coordinates between adjacent frames, then calculates joint angles through skeletal connections, and further differs to obtain angular velocity; the center-of-mass acceleration is obtained by weighted averaging of the accelerations at each keypoint, with weights determined based on the mass distribution of different parts of the human body. These features directly reflect the mechanical state of human motion and have a clear correspondence with the physical quantities measured by inertial measurement units and pressure sensors, providing a foundation for establishing kinematic mappings. The extraction of visual kinematic features reduces the dimensionality of the original skeletal coordinate data while preserving the essential characteristics of motion, facilitating consistency analysis with sensor data from other modalities.

[0035] A kinematic mapping relationship is established between visual kinematic features and auxiliary sensor feature vectors. A consistency loss function is constructed by minimizing the feature differences between the two at the same time point. Kinematic mapping is the core mechanism of multimodal fusion, unifying data measured by different sensors into a common kinematic interpretation framework based on physical principles. The mapping relationship is established using a deep neural network model, with the auxiliary sensor feature vector as input and the predicted visual kinematic features as output. The consistency loss function is designed with two parts: the main part is the mean squared error between the predicted and actual features, reflecting the overall fitting accuracy; the regularization part introduces kinematic constraints, such as the continuity of angle changes, the coordination between center-of-mass acceleration and joint forces, etc. The complete form of the loss function is: ; in, The value of the consistency loss function. Visual kinematic features, Visual features predicted based on auxiliary sensing features, For kinematic constraint regularization, To assist in sensing feature vectors, This is the regularization coefficient.

[0036] This loss function considers both the accuracy of data fitting and the constraints of kinematic knowledge, ensuring that the mapping relationship learned by the model conforms to the principles of human biomechanics and improving the physical rationality of the prediction.

[0037] Multimodal data of patients performing standard rehabilitation movements were collected to train a deep neural network to learn the optimal weight parameters of the consistency loss function. Model training is a crucial process for establishing accurate mapping relationships. Model parameters were optimized using a large amount of real data to ensure accurate prediction of visual features. Training data collection covered various standard rehabilitation movements (such as joint flexion and extension, gait training, and balance exercises) to ensure the model's generalization ability. Dedicated datasets were also collected for different patient groups (such as different age groups and different rehabilitation stages) to improve the model's adaptability to specific populations. The network structure adopted an encoder-decoder architecture, incorporating an attention mechanism to enhance the influence of key features and introducing residual connections to improve gradient propagation. The training process employed mini-batch stochastic gradient descent, combined with learning rate decay and early stopping strategies to prevent overfitting and improve convergence speed. Cross-validation was used to determine the optimal hyperparameters, ensuring the model achieved optimal performance on the validation set and providing a reliable guarantee for subsequent calibration applications.

[0038] The prediction error of the deep neural network on newly acquired data is used as a correction feedback signal. When kinematic inconsistencies occur between the visual skeleton tracking data and the auxiliary sensing feature vector, a correction vector is generated. Correction feedback is a key mechanism for identifying visual tracking errors. Based on the principle of consistency among multimodal data, it identifies potential problems and provides correction information. The workflow first inputs the auxiliary sensing feature vector of the current frame, and then uses a trained network to predict theoretical visual kinematic features. This is then compared with features calculated from actual visual skeleton tracking data, and the degree of difference is calculated. When the difference exceeds a preset threshold, the system determines that an inconsistency exists, further analyzes the direction and magnitude of the difference, and generates a correction vector. The correction vector contains information on the direction and magnitude that need adjustment, directly guiding the subsequent coordinate correction process. This correction mechanism based on prediction error effectively utilizes the complementarity of multimodal data, can identify problems that are difficult to detect with a single modality, and improves the robustness and accuracy of the system.

[0039] Based on the direction and magnitude of the correction vector, spatial displacement correction is performed on the coordinates of skeletal keypoints in visual skeleton tracking data, establishing a multimodal fusion correction model. Spatial displacement correction is the specific step in performing the correction, transforming high-level feature-level correction information into concrete coordinate adjustment operations. The correction process first maps the correction vector from the feature space back to the coordinate space, determining the direction and distance to be adjusted for each keypoint; then, correction sensitivity is assigned based on the confidence level of the keypoint, with keypoints with lower confidence receiving larger adjustment ranges; finally, coordinate correction is performed while maintaining skeleton length constraints and joint range of motion constraints to ensure the biomechanical rationality of the correction results. The entire correction process forms a closed-loop feedback system: visual tracking provides initial estimation, auxiliary sensing provides verification reference, and adjustments are made through the correction model when the two are inconsistent, ultimately obtaining more accurate and reliable skeletal keypoint coordinates. This multimodal fusion correction model improves the system's performance in complex environments and effectively addresses common challenges such as visual occlusion and changes in lighting.

[0040] In this embodiment of the invention, the detailed implementation steps for error compensation of skeletal keypoint coordinates in erroneously identified candidate frames based on a multimodal fusion correction model include: The N keypoints with the lowest confidence levels are extracted from candidate frames of incorrect identification as the main error keypoints, where N is adaptively determined based on the proportion of low confidence levels within the frame. Identifying the main error keypoints is a crucial step in accurately locating the source of the problem, avoiding unnecessary adjustments to the overall skeletal structure. The extraction process first sorts all keypoints in the current frame according to the confidence matrix, identifying the candidate set with the lowest confidence levels; then, the value of N is dynamically determined based on the proportion of low confidence levels within the frame. A higher proportion indicates a more prevalent problem, requiring the processing of more keypoints. The formula for calculating the value of N is: ; in, This represents the total number of key points in the skeletal model. The proportion of low confidence within the frame. This is the adjustment coefficient (usually taken as 1.2-1.5).

[0041] This adaptive determination strategy ensures that the system can adjust the scope of intervention according to the severity of the problem, neither omitting key errors nor over-intervening in correctly identified parts, thus maintaining the accuracy and efficiency of the compensation process.

[0042] The key error points are input into the multimodal fusion correction model to obtain theoretical coordinates calculated based on auxiliary sensor feature vectors. This theoretical coordinate calculation forms the basis of error compensation, providing a reference position based on multimodal information fusion. The calculation process utilizes the aforementioned multimodal fusion correction model, taking the auxiliary sensor feature vector of the current frame as input, to predict the ideal visual kinematic features; then, through inverse mapping, the predicted kinematic features are converted into spatial coordinates of the key points. This calculation fully utilizes data from inertial and pressure sensors, providing alternative position estimates when visual tracking is unreliable. The theoretical coordinates have strong physical consistency, and even in cases of complete visual failure, they can still provide reasonable position estimates based on the patient's actual motion state, providing a reliable reference for subsequent compensation.

[0043] The spatial deviation vector between the detected coordinates and theoretical coordinates of key error points is calculated, and outliers exceeding three standard deviations are removed, retaining only the valid deviation vectors. Deviation analysis is a crucial step in quantifying the degree of error, identifying the specific direction and magnitude requiring compensation. The analysis process first calculates the three-dimensional vector difference between the actual detected coordinates and the theoretically predicted coordinates of each key error point to obtain the original deviation vector; then, statistical analysis is performed on the deviation vectors of all key points to calculate the mean and standard deviation of the deviation distribution; finally, based on the "three standard deviations" criterion, abnormal deviations are identified, and deviations exceeding the normal range are considered outliers and removed. This statistical screening mechanism avoids overreaction to extreme errors, improving the stability and reliability of the compensation process. The retained valid deviation vectors accurately reflect the direction and magnitude requiring compensation, providing a quantitative basis for subsequent precise adjustments.

[0044] Based on the hierarchical topology of the human skeleton, the main error key points are divided into root node key points and terminal node key points, with priority given to compensating the root node key points. This hierarchical compensation strategy is a key technique that considers the characteristics of skeletal structure, avoiding structural inconsistencies caused by simple independent adjustments. The classification process is based on the hierarchical structure of the human skeleton, dividing key points into root nodes (such as core areas like the trunk, pelvis, and scapula) and terminal nodes (such as the wrist and ankle). Root nodes are located upstream in the skeletal chain, and their positional errors will propagate and amplify to downstream nodes. Therefore, prioritizing root nodes can solve the cascading error problem and improve the overall compensation effect. This topology-based hierarchical processing strategy conforms to the biomechanical principles of human movement, ensuring the structural rationality and overall coordination of the compensation results.

[0045] For the root node keypoints, the inverse vector of the effective deviation vector is processed by a temporal smoothing filter to generate a temporal consistency compensation vector, which is then superimposed onto the original coordinates. Root node compensation is the first step in error correction and directly affects the accuracy of the entire skeletal chain. The compensation process first takes the inverse vector of the effective deviation vector as the basic compensation direction, indicating the need to move towards the theoretical coordinate direction; then, a temporal smoothing filter (such as an exponential moving average filter) is used to smooth the compensation vectors of consecutive frames, reducing the abruptness of the compensation action; finally, the smoothed compensation vector is superimposed onto the original coordinates to complete the position correction of the root node. Temporal smoothing ensures the continuity of the compensation process, avoids visual jumps caused by drastic adjustments in a single frame, and makes the corrected skeletal motion more natural and smooth. Accurate compensation of the root node lays the foundation for the subsequent adjustment of the end nodes and is a key link in the overall compensation effect.

[0046] Based on the compensated root node coordinates, the inverse kinematics propagation algorithm is used to calculate the compensated coordinates of the key points of the end nodes. This algorithm constrains the bone length to be constant and the joint angles to conform to physiological ranges. Inverse kinematics propagation is the core technology for handling end nodes, calculating the reasonable positions of a series of joints based on the connection relationships of the skeletal chain. The algorithm first establishes a skeletal chain model from the root node to the end nodes, defining the degrees of freedom and range of motion of each joint; then, it sets the target position of the end nodes, i.e., the theoretical coordinate position; finally, through iterative optimization, while keeping the bone length constant, it adjusts the angles of each joint to bring the end nodes closer to the target position, while ensuring that all joint angles are within the physiological range of motion. When the target position requirements cannot be fully met, the algorithm searches for the optimal solution under the constraints, ensuring the biomechanical rationality of the result. This propagation method based on inverse kinematics considers the integrity and connection constraints of the human skeleton, producing compensation results that are both accurate and natural, avoiding the bone deformation problems that may result from simple interpolation.

[0047] The prediction error of the compensated coordinates is verified through a multimodal fusion correction model. When the prediction error exceeds a preset error threshold, the filter parameters and constraint weights are iteratively adjusted until the threshold requirement is met. Verification and iterative optimization are the guarantee mechanisms for ensuring the quality of compensation, continuously improving compensation accuracy through closed-loop feedback. In the verification process, the compensated coordinates are re-input into the multimodal fusion correction model to calculate the consistency with the auxiliary sensing feature vectors and evaluate the compensation effect. When the prediction error exceeds the preset threshold, the system enters the iterative optimization stage, adjusting the compensation parameters and re-executing the compensation process. The parameters adjusted iteratively mainly include the smoothing coefficient of the temporal smoothing filter and the weight coefficients of various constraints in the inverse kinematics, which are optimized specifically according to the nature of the prediction error. The iteration terminates when the prediction error is below the threshold or the maximum number of iterations is reached. This model-verification-based iterative optimization mechanism ensures the reliability and accuracy of the compensation process, producing high-quality compensation results even in complex and variable rehabilitation scenarios.

[0048] In this embodiment of the invention, the detailed implementation steps for extracting the temporal features of joint angles and the temporal features of velocity from visual skeleton tracking data to construct a skeleton motion pattern memory bank include: This study uses computational vision skeletal tracking data to obtain the 3D angle sequences of major joints. A Discrete Fourier Transform (DFT) is performed on these sequences to extract the dominant frequency and harmonic components in the frequency domain, serving as the frequency domain features of the joint angles. Frequency domain feature extraction is a key technique for capturing periodic movement patterns, converting time-domain signals into frequency components to facilitate the identification of basic movement patterns. The computation process first calculates the 3D angle time-series data of major joints (such as shoulder, elbow, wrist, hip, knee, and ankle) from the coordinates of key skeletal points. Then, a Fast Fourier Transform (FFT) is applied to each angle sequence to convert the time-domain signal to the frequency domain. Finally, the main features are extracted from the spectrum, including the dominant frequency (reflecting the basic movement cycle), the dominant frequency amplitude (reflecting the main movement intensity), and the first N harmonic components (reflecting the complexity and morphology of the movement). These frequency domain features effectively capture the periodicity and rhythmicity of rehabilitation movements, which is particularly important for recognizing standard rehabilitation movements and assessing movement quality. The frequency domain representation also possesses scale invariance and time-shift invariance, making pattern recognition unaffected by changes in speed and starting point, thus improving the generalization ability of the features.

[0049] The temporal differentiation of the positional coordinates of key skeletal points yields their instantaneous velocities. The mean, variance, and peak value of these instantaneous velocities are then statistically analyzed as velocity statistical features. Velocity statistical features are crucial indicators describing the dynamics of movement, reflecting the smoothness and controllability of rehabilitation actions. The calculation process begins with numerical differentiation of the positional coordinates of the key skeletal points (typically using the central difference method to improve accuracy) to obtain the instantaneous velocities of each point in each direction. Then, the statistical features of the velocities are calculated, including the mean (reflecting overall velocity level), variance (reflecting velocity stability), peak value (reflecting maximum explosive force), and peak-to-valley ratio (reflecting acceleration and deceleration characteristics). Finally, the statistical features of each key point are combined into a unified velocity feature vector. These velocity features are directly related to the patient's motor control ability, effectively distinguishing between healthy and pathological movements, and are of great significance for assessing rehabilitation progress and identifying potential risks. Velocity features and angular features complement each other, together forming a comprehensive description of movement patterns, providing a rich feature base for subsequent analysis.

[0050] Joint angle frequency domain features and velocity statistics are input into a Long Short-Term Memory (LSTM) network to extract temporal encoding vectors for different rehabilitation movements. Temporal encoding is a crucial step in compressing high-dimensional temporal features into low-dimensional representations, capturing the essential features of the movement and facilitating subsequent storage and matching. The encoding process employs a LSTM network architecture, which is particularly well-suited for processing temporal data and can learn long-term dependencies. The network input is a sequence of temporal features, including joint angle frequency domain features and velocity statistics; the network output is a fixed-dimensional encoding vector (typically 64-128 dimensions), representing an abstract expression of the movement. Network training uses an autoencoder structure, optimizing encoding quality by reconstructing the original sequence to ensure that the encoding vector retains the key information of the original movement. The temporal encoding vectors possess high information density and discriminative power, effectively distinguishing different types of rehabilitation movements and different quality levels of execution, providing ideal feature representations for subsequent pattern classification and matching.

[0051] Cluster analysis is performed on temporal encoded vectors to group similar movement patterns into the same category, with each category corresponding to a movement pattern prototype. Cluster analysis is a key step in identifying typical movement patterns, discretizing the continuous feature space into meaningful pattern categories. The analysis process employs an improved K-means++ algorithm, finding the optimal clustering result by minimizing intra-cluster distance and maximizing inter-cluster distance. The number of clusters is automatically determined using evaluation metrics such as silhouette coefficient, typically corresponding to the number of common rehabilitation movement types. Each cluster center is defined as a movement pattern prototype, representing the standard form of that type of movement. In addition to cluster centers, the system also calculates the covariance matrix and boundary samples for each category, describing the variation range and boundary characteristics of that type of pattern. This data-driven pattern discovery method can automatically identify typical movement patterns in the patient's rehabilitation process without pre-defining all possible movement types, exhibiting strong adaptability and scalability.

[0052] A skeletal movement pattern memory bank is constructed by storing prototype movement patterns and their corresponding joint angle change curves, velocity change curves, and duration distributions. Memory bank construction is a crucial step in achieving pattern memorization and reuse, systematically storing discovered patterns and establishing a retrieval mechanism. The construction process first assigns a unique identifier to each movement pattern prototype, establishing an index structure; then, it stores the detailed features of the pattern, including temporal encoding vectors, original joint angle curves, velocity change curves, and duration distributions; finally, it establishes an efficient similarity calculation and retrieval mechanism to support subsequent pattern matching operations. The memory bank adopts a hierarchical storage structure, organizing patterns according to functional categories (such as upper limb training, lower limb training, balance training, etc.) and complexity levels, facilitating precise location of target patterns. For each pattern, the memory bank also records its adaptability and variability in different patient groups, providing a reference for subsequent personalized matching. The skeletal movement pattern memory bank is the core knowledge base of the system, continuously expanding and optimizing with system use to continuously improve the ability to recognize and reconstruct various rehabilitation movements.

[0053] In this embodiment of the invention, the detailed implementation steps for reconstructing missing skeletal keypoint data by combining a skeletal motion pattern memory and a bidirectional temporal prediction algorithm for frames with potential tracking interruption risk include: Identifying valid tracking frames before and after frames with potential tracking interruption risk, and extracting skeletal motion features from these valid frames is a prerequisite for data reconstruction. Valid frame identification is crucial for establishing a reliable range of reference data, providing a foundation for subsequent predictions. The identification process begins by filtering frames with high confidence levels within a time window (typically ±2 seconds) before and after the frame with potential tracking interruption risk, using a confidence matrix as candidate valid frames. Then, continuity verification is applied to these candidate valid frames to ensure that the skeletal motion between frames conforms to the principle of physical continuity, eliminating potentially abnormal frames. Finally, from the verified frames, N frames (typically 10-20 frames, adjusted according to motion complexity) before and after the interruption are selected as the final valid tracking frames. Skeletal motion features, including joint angle sequences, velocity sequences, and acceleration sequences, are extracted from these valid frames to provide input data for pattern matching and bidirectional prediction. Accurate identification of valid tracking frames ensures the quality of reference data for the reconstruction process and is a key prerequisite for high-quality reconstruction.

[0054] The system performs similarity matching between skeletal motion features and motion pattern prototypes in a skeletal motion pattern memory database to retrieve the most similar historical motion pattern. Pattern matching is a crucial step in determining the current motion type, providing prior knowledge guidance for subsequent predictions. The matching process first converts the motion features of valid tracking frames into temporal encoded vectors using the same method as the database construction. Then, it calculates the similarity between this encoded vector and each motion pattern prototype in the database, employing cosine similarity or Mahalanobis distance as metrics. Finally, it selects the K most similar patterns (typically K=3-5) as candidate historical motion patterns. For these K candidate patterns, the system further considers their applicability context (e.g., patient type, rehabilitation stage) and their compatibility with the current situation to determine the final most similar historical motion pattern. This content-based similarity retrieval method accurately identifies the type of rehabilitation action currently being performed, providing targeted prior knowledge for subsequent predictions and significantly improving the accuracy and specificity of reconstruction.

[0055] Based on joint angle change curves from historical motion patterns, a forward temporal prediction model is constructed to predict the skeletal keypoint positions of interrupted frames, starting from the frame before the interruption. Forward prediction is an inference process from the past to the future, predicting the state during the interruption based on the motion trend before the interruption. The prediction model adopts a hybrid strategy of fusion of autoregressive moving average (ARMA) and historical patterns: on the one hand, the ARMA model is used to capture short-term dynamic trends, and an autoregressive model is built based on the temporal data of N frames before the interruption; on the other hand, prior knowledge of historical motion patterns is incorporated, especially the expected trends and characteristics of the current stage. The two predictions are fused by weighting to form the final forward prediction result, with the weights dynamically determined according to the goodness of fit of the autoregressive model and the matching degree of the historical pattern. Forward prediction directly generates joint angle predictions for interrupted frames, which are then converted into skeletal keypoint coordinates through forward kinematics methods, completing the preliminary position prediction. This hybrid prediction strategy, combining data-driven and knowledge-guided approaches, can effectively balance short-term trends and long-term patterns, improving the accuracy and robustness of predictions.

[0056] Simultaneously, a backward temporal prediction model is constructed to predict the skeletal keypoint positions of the interrupted frames from the frames after the interruption. Backward prediction is a reverse inference from the future to the past, inferring changes during the interruption based on the state after the interruption. The prediction process is similar to forward prediction, but the processing direction is reversed. A backward autoregressive model is built using the effective frames after the interruption, while considering the reverse evolution of historical motion patterns. Backward prediction pays special attention to final state constraints, ensuring that the prediction results can smoothly transition to the known state after the interruption, avoiding abrupt changes between the prediction and reality. Backward prediction also generates joint angle predictions and converts them into skeletal keypoint coordinates, forming a position estimate from another temporal direction. Forward and backward predictions are independent and complementary, improving the overall prediction accuracy through bidirectional constraints, especially significantly enhancing the ability to handle longer interruption intervals.

[0057] The prediction results of the forward and backward temporal prediction models are weighted and fused. The weights are inversely distributed based on the temporal distance between the prediction start point and the interrupted frame to generate reconstructed skeletal keypoint data. Weighted fusion is a key step in integrating the advantages of bidirectional prediction, adaptively adjusting the influence of predictions in different directions based on temporal distance. During the fusion process, for each frame within the interruption interval, the temporal distance to the last valid frame before the interruption and the first valid frame after the interruption are calculated separately. The closer the distance, the greater the prediction weight for the direction. The weight calculation formula is as follows: ; ; in, and These are the weights for forward prediction and backward prediction, respectively. and These are the time distances from the forward prediction start point and the backward prediction start point, respectively.

[0058] Weighted fusion is performed simultaneously in both joint angle space and coordinate space to ensure consistency and smoothness of the results in both spaces. The fused skeletal keypoint data takes into account both the continuity and inertia of motion and satisfies the boundary constraints of the start and end points, forming a natural and smooth transition sequence that effectively eliminates the data discontinuity problem caused by interruptions.

[0059] The measured values ​​of the auxiliary sensor feature vector at the moment of interruption are used as physical constraints to perform kinematic rationality verification and fine-tuning on the reconstructed skeletal keypoint data. Physical constraint verification is the final guarantee to ensure reconstruction quality, utilizing the complementarity of multimodal data for cross-validation. The verification process first calculates kinematic features, such as joint angular velocity and center of mass acceleration, from the reconstructed skeletal keypoints; then, it compares these with the actual measured values ​​of the auxiliary sensor feature vector at the same moment to calculate the degree of kinematic consistency; when the consistency is lower than a preset threshold, the system initiates a fine-tuning process, iteratively optimizing and adjusting the reconstruction results until the kinematic constraints are met. The fine-tuning process maintains the overall motion trend unchanged, optimizing only local details to ensure that the reconstruction results conform to both the overall trend of the time-series prediction and the physical constraints of the auxiliary sensor data. This multimodal cross-validation mechanism significantly improves the physical rationality of the reconstruction results, ensuring that even in the absence of complete visual data, the system can still generate skeletal keypoint data that matches the patient's actual motion state based on the auxiliary sensor data.

[0060] In this embodiment of the invention, the detailed implementation steps for calculating the patient's time-segmented activity levels and rehabilitation movement completion scores based on a full-day skeletal movement sequence include: The 24-hour skeletal movement sequence is segmented according to a preset time window. For each segment, the cumulative movement distance of key skeletal points and the range of motion of joints are calculated. Time segmentation is a fundamental step in activity volume analysis, breaking down continuous data into manageable and comparable units. The segmentation process is based on a preset time window (usually 15-30 minutes, or a specific interval determined according to the rehabilitation training plan), dividing the 24-hour data into multiple continuous, non-overlapping time periods. Within each time period, the system calculates two basic activity volume indicators: cumulative movement distance, which is the sum of the displacements of each key skeletal point in three-dimensional space, reflecting overall activity volume; and joint range of motion, which is the maximum angular change range of major joints within that time period, reflecting joint flexibility. These basic indicators directly quantify the patient's activity level and intensity at different times, providing an objective basis for subsequent activity classification and scoring. Time segmentation analysis can reveal the temporal distribution pattern of the patient's 24-hour activity, identify high-activity and low-activity periods, and provide a reference for optimizing the timing of the rehabilitation plan.

[0061] The system statistically analyzes the duration of patients in different postures (sitting, standing, walking, and rehabilitation training) within each time period, identifying each posture type using a posture classifier. Posture classification is a crucial step in activity type analysis, distinguishing different types of activities and quantifying their temporal distribution. The classification process employs a machine learning model, based on the spatial distribution and dynamic changes of skeletal key points, to identify common posture types. The classifier typically uses decision tree ensembles or convolutional neural network structures, trained on a large amount of labeled data, and can accurately distinguish between static postures (such as sitting, standing, and lying down) and dynamic activities (such as walking, climbing stairs, and rehabilitation training movements). For rehabilitation training movements, the system further identifies specific training types, such as upper limb training, lower limb training, and balance training, and matches them with a pre-set rehabilitation plan. The posture classification results provide a detailed activity composition for each time period, quantifying not only the duration of various activities but also revealing the transition patterns between activities, providing rich information for a comprehensive assessment of the patient's daily functional status.

[0062] Energy consumption coefficients are assigned to different posture types. The cumulative distance traveled, range of motion of joints, and energy consumption coefficients are weighted and summed to calculate the activity level index for different time periods. Calculating the activity level index is a crucial step in integrating activity characteristics from different dimensions into a unified evaluation standard. The calculation process first assigns energy consumption coefficients to various posture activities, determined based on exercise physiology research, reflecting the energy consumption level per unit time. For example, sitting is assigned a baseline value of 1.0, standing 1.2-1.5, walking 2.0-3.0, and rehabilitation training is set at 2.0-5.0 depending on the intensity. Then, the cumulative distance traveled and range of motion of joints are standardized to eliminate dimensional differences. Finally, a comprehensive activity level index is calculated using a weighted summation formula. ; in, For activity level indicators, For posture Duration, For posture The energy consumption coefficient, For standardized cumulative distance of movement, Standardized range of motion for joints. , , These are the weighting coefficients, and .

[0063] The weights are dynamically adjusted according to rehabilitation goals, with an emphasis on endurance training. Higher, when focusing on activity level Higher, when focusing on joint function The activity level is relatively high. The activity level index at different time periods directly reflects the intensity and quality of the patient's activities at each time period, making it easy to compare different time periods horizontally and track rehabilitation progress vertically.

[0064] Movement segments matching pre-defined rehabilitation movement templates are extracted from the 24-hour skeletal motion sequence, and the dynamic time warping distance between the joint angle trajectories of the movement segments and the template is calculated. Rehabilitation movement matching is a crucial step in assessing the quality of treatment execution, identifying the rehabilitation training actually performed by the patient and evaluating its accuracy. The matching process first searches for specific pattern features in the 24-hour sequence to initially locate possible rehabilitation movement segments; then, candidate segments are precisely compared with pre-defined rehabilitation movement templates to confirm movement types and boundaries; finally, the similarity distance between the joint angle trajectories and the template is calculated using the dynamic time warping (DTW) algorithm. The DTW algorithm can handle movement comparison at different speeds and rhythms, finding the optimal matching path through non-linear time alignment to obtain a more accurate similarity assessment. The smaller the similarity distance, the closer the movement execution is to the standard, and the higher the quality. This template-matching-based method can accurately identify all rehabilitation training movements performed by patients in their daily activities and provide an objective quality assessment, offering medical personnel a reliable basis for understanding patient treatment compliance and execution quality.

[0065] Based on the dynamic time warping distance (DTW), the duration deviation of movement completion, and the joint range of motion deviation, a rehabilitation movement completion score is calculated, and a personalized rehabilitation suggestion report is generated. The completion score is a comprehensive assessment of the rehabilitation treatment effect, quantifying the accuracy and effectiveness of the patient's execution of rehabilitation movements. The score calculation comprehensively considers three key dimensions: the accuracy of the movement trajectory (assessed via DTW distance), the rationality of the movement rhythm (assessed via duration deviation), and the adequacy of the movement range (assessed via joint range of motion deviation). The scoring formula is designed as follows: ; in, Rate the completion level (0-100 points). For normalized DTW distance, For duration deviation, For standard duration, This is due to a deviation in the range of motion of the joint. For standard range of motion, , , This is a weighting coefficient, which is adjusted according to the focus of rehabilitation treatment.

[0066] The scoring results are summarized by movement type and time period to form a rehabilitation movement completion report. Simultaneously, based on the scoring results and the patient's rehabilitation stage, the system generates targeted improvement suggestions, such as adjusting movement range, optimizing movement rhythm, and enhancing specific joint mobility. This combination of quantitative completion scoring and personalized suggestions provides medical personnel with objective assessments of treatment effectiveness and optimization directions, while also providing patients with clear progress feedback and improvement goals, significantly improving the targetedness and effectiveness of rehabilitation treatment.

[0067] In this embodiment of the invention, the detailed implementation steps for verifying anatomical constraints on key skeletal points in abnormal jump frames and calculating whether joint angles exceed the range of human physiological activity include: Based on the coordinates of key skeletal points in anomalous jump frames, the flexion-extension, abduction-inversion, and rotation angles of major joints are calculated. Joint angle calculation is a fundamental step in anatomical verification, converting spatial coordinates into physiologically meaningful angular parameters. The calculation process is based on a three-dimensional anatomical coordinate system. First, a local coordinate system is established for each major joint (e.g., shoulder, elbow, wrist, hip, knee, ankle) to accurately reflect its anatomical orientation. Then, the three degrees of freedom of the joint in this coordinate system are calculated: the flexion-extension angle reflects the joint's movement in the sagittal plane, the abduction-inversion angle reflects its movement in the coronal plane, and the rotation angle reflects its rotation in the transverse plane. Euler angles or quaternions are used for angle calculation to ensure the accuracy and continuity of the values. This anatomically based angle decomposition method allows the verification results to directly correspond to medical standards, facilitating clinical interpretation and application. The calculated joint angles comprehensively reflect the current state of the skeletal system, providing accurate input for subsequent physiological range verification.

[0068] The system queries a pre-defined database of human joint range of motion to obtain the upper and lower limits of physiological angles for each joint under different genders, ages, and rehabilitation stages. Obtaining the physiological range is a crucial step in setting verification standards, ensuring the scientific rigor and personalization of the judgment. The database is constructed based on human kinesiology research and clinical rehabilitation medicine data, recording the normal range of motion of major joints under different conditions. The database structure includes multi-level parameters: basic parameters such as joint type, direction of movement, and gender differences; and personalized parameters such as age group, physical condition, and rehabilitation stage. The query process first filters the applicable range based on the patient's basic information (such as gender and age); then it is precisely adjusted according to the rehabilitation stage and specific conditions (such as postoperative status and specific diseases); finally, personalized physiological angle limits are obtained for the current patient. This personalized range setting ensures that the verification standards conform to medical norms while adapting to the specific circumstances of the patient, avoiding one-size-fits-all judgment errors. The upper and lower limits of physiological angles provide accurate references for subsequent judgments of exceeding these limits and are the core basis for anatomical verification.

[0069] The process involves determining whether the flexion / extension, abduction / inversion, and rotation angles are within the physiological limits, and counting the number of joints exceeding these limits. Boundary judgment is a core step in anatomical verification, identifying abnormal joint states that violate physiological laws. The judgment process compares the three degrees of freedom of each joint with their corresponding physiological limits to determine if they exceed the range. For angles exceeding the range, the degree of exceedance (i.e., the difference from the boundary) is further calculated to assess the severity of the abnormality. Finally, the total number of joints exceeding the limits throughout the body is counted to quantify the overall degree of anatomical abnormality. The judgment process considers the coupling relationship between angles; some angle combinations, while individually within the range, may not conform to physiological possibilities, and these are also marked as abnormal. The number of joints exceeding the limits is a key indicator for assessing the anatomical rationality of the entire frame of data; a higher number indicates a higher degree of abnormality and is more likely to be a misidentification. This verification method based on medical knowledge can accurately identify erroneous data that appears numerically reasonable but actually violates human physiological laws.

[0070] The rate of change of skeletal chain length between the abnormal jump frame and the preceding and following frames is calculated. When the rate of change of skeletal chain length exceeds a preset deformation threshold, it is determined to be a skeletal structure distortion. Skeletal structure verification is an important supplement to anatomical constraints, identifying physically impossible skeletal deformations based on the principle of constant human skeletal length. The verification process first defines multiple skeletal chains (such as upper arm-forearm chain, thigh-lower leg chain, etc.) according to skeletal connections; then, the lengths of the corresponding skeletal chains in the current frame and the preceding and following frames are calculated to obtain the rate of change of length; finally, it is compared with a preset deformation threshold (usually 1%-3%, considering measurement error) to determine whether there is abnormal deformation. The bone length should remain relatively constant within a short period of time; a rate of change exceeding the threshold clearly violates physical laws, indicating an identification error. This verification method based on rigid body constraints can effectively identify errors where the joint angles are reasonable but the spatial structure is distorted, and is an important part of anatomical verification.

[0071] Frames with a number of joints exceeding the physiological limits (above a preset threshold) or exhibiting skeletal structural distortion are identified as candidate frames for error recognition. The final decision involves a comprehensive evaluation of all verification results to determine the target frame requiring error compensation. The decision uses an "OR" logic: a frame is considered a candidate frame for error recognition if either of the following conditions is met: the joint angle condition is that the number of joints exceeding the physiological range exceeds a preset threshold (typically 15%-25% of the total number of joints, adjusted according to model complexity); the skeletal structure condition is the presence of skeletal chains with a length change rate exceeding a threshold. This multi-condition decision-making strategy improves the detection rate of errors, ensuring that potentially problematic data is not overlooked. Data identified as candidate frames for error recognition will proceed to the subsequent error compensation process, where a multimodal fusion correction model will be used to correct anatomical abnormalities and restore reasonable skeletal structure and joint states. This rigorous anatomical verification mechanism is a crucial part of the system's data quality control, ensuring that subsequent analyses are based on physiologically sound skeletal data.

[0072] In this embodiment of the invention, the detailed implementation steps for spatial displacement correction of the coordinates of key skeletal points in visual skeleton tracking data based on the direction and magnitude of the correction vector include: The correction vector is decomposed into three components along the anterior-posterior axis, left-right axis, and superior-inferior axis of the human anatomical coordinate system. Vector decomposition is a crucial step in adapting to the characteristics of human anatomy, making the correction process more biomechanically sound. The decomposition process first establishes a patient-centered anatomical coordinate system: the anterior-posterior axis points to the front and back of the patient, the left-right axis points to the left and right sides, and the superior-inferior axis points to the top of the head and the soles of the feet. Then, the original correction vector is transformed from the world coordinate system to this anatomical coordinate system, obtaining the component values ​​in the three directions. This anatomically based decomposition method makes the correction process more intuitive and controllable: the anterior-posterior component mainly affects sagittal plane motion, the left-right component mainly affects coronal plane motion, and the superior-inferior component mainly affects vertical motion. Decomposition in the anatomical coordinate system also facilitates the setting of directional correction constraints; for example, if the movement of certain joints is restricted in a specific direction, the corresponding components can be adjusted accordingly. Vector decomposition lays the foundation for subsequent personalized corrections, enabling the system to make precise adjustments based on the characteristics of different joints.

[0073] Based on the confidence matrix of skeletal keypoints, a correction sensitivity coefficient is assigned to each keypoint; the lower the confidence level, the higher the correction sensitivity coefficient. Sensitivity assignment is the key mechanism for achieving adaptive correction, adjusting the correction strength according to data reliability. The assignment process is based on the previously generated skeletal keypoint confidence matrix, calculating the correction sensitivity coefficient through an inverse proportional relationship: ; in, Key point The corrected sensitivity coefficient, Let i be the confidence value for key point i. This is the scaling factor (usually between 0.5 and 1.5). This is a smoothing factor (usually set to 0.1-0.3 to avoid the denominator being close to zero).

[0074] This formula assigns a higher sensitivity coefficient to keypoints with lower confidence levels, resulting in a more significant correction effect; while keypoints with higher confidence levels receive smaller corrections, preserving the advantages of the original data. The sensitivity coefficient typically ranges from [0, 5], with the upper limit appropriately increased for extremely low confidence levels. This adaptive correction mechanism avoids excessive intervention in reliable data while ensuring sufficient correction of unreliable data, thus optimizing the overall correction effect.

[0075] The three components of the correction vector are multiplied by the correction sensitivity coefficient of the corresponding skeletal keypoint to obtain the personalized correction vector for each keypoint. Personalized correction is a precise adjustment based on the characteristics of different keypoints, improving the adaptability and accuracy of the correction. The calculation process multiplies the three components in the anatomical coordinate system by the correction sensitivity coefficient to obtain the adjusted component values; then, they are recombined into a three-dimensional correction vector, which serves as the personalized correction vector for that keypoint. The personalized correction vector considers both the overall correction direction and the reliability characteristics of the keypoint itself, achieving a fine-grained correction strategy of "same direction, different intensity". This personalized correction method avoids the inconsistencies caused by simple, one-size-fits-all displacement correction, resulting in a more natural and harmonious final correction, especially superior in complex postures and movements.

[0076] A skeletal chain constraint is applied to the personalized correction vector to ensure that the distance between adjacent keypoints after correction remains within a reasonable bone length range. The skeletal chain constraint is a key mechanism for maintaining biomechanical rationality, preventing the correction process from compromising the integrity of the skeletal structure. The constraint process first defines the connection relationships between keypoints based on the skeletal model, marking each pair of adjacent keypoints; then, it predicts the positions of the corrected keypoints and calculates the new distances between adjacent points; the new distances are compared with the standard bone length to check if they exceed the allowable range (typically ±3% of the standard length); for correction vectors that violate the constraint, scaling or orientation adjustments are made until the bone length constraint is met. The constraint processing employs an iterative optimization strategy, adjusting step-by-step from the root node of the skeletal chain to the ends to ensure the rationality of the overall structure. This skeletal chain-based constraint processing ensures that the correction process does not introduce new biomechanical errors, maintaining the integrity and rationality of the skeletal structure, and is an important guarantee for high-quality correction.

[0077] The personalized correction vector is superimposed onto the original skeletal keypoint coordinates to complete spatial displacement correction and update the skeletal keypoint confidence matrix. Displacement execution is the final step in completing the correction, applying the calculated correction vector to the actual coordinates. The execution process involves vector addition between the constrained personalized correction vector and the original keypoint coordinates to obtain the corrected new coordinates; simultaneously, the skeletal keypoint confidence matrix is ​​updated, increasing the confidence value of the corrected keypoints to reflect the quality improvement brought about by the correction. The confidence update formula is: ; in, This is the updated confidence level. The original confidence level value. For the correction vector of the application, To the maximum allowable correction range, This is an enhancement factor (usually taken as 0.3-0.6).

[0078] The increase in confidence level is directly proportional to the magnitude of the correction; the more significant the correction, the more serious the problem, and the more significant the quality improvement after correction. The updated confidence matrix provides a more accurate quality assessment for subsequent analysis, while also recording the correction history to facilitate continuous system learning and optimization. The skeletal keypoint coordinates after spatial displacement correction not only maintain the basic characteristics of the original data but also eliminate obvious errors and anomalies, providing high-quality skeletal motion data for subsequent analysis.

[0079] This invention acquires real-time visual skeletal tracking data, inertial measurement unit data, and pressure distribution data during patient rehabilitation. It constructs a precise skeletal keypoint confidence assessment mechanism and a multimodal fusion correction model, and combines this with a skeletal motion pattern memory database to intelligently compensate for and reconstruct potential tracking interruptions and errors. Through continuous and complete 24-hour skeletal motion sequences, it achieves precise monitoring and assessment of the patient's rehabilitation process. It is highly robust, capable of handling sensor data interruptions and errors in complex rehabilitation environments, improving the continuity and accuracy of rehabilitation monitoring, and providing medical personnel with reliable rehabilitation assessment data.

[0080] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0081] It should be noted that all formulas in this manual are calculated by removing dimensions and taking their numerical values. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0082] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A rehabilitation therapy monitoring system based on multi-modal sensing and AI evaluation, characterized in that, The method comprises the following steps: a data acquisition module is used to acquire visual bone tracking data, inertial measurement unit data and pressure distribution data during patient rehabilitation; a confidence assessment module is used to perform multi-dimensional confidence assessment on image features of each bone key point in the visual bone tracking data, and generate a bone key point confidence matrix; an interruption risk identification module is used to identify potential tracking interruption risk frames and error identification candidate frames based on the time sequence change rate of the bone key point confidence matrix; a multi-modal data fusion module is used to timestamp align the inertial measurement unit data and the pressure distribution data, construct auxiliary sensing feature vectors, and establish a multi-modal fusion correction model according to the kinematic consistency constraint of the auxiliary sensing feature vectors and the visual bone tracking data; a reconstruction module is used to perform error compensation on the bone key point coordinates in the error identification candidate frames based on the multi-modal fusion correction model, extract joint angle time sequence features and velocity time sequence features of the visual bone tracking data, and construct a bone motion pattern memory bank; for the potential tracking interruption risk frames, the bone motion pattern memory bank and a bidirectional time sequence prediction algorithm are combined to reconstruct missing bone key point data; the compensated bone key point coordinates and the reconstructed bone key point data are spliced to generate a continuous and complete whole-day bone motion sequence; a rehabilitation evaluation module is used to calculate the patient's time-periodic activity index and rehabilitation action completion score based on the whole-day bone motion sequence.

2. The system of claim 1, wherein, The method comprises the following steps: extract the local gradient intensity, edge response value and texture contrast of each bone key point in the image as the image stability feature of the key point; calculate the bone length ratio between the bone key point and the adjacent key point, compare the bone length ratio with the preset human body bone anatomy ratio, and obtain a bone structure consistency score; statistically analyze the position jitter amplitude of the bone key point in the continuous frames, predict the deviation between the position and the actual detection position through Kalman filtering, and calculate the time sequence stability score of the key point; estimate the occlusion probability of the pixel region around the bone key point, and obtain the visibility score of the key point by analyzing the continuity of the color histogram and the depth gradient mutation; weight and fuse the image stability feature, the bone structure consistency score, the time sequence stability score and the visibility score to generate the bone key point confidence matrix.

3. The system of claim 1, wherein, The method comprises the following steps: calculate the confidence decrease gradient of the bone key point confidence matrix between adjacent frames, denoted as the confidence decay rate; statistically analyze the proportion of key points with confidence lower than the preset confidence threshold in a single frame, denoted as the intra-frame low confidence proportion; mark the frames with the confidence decay rate greater than the preset decay threshold and the intra-frame low confidence proportion greater than the preset proportion threshold as the potential tracking interruption risk frames. Detecting a position jump distance of the skeleton key point between continuous frames, and marking as an abnormal jump frame when the position jump distance exceeds a reasonable motion range calculated based on a historical motion speed; Performing an anatomical constraint verification on the skeleton key point in the abnormal jump frame, and calculating whether a joint angle exceeds a human physiological activity range, and marking a frame violating the physiological activity range as the error recognition candidate frame.

4. The system of claim 1, wherein, The time stamp alignment of the inertial measurement unit data and the pressure distribution data, construction of an auxiliary sensing feature vector, includes: Cross-correlation analysis of the sampling time stamp of the inertial measurement unit data and the frame time stamp of the visual skeleton tracking data to identify a system time delay offset; According to the system time delay offset, time stamp correction of the inertial measurement unit data and the pressure distribution data is performed to realize synchronization with the visual skeleton tracking data; From the inertial measurement unit data, three-axis acceleration peak value, angular velocity change rate and attitude quaternion are extracted, and plantar pressure center trajectory calculation and gait phase recognition are performed on the pressure distribution data; The three-axis acceleration peak value, the angular velocity change rate, the attitude quaternion, the plantar pressure center trajectory and the gait phase are spliced according to a time window to generate the auxiliary sensing feature vector.

5. The system of claim 1, wherein, The multi-modal fusion correction model is established according to the kinematic consistency constraint of the auxiliary sensing feature vector and the visual skeleton tracking data, including: The joint angular velocity vector and the center of mass acceleration vector are calculated from the visual skeleton tracking data as visual kinematic features; A kinematic mapping relationship between the visual kinematic features and the auxiliary sensing feature vector is established, and a consistency loss function is constructed by minimizing the feature difference at the same time; Multi-modal data of patients under standard rehabilitation actions are collected, and a deep neural network is trained to learn the optimal weight parameters of the consistency loss function; The prediction error of the deep neural network on newly collected data is taken as a correction feedback signal, and when the visual skeleton tracking data and the auxiliary sensing feature vector are inconsistent in kinematics, a correction vector is generated; Based on the direction and amplitude of the correction vector, the spatial displacement of the skeleton key point coordinates in the visual skeleton tracking data is corrected, and the multi-modal fusion correction model is established.

6. The system of claim 1, wherein, The joint angle time sequence feature and the speed time sequence feature of the visual skeleton tracking data are extracted to construct a skeleton motion pattern memory bank, including: The three-dimensional angle sequence of the main joints in the visual skeleton tracking data is calculated, and the main frequency component and the harmonic component in the frequency domain are extracted by discrete Fourier transform on the three-dimensional angle sequence as the joint angle frequency domain feature; The position coordinates of the skeleton key point are time-differentiated to obtain the instantaneous speed of the key point, and the mean, variance and peak value of the instantaneous speed are counted as the speed statistical feature; The joint angle frequency domain feature and the speed statistical feature are input into a long short-term memory network to extract the time sequence encoding vector of different rehabilitation actions; performing cluster analysis on the time series coding vectors to classify similar motion patterns into the same category, each category corresponding to a motion pattern prototype; storing the motion pattern prototypes and their corresponding joint angle change curves, velocity change curves, and duration distributions to construct the skeletal motion pattern memory bank.

7. The system of claim 1, wherein, The method comprises the following steps: combining the skeletal motion pattern memory bank and the bidirectional time series prediction algorithm to reconstruct the missing skeletal key point data for the potential tracking interruption risk frame, including: identifying the effective tracking frames before and after the potential tracking interruption risk frame, and extracting the skeletal motion features of the effective tracking frames; performing similarity matching between the skeletal motion features and the motion pattern prototypes in the skeletal motion pattern memory bank to retrieve the most similar historical motion pattern; based on the joint angle change curve of the historical motion pattern, constructing a forward time series prediction model to predict the skeletal key point positions of the interruption frame backward from the frame before the interruption; simultaneously constructing a backward time series prediction model to predict the skeletal key point positions of the interruption frame forward from the frame after the interruption; performing weighted fusion on the prediction results of the forward time series prediction model and the backward time series prediction model, and assigning the weights inversely proportional to the time distance from the prediction starting point to the interruption frame to generate the reconstructed skeletal key point data; 8. The system of claim 1, wherein, taking the measurement value of the auxiliary sensing feature vector at the interruption moment as a physical constraint to perform kinematic rationality verification and fine-tuning on the reconstructed skeletal key point data. The method comprises the following steps: segmenting the all-day skeletal motion sequence according to a preset time window, calculating the cumulative motion distance and joint range of motion of the skeletal key points in each segment; statistically calculating the duration of the patient in different postures in each time period, and identifying the posture types through a posture classifier; setting energy consumption coefficients for different posture types, and performing weighted summation on the cumulative motion distance, the joint range of motion, and the energy consumption coefficients to calculate the time-period activity amount index; extracting the action segment matching the preset rehabilitation action standard template from the all-day skeletal motion sequence, calculating the dynamic time warping distance between the joint angle trajectory of the action segment and the standard template; 9. The system of claim 3, wherein, calculating the rehabilitation action completion degree score according to the dynamic time warping distance, the duration deviation of action completion, and the joint activity amplitude deviation, and generating a personalized rehabilitation recommendation report. The method comprises the following steps: calculating the flexion angle, abduction angle, and rotation angle of the main joints according to the skeletal key point coordinates of the abnormal jump frame; querying a preset human joint activity range database to obtain the upper and lower limits of the physiological angles of each joint under different genders, ages, and rehabilitation stages; judging whether the flexion angle, abduction angle, and rotation angle are within the upper and lower limits of the physiological angles, and counting the number of joints exceeding the limits; Calculate the rate of change of the bone chain length of the abnormal jump frame and the frames before and after the frame, and determine that the bone structure is distorted when the rate of change of the bone chain length exceeds a preset deformation threshold; Frames in which the number of joints exceeding the upper and lower limits of the physiological angle is greater than a preset number threshold or the bone structure is distorted are confirmed as the error recognition candidate frames.

10. The system of claim 5, wherein, The spatial displacement correction of the skeletal key point coordinates in the visual skeletal tracking data based on the direction and amplitude of the correction vector includes: Decompose the correction vector into three components along the front-back axis, left-right axis, and up-down axis of the human anatomical coordinate system; Assign a correction sensitivity coefficient to each skeletal key point according to the skeletal key point confidence matrix, and the lower the confidence of a key point, the higher the correction sensitivity coefficient thereof; Multiply the three components of the correction vector by the correction sensitivity coefficient of the corresponding skeletal key point to obtain an individualized correction vector for each key point; Apply a bone chain constraint to the individualized correction vector to ensure that the distance between the corrected adjacent key points remains within a reasonable bone length range; Superimpose the individualized correction vector on the original skeletal key point coordinates to complete the spatial displacement correction and update the skeletal key point confidence matrix.

Citation Information

Patent Citations

  • Patient rehabilitation training data acquisition method and system based on visual identification

    CN119446391A

  • Multi-mode optical non-contact health monitoring system and method

    CN120392025A

  • Human motion capture and intelligent rehabilitation training evaluation method and system

    CN120544257A

  • Physical ability evaluation method and system based on human skeleton trajectory tracking

    CN120837061A

  • Orthopedic rehabilitation-oriented multi-modal data fusion and real-time monitoring method and system

    CN120932876A

Cited By

  • Clinical nursing aid decision-making method

    CN121460066A