Rehabilitation treatment monitoring system based on multi-modal sensing and ai evaluation

The rehabilitation monitoring system, which utilizes multimodal sensing and AI assessment, solves the problems of data interruption and misjudgment in complex environments of existing systems. It achieves stable monitoring and accurate rehabilitation assessment around the clock, supports the development of personalized treatment plans, and improves rehabilitation efficacy and patient compliance.

CN121237387BActive Publication Date: 2026-02-27LEDETANG (SHANGHAI) DIGITAL MEDICAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511796876.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-02-27
Estimated Expiration
2045-12-02

AI Technical Summary

Technical Problem

Existing rehabilitation monitoring systems are prone to data interruption or misjudgment in complex environments, leading to distorted rehabilitation assessments. They are unable to adapt to the unique movement patterns of different patient groups and lack effective data quality assessment and correction mechanisms, which affects the efficacy of rehabilitation and the formulation of treatment plans.

Method used

The rehabilitation monitoring system employs multimodal sensing and AI assessment. Through comprehensive analysis of visual skeletal tracking, inertial measurement units, and pressure distribution data, it performs confidence assessment, multimodal data fusion, and error compensation to generate a continuous and complete skeletal motion sequence, providing accurate rehabilitation assessment.

Benefits of technology

It improves the reliability and continuity of rehabilitation treatment monitoring, ensures stable monitoring around the clock, provides objective assessment of rehabilitation progress, supports the development of personalized treatment plans, and enhances patient compliance and treatment efficacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237387B_ABST
    Figure CN121237387B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of medical rehabilitation, and discloses a rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation; three types of sensing data, namely visual skeleton tracking, inertial measurement unit and pressure distribution, are integrated to establish a multi-dimensional skeleton key point confidence evaluation mechanism, and potential tracking interruption risk frames and error recognition candidate frames are recognized in real time. A multi-modal fusion correction model is used to accurately compensate for the error recognition, and the missing data is reconstructed by combining a skeleton motion mode memory bank and a bidirectional time sequence prediction algorithm to generate a continuous and complete all-day skeleton motion sequence. Based on this, the patient's time period activity index and rehabilitation action completion score are calculated to provide objective and quantitative rehabilitation evaluation basis. The application improves the stability and continuity of skeleton tracking in a complex rehabilitation environment, enables the medical team to obtain the real and complete all-day activity image of the patient, accurately locates the rehabilitation bottleneck, formulates a personalized treatment plan, and promotes the functional recovery of the patient.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical rehabilitation, more particularly, the present application relates to a rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation. BACKGROUND

[0002] The existing rehabilitation treatment monitoring system faces some technical challenges in actual clinical application. In a typical rehabilitation environment, the visual monitoring system frequently encounters various interference factors: partial obstruction of rehabilitation equipment, visual obstruction caused by the movement of medical staff, self-obstruction of the patient's body, and complex situations such as changes in indoor light over time, resulting in frequent interruption or incorrect identification of skeletal key point tracking. For example, when the patient is doing balance training, the support equipment often obstructs the lower limb key points, causing the system to incorrectly infer the joint position; when a stroke patient is doing upper limb rehabilitation training, due to slow or shaking movements, the system often misjudges normal but slow movements as tracking errors and discards the data. These problems are particularly prominent in long-term monitoring. The traditional visual tracking system has data interruption or serious distortion for some time in all-day monitoring, and lacks effective warning and repair mechanisms. When the interruption occurs, the system usually uses simple linear interpolation or directly discards abnormal data, which completely ignores the biomechanical characteristics and continuity rules of human movement. Especially when the patient is doing complex actions such as turning around and converting from sitting to lying down, the system has a high false alarm rate, resulting in a serious underestimate of the patient's activity ability. More importantly, these data loss and errors will be amplified in the all-day activity volume statistics, making the rehabilitation evaluation distorted: the system cannot distinguish between the patient's true rest and data loss, cannot accurately capture short but important activity peaks (such as independent standing attempts), and lacks a reliable correction mechanism for incorrect identification, ultimately causing the medical team to make treatment decisions based on incomplete or inaccurate data, directly affecting the rehabilitation efficacy evaluation and the development of individualized treatment plans. The existing system also generally lacks real-time quantitative evaluation of data quality, cannot take differentiated processing strategies for data of different confidence levels, and is difficult to adapt to the unique movement patterns and rehabilitation needs of different patient groups (such as the elderly, children, and different disease types), making the system less practical in complex and variable clinical environments.

[0003] In view of this, the present application proposes a rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation to solve the above problems. SUMMARY

[0004] In order to overcome the above-mentioned defects of the prior art, in order to achieve the above-mentioned purposes, the present application provides the following technical scheme: a rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation, comprising:

[0005] A data acquisition module for acquiring visual skeletal tracking data, inertial measurement unit data, and pressure distribution data during the patient's rehabilitation process;

[0006] a confidence evaluation module configured to perform multi-dimensional confidence evaluation on image features of each skeletal key point in the visual skeletal tracking data, and generate a skeletal key point confidence matrix;

[0007] an interruption risk identification module configured to identify potential tracking interruption risk frames and error identification candidate frames based on a time sequence change rate of the skeletal key point confidence matrix;

[0008] a multi-modal data fusion module configured to timestamp-align inertial measurement unit data and pressure distribution data, and construct an auxiliary sensing feature vector; and establish a multi-modal fusion correction model according to kinematic consistency constraints of the auxiliary sensing feature vector and the visual skeletal tracking data;

[0009] a reconstruction module configured to perform error compensation on skeletal key point coordinates in the error identification candidate frames based on the multi-modal fusion correction model; extract joint angle time sequence features and velocity time sequence features of the visual skeletal tracking data, and construct a skeletal motion pattern memory bank;

[0010] for the potential tracking interruption risk frames, the skeletal key point data missing is reconstructed in combination with the skeletal motion pattern memory bank and a bidirectional time sequence prediction algorithm;

[0011] the compensated skeletal key point coordinates and the reconstructed skeletal key point data are spliced to generate a continuous and complete all-day skeletal motion sequence;

[0012] a rehabilitation evaluation module configured to calculate time-periodic activity amount indexes and rehabilitation action completion degree scores of a patient based on the all-day skeletal motion sequence;

[0013] The above various modules are connected through wired and / or wireless modes to realize data transmission between the modules.

[0014] The rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation has the following technical effects and advantages:

[0015] The application improves the reliability and continuity of rehabilitation treatment monitoring. Through the innovative multi-dimensional monitoring mechanism, stable all-weather monitoring capability is maintained even in the case of partial obstruction of the patient by the rehabilitation device, interference by the medical staff walking around, or changes in light conditions, ensuring the integrity and consistency of rehabilitation evaluation data. This high robustness enables the clinical medical team to obtain a true and complete all-day activity profile of the patient, accurately distinguish between rest and activity periods, capture short but clinically significant functional motion attempts, and fully reflect the dynamic changes in rehabilitation progress. In terms of rehabilitation evaluation dimensions, the application provides a multi-level evaluation system including activity intensity, quality, distribution, and rehabilitation special motion completion, enabling the medical team to accurately locate rehabilitation bottlenecks and adjust treatment plans accordingly. For patients, the intuitive progress feedback provided by the application significantly improves rehabilitation compliance and confidence, replacing subjective feelings with objective and quantifiable results, and inspiring the motivation to continue participating in rehabilitation training. In clinical practice, the application reduces the observation and recording burden of medical staff, freeing up more time for treatment itself, while providing detailed data support to enable the development and adjustment of rehabilitation programs based on objective data rather than traditional experience-based judgments. The application optimizes the allocation of rehabilitation resources, shortens the patient rehabilitation period, and improves the quality of functional recovery. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 A rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation according to the present application. DETAILED DESCRIPTION

[0017] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application.

[0018] The present application provides a rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation. The execution subject of the system includes but is not limited to the following: rehabilitation medical devices, intelligent wearable terminals, rehabilitation training monitoring platforms, patient rehabilitation data analysis centers, and other general computing nodes of the present application. The AI evaluation system includes but is not limited to the following: cloud-based rehabilitation evaluation engines, distributed multi-modal data processing systems, and intelligent rehabilitation motion analyzers.

[0019] Please refer to Figure 1 In the embodiments of the present application, the rehabilitation treatment monitoring system based on multi-modal sensing and AI evaluation includes:

[0020] The data acquisition module is configured to acquire visual skeleton tracking data, inertial measurement unit data, and pressure distribution data during the rehabilitation process of a patient. The visual skeleton tracking data is collected by a depth camera or a general camera to obtain a sequence of bone key point coordinates during the movement of the patient. The inertial measurement unit data is recorded by a wearable sensor to record the acceleration, angular velocity, and attitude parameters of the joint movement of the patient. The pressure distribution data is recorded by a smart floor mat or insole to record the plantar pressure distribution of the patient during standing and walking. The three types of data complement each other to form a multi-modal comprehensive monitoring system and provide rich raw data for subsequent analysis.

[0021] The confidence evaluation module is configured to perform multi-dimensional confidence evaluation on the image features of each bone key point in the visual skeleton tracking data to generate a bone key point confidence matrix. The module analyzes the image stability features, skeletal structure consistency, time sequence stability, and visibility of the key points to comprehensively evaluate the reliability of each key point and form a quantitative confidence matrix. This confidence evaluation provides a key basis for subsequent error identification and compensation, ensuring that the system can accurately determine which key point data needs to be optimized.

[0022] The interruption risk identification module is configured to identify potential tracking interruption risk frames and error identification candidate frames based on the time sequence change rate of the bone key point confidence matrix. The module analyzes the time sequence attenuation rate of the confidence, the proportion of low-confidence key points, abnormal position jumps, and anatomical constraint verification to accurately identify frames that may have problems in the visual tracking system, providing targets for subsequent compensation and reconstruction. This risk identification mechanism can actively discover data quality problems and avoid the transmission of error data to subsequent analysis links.

[0023] The multi-modal data fusion module is configured to time-stamp align the inertial measurement unit data and the pressure distribution data to construct an auxiliary sensing feature vector, and to establish a multi-modal fusion correction model based on the kinematic consistency constraint between the auxiliary sensing feature vector and the visual skeleton tracking data. The module realizes accurate alignment of multi-source data through a time synchronization algorithm and establishes a mapping relationship between different modal data based on kinematic principles to provide a correction mechanism for errors in visual skeleton tracking. This fusion strategy effectively utilizes the complementarity of various sensing data to improve the robustness and accuracy of the system.

[0024] The reconstruction module is configured to perform error compensation on the skeletal key point coordinates in the error-identified candidate frame based on a multi-modal fusion correction model; extract joint angle time sequence features and speed time sequence features of the visual skeletal tracking data, and construct a skeletal motion pattern memory bank; for a frame with a potential tracking interruption risk, combine the skeletal motion pattern memory bank and a bidirectional time sequence prediction algorithm to reconstruct the missing skeletal key point data; and splice the compensated skeletal key point coordinates and the reconstructed skeletal key point data to generate a continuous and complete all-day skeletal motion sequence. The module repairs the identified problem data through various intelligent algorithms, ensures the continuity and integrity of the skeletal motion sequence, and provides high-quality basic data for rehabilitation assessment.

[0025] The rehabilitation assessment module is configured to calculate the time-period activity index and rehabilitation action completion score of the patient based on the all-day skeletal motion sequence. The module quantifies the activity amount and action quality in the rehabilitation process by analyzing the all-day motion data of the patient, provides objective and quantitative rehabilitation progress assessment for medical personnel, and supports optimization and adjustment of the rehabilitation scheme.

[0026] The above various modules are connected through wired and / or wireless means to realize data transmission between the modules.

[0027] In the embodiment of the application, the image features of each skeletal key point in the visual skeletal tracking data are subjected to multi-dimensional confidence evaluation to generate a skeletal key point confidence matrix, and the detailed implementation steps include:

[0028] The local gradient intensity, edge response value and texture contrast of each skeletal key point in the image are extracted as the image stability features of the key point. The image stability features are basic visual parameters for evaluating the reliability of the key point, and reflect the prominence and stability of the key point in the image. The extraction process first sets a fixed size pixel region (usually 16x16 pixels) around the key point, and then calculates the image features of the region: the local gradient intensity is calculated by the Sobel operator to calculate the average value of the pixel gradient amplitude, reflecting the edge intensity; the edge response value is calculated by the Harris corner response function, reflecting the corner characteristics; and the texture contrast is calculated by the contrast feature of the local region gray level co-occurrence matrix, reflecting the texture richness. The three features form the image stability vector of the key point, and the higher the value, the more stable and reliable the key point in the image, and the easier it is to be accurately identified.

[0029] The bone length proportion between the skeleton key points and the adjacent key points is calculated, and the bone length proportion is compared with the preset human bone anatomy proportion to obtain a bone structure consistency score. The bone structure consistency score is based on human anatomy knowledge to evaluate whether the key point position conforms to the normal human structure. The calculation process first establishes the connection relationship between the skeleton key points, defines the adjacent key point pair; then calculates the distance between each pair of adjacent key points to obtain the bone length; calculates the ratio of each bone length to obtain the bone length proportion; finally, the actual proportion is compared with the standard human bone proportion to calculate the deviation degree. The bone structure consistency score is calculated by the following formula:

[0030] ;

[0031] wherein, is the bone structure consistency score, is the actual bone length proportion, is the preset human bone anatomy proportion, is the number of proportion pairs, is the scaling factor (usually 3-5).

[0032] The score range is [0, 1], and the higher the value, the more the bone structure conforms to the human anatomy characteristics, and the more reliable the key point position. This scoring mechanism effectively identifies incorrect identification that violates the human anatomy structure, such as abnormal elongation or compression of the limbs.

[0033] The position jitter amplitude of the skeleton key points in the continuous frames is counted, the deviation between the predicted position and the actual detected position is predicted by Kalman filtering, and the time sequence stability score of the key points is calculated. The time sequence stability score evaluates the stability of the key points in the time dimension, which is particularly important for identifying jumps and jitter. The evaluation process first establishes a Kalman filter model for each key point, predicts the theoretical position of the current frame based on the historical position and speed; then calculates the Euclidean distance between the predicted position and the actual detected position to obtain the prediction deviation; finally, the time sequence stability score is calculated based on the prediction deviation. The calculation formula is:

[0034] ;

[0035] wherein, TS is the time sequence stability score, d is the Euclidean distance between the predicted position and the actual position, and σ is the standard deviation parameter (usually dynamically set according to the type of the skeleton key point, the core key point such as the trunk is smaller, and the end of the limbs is larger).

[0036] The score range is (0, 1], and the closer the value to 1, the higher the time sequence stability, and the smaller the jitter and jump. This scoring mechanism is particularly sensitive to sudden abnormal motion and can effectively filter rapid jumps that are not reasonable in motion physiology.

[0037] The occlusion probability estimation is performed on the pixel region around the skeleton key point, and the visibility score of the key point is obtained by analyzing the continuity of the color histogram and the depth gradient mutation. The visibility score evaluates whether the key point is occluded by other objects, which is an important dimension for judging the reliability of the data. Different strategies are adopted for depth cameras and ordinary cameras in the evaluation process: for depth cameras, the depth value distribution around the key point is analyzed to detect the mutation area of the depth gradient and identify the possible occlusion boundary; for ordinary cameras, the color histogram of the region around the key point in consecutive frames is analyzed to detect the sudden histogram discontinuity and identify the possible occlusion event. The final visibility score is calculated by combining the two methods, reflecting the complete visibility of the key point. The score range is [0, 1], and the higher the value, the better the visibility and the lower the occlusion degree. This scoring mechanism is particularly effective for detecting partial and complete occlusion, providing an important basis for subsequent data compensation.

[0038] The image stability feature, the skeleton structure consistency score, the time sequence stability score and the visibility score are weighted and fused to generate a skeleton key point confidence matrix. The confidence matrix is a comprehensive expression of multi-dimensional evaluation results, which directly reflects the reliability level of each key point in each frame. The fusion process adopts a weighted average method, and the weights are dynamically adjusted according to the key point type and application scenario: the core joint (such as the hip joint and the shoulder joint) is usually given a higher weight of the skeleton structure consistency; the end joint with large motion amplitude (such as the wrist and the ankle) is given a higher weight of the time sequence stability; the joint that is easily occluded is given a higher weight of the visibility score. The weighted fusion formula is:

[0039] ;

[0040] , is the normalized value of the image stability feature, is the skeleton structure consistency score, is the time sequence stability score, is the visibility score, , , , is the weight coefficient, and .

[0041] The value range of each element of the confidence matrix is [0, 1], and the higher the value, the more reliable the key point. This matrix provides an accurate quantitative evaluation basis for subsequent interruption risk identification and error compensation, and is the core quality control mechanism of the system.

[0042] In the embodiment of the application, based on the time sequence change rate of the skeleton key point confidence matrix, the detailed implementation steps for identifying potential tracking interruption risk frames and error identification candidate frames include:​

[0043] The confidence drop gradient of the computed bone key point confidence matrix between adjacent frames is denoted as the confidence decay rate. The confidence decay rate is a key indicator for detecting tracking quality mutations, quantifying the temporal trend of confidence. The calculation process first obtains the confidence values of the corresponding key points in the adjacent two frames, then calculates the difference value and divides it by the inter-frame time interval to obtain the confidence change rate per unit time. For a system with stable frame rate, the confidence difference value can be directly calculated. A negative confidence decay rate indicates a decrease in confidence, and the larger the absolute value, the more dramatic the decrease and the higher the potential risk. The system calculates the confidence decay rate for each key point separately, and different alert thresholds can be set according to the importance of the key points. The alert threshold for core joints is usually lower and more sensitive to quality changes.

[0044] The proportion of key points with confidence below the pre-set confidence threshold in a single frame is calculated and denoted as the intra-frame low confidence proportion. The intra-frame low confidence proportion is a comprehensive indicator for evaluating the quality of the entire frame, reflecting the overall level of tracking quality in the current frame. The calculation process first sets a pre-set confidence threshold (usually 0.4-0.6, adjusted according to the application scenario), then counts the number of key points in the current frame with confidence below the threshold, and finally divides the total number of key points to obtain the proportion value. The higher the proportion value, the worse the quality of the current frame, and the more likely there is a risk of tracking interruption. The system can dynamically adjust the threshold and evaluation strategy according to different skeletal models (such as the full-body 25-point model, the upper-body 15-point model, etc.), ensuring the relevance and accuracy of risk identification.

[0045] Frames with a confidence decay rate greater than the pre-set decay threshold and an intra-frame low confidence proportion greater than the pre-set proportion threshold are marked as potential tracking interruption risk frames. This step considers both the temporal changes and the current state, accurately identifying risk frames that may soon experience tracking interruption. The marking process uses a double-threshold decision strategy: the pre-set decay threshold is usually set to -0.1 to -0.2, indicating a 10%-20% decrease in single-frame confidence; the pre-set proportion threshold is usually set to 0.3-0.4, indicating that 30%-40% of the key points are in a low confidence state. When both conditions are met, the system determines that the current frame has a high risk of tracking interruption and initiates the defense mechanism in advance to prepare for subsequent data reconstruction. This early warning mechanism significantly improves the system's ability to respond to tracking interruptions and reduces the impact of data loss.

[0046] Detect the position jump distance of the skeleton key points between consecutive frames, and when the position jump distance exceeds the reasonable motion range calculated based on the historical motion speed, mark it as an abnormal jump frame. Abnormal jump detection is a key step to find sudden error recognition. Based on the continuity principle of human motion, identify physically unreasonable position changes. The detection process first calculates the displacement distance of each key point between adjacent frames, and then establishes a reasonable motion range model for the key point based on the average speed and standard deviation of the last N frames (usually 5-10 frames):

[0047] Reasonable motion range ;

[0048] wherein, is the historical average speed, is the speed standard deviation, is the inter-frame time interval, is the expansion coefficient (usually 2.5-3.5).

[0049] When the actual displacement exceeds this range, the system determines that it is an abnormal jump. This adaptive historical speed-based determination method can adapt to the motion characteristics of different patients and the speed differences of different rehabilitation movements, improving the accuracy and adaptability of abnormal detection.

[0050] Anatomical constraint verification is performed on the skeleton key points in the abnormal jump frame to calculate whether the joint angle exceeds the human physiological activity range, and the frame that violates the physiological activity range is marked as an error recognition candidate frame. Anatomical constraint verification is the final determination step to confirm whether the abnormal jump is an error recognition. Based on the medical knowledge of human joint activity range, it identifies physiologically impossible postures. The verification process first calculates the three-dimensional angles of the main joints (shoulder, elbow, wrist, hip, knee, ankle, etc.) according to the key point coordinates, and then queries the pre-set human joint activity range database to determine whether the joint angles are within a reasonable range. The activity range in the database is personalized according to the patient's age, gender and rehabilitation stage to ensure the adaptability of the determination standard. When the number of joints exceeding the reasonable range is greater than the pre-set threshold, or when the length of the skeleton changes significantly, the system confirms that it is an error recognition candidate frame and marks it as a target that needs error compensation. This biomechanics and anatomy-based verification mechanism greatly improves the reliability of error recognition determination and avoids false intervention on normal but large amplitude movements.

[0051] In the embodiment of the application, the inertial measurement unit data and the pressure distribution data are time-stamped and aligned, and the detailed implementation steps of constructing the auxiliary sensing feature vector include:

[0052] The sampling timestamps of the inertial measurement unit data and the frame timestamps of the visual skeleton tracking data are cross-correlated to identify a system latency offset. The latency offset is a key parameter for accurate alignment of multi-modal data, quantifying the time difference between different sensing systems. The analysis process is based on cross-correlation techniques in signal processing, finding the time offset value corresponding to the maximum similarity by calculating the similarity of two time series signals with respect to time offset. In specific implementation, explicit movement events of the patient (such as standing, turning, lifting arms, etc.) are selected as feature points, and the corresponding feature responses in each modal data are extracted, the time difference between the responses is calculated, and the stable system latency offset is obtained by averaging multiple samples. This offset is usually in the order of milliseconds, but it is crucial for accurate motion analysis, especially in high-speed motion and fine motion evaluation scenarios.

[0053] According to the system latency offset, the inertial measurement unit data and the pressure distribution data are timestamp corrected to realize synchronization with the visual skeleton tracking data. Timestamp correction is a preprocessing step for data fusion, ensuring that data from different sources are aligned in the time dimension, providing a basis for subsequent feature extraction. The correction process uses a linear mapping method to adjust the timestamps of the inertial measurement unit and pressure distribution data according to the system latency offset, and unifies them to the time reference of the visual skeleton tracking data. For data sources with different sampling frequencies, interpolation or downsampling techniques are used to match the reference frequency, ensuring one-to-one correspondence of data points. The corrected data meets the time synchronization requirement and can accurately reflect the multi-dimensional state information of the patient at the same time, providing reliable data for kinematic consistency analysis.

[0054] From the inertial measurement unit data, three-axis acceleration peak, angular velocity rate of change and attitude quaternion are extracted, and the center of pressure trajectory calculation and gait phase identification are performed on the pressure distribution data. This step is a key process for extracting high-value features from raw sensor data, converting low-level sensing signals into biomechanically meaningful feature parameters. For inertial measurement unit data, three-axis acceleration peak reflects motion intensity, angular velocity rate of change reflects motion conversion characteristics, and attitude quaternion reflects spatial direction; for pressure distribution data, center of pressure trajectory calculation reflects center of gravity changes, and gait phase identification (such as ground contact period, swing period) reflects walking cycle characteristics. These extracted feature parameters have clear biomechanical interpretation and can be directly used for subsequent kinematic consistency analysis and data correction, improving the utilization efficiency of auxiliary sensing data.

[0055] The triaxial acceleration peak value, angular velocity change rate, attitude quaternion, foot pressure center trajectory and gait phase are spliced according to a time window to generate an auxiliary sensing feature vector. Feature splicing is the last step of forming a comprehensive auxiliary sensing feature, which combines features of different dimensions and physical meanings into a unified feature vector, facilitating subsequent model processing. The splicing process first sets a fixed size time window (usually 0.5-2 seconds, adjusted according to motion characteristics), then calculates the statistics (such as mean, standard deviation, peak value, etc.) of each feature in each window, and finally combines these statistics in a predetermined order to form a fixed-dimensional feature vector. For different types of features, appropriate normalization processing is used to ensure dimensional consistency and avoid a dimension feature dominating the model behavior due to a large value range. The generated auxiliary sensing feature vector contains rich kinematic information and has a structured and standardized characteristic, providing ideal input data for subsequent multi-modal fusion correction models.

[0056] In the embodiments of the present application, the detailed implementation steps of establishing a multi-modal fusion correction model according to the kinematic consistency constraint between the auxiliary sensing feature vector and the visual skeleton tracking data include:

[0057] The joint angular velocity vector and the center of mass acceleration vector are calculated from the visual skeleton tracking data as visual kinematic features. Visual kinematic features are high-level representations of skeletal motion, reflecting the core dynamics of motion. The calculation process first obtains velocity information based on the difference in key point coordinates between adjacent frames, then calculates the joint angle through the skeletal connection relationship, and further differentiates to obtain the angular velocity; the center of mass acceleration is obtained by weighted averaging the accelerations of each key point, and the weight is determined according to the mass distribution of each part of the human body. These features directly reflect the mechanical state of human motion and have a clear correspondence with the physical quantities measured by the inertial measurement unit and the pressure sensor, providing a basis for establishing kinematic mapping. The extraction of visual kinematic features reduces the dimensionality of the original skeletal coordinate data while retaining the essential characteristics of motion, facilitating consistency analysis with other modal sensing data.

[0058] A kinematic mapping relationship between the visual kinematic features and the auxiliary sensing feature vector is established, and a consistency loss function is constructed by minimizing the feature difference between the two at the same time. Kinematic mapping is the core mechanism of multi-modal fusion, which unifies the data measured by different sensors into a common kinematic interpretation framework based on physical principles. The mapping relationship is established using a deep neural network model, with the auxiliary sensing feature vector as input and the predicted visual kinematic features as output. The consistency loss function is designed in two parts: the main part is the mean square error between the predicted features and the actual features, reflecting the overall fitting accuracy; the regularization part introduces kinematic constraint conditions such as the continuity of angle change, the coordination of center of mass acceleration and joint force, etc. The complete loss function form is:

[0059] ;

[0060] wherein, is the consistency loss function value, is the visual kinematics feature, is the visual feature predicted based on the auxiliary sensor feature, is the kinematics constraint regularization term, is the auxiliary sensor feature vector, is the regularization coefficient.

[0061] This loss function not only considers data fitting accuracy, but also incorporates kinematics knowledge constraints, ensuring that the mapping relationship learned by the model conforms to human biomechanics principles, improving the physical rationality of prediction.

[0062] Collect multi-modal data of patients under standard rehabilitation actions, train deep neural networks to learn the optimal weight parameters of the consistency loss function. Model training is the key process to establish accurate mapping relationship, through a large number of real data to optimize model parameters, so that it can accurately predict visual features. Training data collection covers a variety of standard rehabilitation actions (such as joint flexion and extension, gait training, balance exercises, etc.), ensuring the generalization ability of the model; at the same time, special data sets are collected for different patient groups (such as different age groups, different rehabilitation stages), to improve the adaptability of the model to specific groups. The network structure adopts an encoder-decoder architecture, with attention mechanism added in the middle to enhance the influence of key features, and residual connection introduced to improve gradient propagation. The training process uses the mini-batch stochastic gradient descent method, combined with learning rate decay and early stopping strategy, to prevent overfitting and improve convergence speed. The best hyperparameters are determined through cross-validation to ensure that the model achieves optimal performance on the validation set, providing reliable protection for subsequent correction applications.

[0063] The prediction error of the deep neural network on newly collected data is used as a correction feedback signal, and when there is kinematic inconsistency between the visual skeleton tracking data and the auxiliary sensor feature vector, a correction vector is generated. Correction feedback is the key mechanism to identify visual tracking errors, based on the consistency principle between multi-modal data, to find potential problems and provide correction information. The workflow first inputs the auxiliary sensor feature vector of the current frame, predicts the theoretical visual kinematics feature through the trained network; then compares it with the feature calculated from the actual visual skeleton tracking data to calculate the difference; when the difference exceeds the preset threshold, the system determines that there is inconsistency, further analyzes the direction and amplitude of the difference, and generates a correction vector. The correction vector contains the direction and size information that needs to be adjusted, directly guiding the subsequent coordinate correction process. This correction mechanism based on prediction error effectively utilizes the complementarity of multi-modal data, can identify problems that single modal cannot find, and improves the robustness and accuracy of the system.

[0064] Based on the direction and amplitude of the correction vector, the spatial displacement correction is performed on the skeletal key point coordinates in the visual skeleton tracking data, and a multi-modal fusion correction model is established. The spatial displacement correction is a specific step of performing correction, which converts the correction information at the high-level feature level into specific coordinate adjustment operations. The correction process first maps the correction vector from the feature space back to the coordinate space to determine the direction and distance that each key point needs to adjust; then the correction sensitivity is allocated according to the confidence of the key point, and the key point with low confidence obtains a larger adjustment amplitude; finally, the coordinate correction is performed while maintaining the skeletal length constraint and joint range of motion constraint to ensure the biomechanical rationality of the correction result. The whole correction process forms a closed-loop feedback system: visual tracking provides initial estimation, auxiliary sensing provides verification reference, when the two are inconsistent, adjustment is made through the correction model, and finally more accurate and reliable skeletal key point coordinates are obtained. This multi-modal fusion correction model improves the performance of the system in complex environments and effectively deals with common challenges such as visual occlusion and light changes.

[0065] In the embodiment of the application, the detailed implementation steps of error compensation for the skeletal key point coordinates in the error recognition candidate frame based on the multi-modal fusion correction model include:

[0066] The N key points with the lowest confidence are extracted from the error recognition candidate frame as the main error key points, wherein N is adaptively determined according to the low confidence ratio in the frame. The identification of the main error key points is a key step to accurately locate the problem source, avoiding unnecessary adjustment of the overall skeletal structure. The extraction process first sorts all key points in the current frame according to the confidence matrix to identify the candidate set with the lowest confidence; then the value of N is dynamically determined based on the low confidence ratio in the frame, and the higher the ratio, the more common the problem and the more key points need to be processed. The calculation formula of N value is:

[0067] ;

[0068] wherein, is the total number of key points in the skeletal model, is the low confidence ratio in the frame, is the adjustment coefficient (usually 1.2-1.5).

[0069] This adaptive determination strategy ensures that the system can adjust the intervention range according to the severity of the problem, neither missing critical errors nor over-intervening correctly identified parts, maintaining the accuracy and efficiency of the compensation process.

[0070] The main error key points are input into the multi-modal fusion correction model to obtain the theoretical coordinates calculated based on the auxiliary sensor feature vectors. Theoretical coordinate calculation is the basis for error compensation, providing a reference position based on multi-modal information fusion. The calculation process uses the multi-modal fusion correction model established above, with the auxiliary sensor feature vectors of the current frame as input, to predict the visual kinematic features under ideal conditions. Then, through reverse mapping, the predicted kinematic features are converted into the spatial coordinates of the key points. This calculation makes full use of the data from the inertial sensor and the pressure sensor, providing an alternative position estimate when visual tracking is unreliable. The theoretical coordinates have strong physical consistency, and even in the case of complete visual failure, they can still provide reasonable position estimates based on the actual motion state of the patient, providing a reliable reference for subsequent compensation.

[0071] The spatial deviation vector between the detected coordinates and the theoretical coordinates of the main error key points is calculated, and the outliers with a deviation amplitude exceeding three standard deviations are removed, leaving the effective deviation vector. Deviation analysis is a key step in quantifying error levels, identifying specific directions and amplitudes that need to be compensated. The analysis process first calculates the three-dimensional vector difference between the actual detected coordinates and the theoretically predicted coordinates of each error key point, obtaining the original deviation vector. Then, statistical analysis is performed on the deviation vectors of all key points to calculate the mean and standard deviation of the deviation distribution. Finally, based on the "three standard deviation" criterion, abnormal deviations are identified and removed. This statistical screening mechanism avoids over-reaction to extreme errors, improving the stability and reliability of the compensation process. The retained effective deviation vector accurately reflects the direction and amplitude that need to be compensated, providing a quantitative basis for subsequent precise adjustment.

[0072] According to the hierarchical topological structure of the human skeleton, the main error key points are divided into root node key points and end node key points, and the root node key points are compensated first. The hierarchical compensation strategy is a key technology that considers the structure of the skeleton, avoiding the structural inconsistency caused by simple independent adjustment. The classification process is based on the hierarchical structure of the human skeleton, dividing the key points into root nodes (such as the torso, pelvis, scapula, etc. core parts) and end nodes (such as the wrist, ankle, etc. limb ends). The root node is located upstream of the skeletal chain, and its position error will be transmitted and amplified to the downstream nodes, so prioritizing the root node can solve the problem of chain errors and improve the overall compensation effect. This hierarchical processing strategy based on topological structure conforms to the biomechanical principles of human motion, ensuring the structural rationality and overall coordination of the compensation results.

[0073] For the root node key point, the inverse of the effective deviation vector is processed by a time series smoothing filter to generate a time series consistency compensation vector and superimposed on the original coordinates. Root node compensation is the first step of error correction, directly affecting the accuracy of the entire bone chain. The compensation process first takes the inverse of the effective deviation vector as the basis for compensation direction, indicating the need to move in the direction of the theoretical coordinates; then the compensation vectors of consecutive frames are smoothed by a time series smoothing filter (such as an exponential moving average filter) to reduce the abruptness of the compensation action; finally, the smoothed compensation vector is superimposed on the original coordinates to complete the position correction of the root node. Time series smoothing ensures the continuity of the compensation process, avoiding visual jumps caused by single-frame drastic adjustments, making the corrected bone movement more natural and smooth. The accurate compensation of the root node lays the foundation for the adjustment of the subsequent end node and is the key link of the overall compensation effect.

[0074] Based on the compensated root node coordinates, the inverse kinematics propagation algorithm is used to calculate the compensation coordinates of the end node key points. The inverse kinematics propagation algorithm constrains the constant length of the bone and the joint angles to meet the physiological range. Inverse kinematics propagation is the core technology for processing end nodes, based on the connection relationship of the bone chain to calculate a series of reasonable positions of the joints. The algorithm process first establishes a bone chain model from the root node to the end node, defines the degrees of freedom and activity range of each joint; then sets the target position of the end node, i.e. the theoretical coordinate position; finally, through iterative optimization method, under the premise of keeping the bone length unchanged, adjusts the joint angles to make the end node approach the target position, while ensuring that all joint angles are within the physiological activity range. When the target position requirements cannot be fully met, the algorithm will find the optimal solution under the constraint conditions to ensure the biomechanical reasonableness of the results. This inverse kinematics-based propagation method considers the integrity and connection constraints of the human body skeleton, and the compensation results are accurate and natural, avoiding the bone deformation problem that may be caused by simple interpolation.

[0075] The prediction error of the compensated coordinates is verified by the multi-modal fusion correction model, and when the prediction error is greater than a preset error threshold, the filtering parameters and constraint weights are iteratively adjusted until the threshold requirement is met. Verification and iterative optimization are the guarantee mechanism to ensure the compensation quality, and the compensation accuracy is continuously improved through closed-loop feedback. The verified and compensated coordinates are re-input into the multi-modal fusion correction model, the consistency degree with the auxiliary sensing feature vector is calculated, and the compensation effect is evaluated; when the prediction error exceeds the preset threshold, the system enters the iterative optimization stage, adjusts the compensation parameters and re-executes the compensation process. The parameters iteratively adjusted include the smoothing coefficient of the time series smoothing filter, the weight coefficient of various constraints in inverse kinematics, etc., which are optimized according to the nature of the prediction error. The iteration termination condition is that the prediction error is lower than the threshold or the maximum iteration number is reached. This iterative optimization mechanism based on model verification ensures the reliability and accuracy of the compensation process, and even in complex and variable rehabilitation scenarios, it can also produce high-quality compensation results.

[0076] In the embodiment of the application, the detailed implementation steps for constructing the skeleton motion pattern memory library include:

[0077] The three-dimensional angle sequence of the main joints in the visual skeleton tracking data is calculated, and the discrete Fourier transform is performed on the three-dimensional angle sequence to extract the main frequency component and the harmonic component in the frequency domain as the joint angle frequency domain feature. Frequency domain feature extraction is a key technology for capturing periodic motion patterns, which converts time domain signals into frequency components, facilitating the identification of basic motion patterns. The calculation process first calculates the three-dimensional angle time series data of the main joints (such as shoulders, elbows, wrists, hips, knees, ankles, etc.) from the skeleton key point coordinates; then applies fast Fourier transform (FFT) to each angle sequence to convert the time domain signal to the frequency domain; finally, the main features are extracted from the frequency spectrum, including the main frequency (reflecting the basic motion period), the main frequency amplitude (reflecting the main motion intensity), and the first N harmonic components (reflecting the complexity and morphology of the motion). These frequency domain features effectively capture the periodicity and rhythm of rehabilitation movements, which are particularly important for identifying standard rehabilitation movements and evaluating movement quality. Frequency domain representation also has scale invariance and time shift invariance, so that pattern recognition is not affected by speed changes and starting point differences, improving the generalization ability of the features.

[0078] The position coordinates of the skeletal key points are time-differentiated to obtain the instantaneous speed of the key points, and the mean, variance and peak value of the instantaneous speed are counted as speed statistical features. The speed statistical features are important indicators for describing the kinetic characteristics of the motion, and reflect the fluency and controllability of the rehabilitation motion. The calculation process first differentiates the position coordinates of the skeletal key points (usually using central difference method to improve accuracy) to obtain the instantaneous speed of each key point in each direction; then calculates the statistical features of the speed, including mean (reflecting the overall speed level), variance (reflecting the speed stability), peak value (reflecting the maximum explosive force), peak-to-valley ratio (reflecting the acceleration and deceleration characteristics), etc.; finally, the statistical features of each key point are combined into a unified speed feature vector. These speed features are directly related to the patient's motor control ability, and can effectively distinguish between healthy and pathological motions, which is of great significance for evaluating rehabilitation progress and identifying potential risks. The speed features and angle features complement each other and together constitute a comprehensive motion pattern description, providing a rich feature basis for subsequent analysis.

[0079] The joint angle frequency domain features and speed statistical features are input into a long short-term memory network to extract time sequence encoding vectors of different rehabilitation motions. Time sequence encoding is a key step to compress high-dimensional time sequence features into low-dimensional representations, capturing the essential features of the motion and facilitating subsequent storage and matching. The encoding process uses a long short-term memory (LSTM) network architecture, which is particularly suitable for processing time sequence data and can learn long-term dependencies. The network input is a time sequence feature sequence, including joint angle frequency domain features and speed statistical features; the network output is a fixed-dimensional encoding vector (usually 64-128 dimensions), representing the abstract expression of the motion. The network training uses an autoencoder structure to optimize the encoding quality by reconstructing the original sequence, ensuring that the encoding vector retains the key information of the original motion. The time sequence encoding vector has high information density and discriminability, and can effectively distinguish between different types of rehabilitation motions and different quality levels of execution, providing an ideal feature representation for subsequent pattern classification and matching.

[0080] The time series encoding vectors are subjected to cluster analysis, and similar motion patterns are classified into the same category, and each category corresponds to a motion pattern prototype. The cluster analysis is a key step for identifying typical motion patterns, which discretizes the continuous feature space into meaningful pattern categories. The analysis process adopts an improved K-means++ algorithm, and the optimal clustering result is found by minimizing the intra-class distance and maximizing the inter-class distance. The number of clusters is automatically determined by evaluation indexes such as the silhouette coefficient, and usually corresponds to the number of common rehabilitation action types. Each cluster center is defined as a motion pattern prototype, representing the standard form of the action in this category. In addition to the cluster center, the system also calculates the covariance matrix and boundary examples of each category, describing the variation range and boundary characteristics of the pattern in this category. This data-driven pattern discovery method can automatically identify typical action patterns in the patient's rehabilitation process without predefining all possible action types, and has strong adaptability and scalability.

[0081] The motion pattern prototypes and their corresponding joint angle change curves, velocity change curves and duration distributions are stored to construct a skeletal motion pattern memory library. Memory library construction is a key step to realize pattern memory and reuse, which systematically stores the discovered patterns and establishes a retrieval mechanism. The construction process first assigns a unique identifier to each motion pattern prototype and establishes an index structure; then stores the detailed features of the pattern, including the time series encoding vector, the original joint angle curve, the velocity change curve and the duration distribution; finally, an efficient similarity calculation and retrieval mechanism is established to support subsequent pattern matching operations. The memory library adopts a hierarchical storage structure, which organizes patterns by function category (such as upper limb training, lower limb training, balance training, etc.) and complexity level, facilitating precise positioning of target patterns. For each pattern, the memory library also records its adaptability and variation in different patient groups, providing a reference for subsequent personalized matching. The skeletal motion pattern memory library is the core knowledge base of the system, and with continuous expansion and optimization of system use, it continuously improves the recognition and reconstruction ability of various rehabilitation actions.

[0082] In the embodiment of the application, for the potential tracking interruption risk frame, the detailed implementation steps for reconstructing the missing skeletal key point data include:

[0083] Identify the effective tracking frames before and after the potential tracking interruption risk frame, and extract the skeletal motion features of the effective tracking frames. Effective frame identification is a prerequisite step for data reconstruction, determining a reliable reference data range and providing a basis for subsequent prediction. The identification process first filters the frames with higher confidence in the time window (usually ±2 seconds) before and after the potential tracking interruption risk frame based on the confidence matrix, as candidate effective frames; then applies continuity verification to the candidate effective frames to ensure that the skeletal motion between frames meets the physical continuity principle and eliminates possible abnormal frames; finally, from the verified frames, select N frames (usually 10-20 frames, adjusted according to the complexity of the action) before and after the interruption as the final effective tracking frames. Extract skeletal motion features from these effective frames, including joint angle sequences, velocity sequences, and acceleration sequences, to provide input data for pattern matching and bidirectional prediction. Accurate identification of effective tracking frames ensures the quality of reference data for the reconstruction process and is a key prerequisite for high-quality reconstruction.

[0084] Match the skeletal motion features with the motion pattern prototypes in the skeletal motion pattern memory library to retrieve the most similar historical motion patterns. Pattern matching is a key step in determining the current motion type, providing prior knowledge guidance for subsequent prediction. The matching process first converts the motion features of the effective tracking frames into time sequence encoding vectors through the same method as the memory library construction; then calculates the similarity of the encoding vector and each motion pattern prototype in the memory library, using cosine similarity or Mahalanobis distance measurement; finally, select the K patterns with the highest similarity (usually K=3-5) as candidate historical motion patterns. For these K candidate patterns, the system further considers the matching degree of their applicable background (such as patient type, rehabilitation stage) with the current situation to determine the final most similar historical motion pattern. This content-based similarity retrieval method can accurately identify the type of rehabilitation action being performed, providing targeted prior knowledge for subsequent prediction and greatly improving the accuracy and specificity of reconstruction.

[0085] Based on the historical motion pattern of joint angle change curve, a forward time series prediction model is constructed to predict the skeletal key point position of the interrupted frame from the frame before interruption. Forward prediction is a process of inferring from the past to the future, which predicts the state during interruption based on the motion trend before interruption. The prediction model adopts a hybrid strategy of combining autoregressive moving average (ARMA) with historical patterns: on the one hand, the ARMA model is used to capture the short-term dynamic trend, and an autoregressive model is established based on the time series data of the last N frames before interruption; on the other hand, the prior knowledge of historical motion pattern is integrated, especially the change trend and characteristics that should be in the current stage. The two parts of prediction are fused by weighting to form the final forward prediction result, and the weight is dynamically determined according to the fitting degree of the autoregressive model and the matching degree of the historical pattern. Forward prediction directly generates joint angle prediction of the interrupted frame, and then converts it into skeletal key point coordinates through forward kinematics to complete the preliminary position prediction. This hybrid prediction strategy combining data-driven and knowledge-guided can effectively balance short-term trend and long-term pattern, and improve the accuracy and robustness of prediction.

[0086] At the same time, a backward time series prediction model is constructed to predict the skeletal key point position of the interrupted frame from the frame after interruption. Backward prediction is a reverse inference from the future to the past, which inversely deduces the change during interruption based on the state after interruption. The prediction process is similar to forward prediction, but the processing direction is opposite, and a backward autoregressive model is established using the effective frames after interruption, while considering the inverse evolution rule of historical motion pattern. Backward prediction pays special attention to the final state constraint, that is, to ensure that the prediction result can smoothly transition to the known state after interruption, avoiding the mutation between prediction and reality. Backward prediction also generates joint angle prediction and converts it into skeletal key point coordinates to form position estimation from another time direction. Forward prediction and backward prediction are independent and complementary to each other, and the accuracy of overall prediction is improved through bidirectional constraint, especially the processing capability for longer interruption interval is significantly enhanced.

[0087] The prediction results of forward time series prediction model and backward time series prediction model are weighted and fused, and the weight is inversely proportional to the time distance between the prediction starting point and the interrupted frame, to generate the reconstructed skeletal key point data. Weighted fusion is a key step to integrate the advantages of bidirectional prediction, and the influence of different direction prediction is adaptively adjusted according to the time distance. In the fusion process, for each frame in the interruption interval, the time distance from the last effective frame before interruption and the first effective frame after interruption is calculated respectively, and the closer the distance, the greater the prediction weight of the direction. The weight calculation formula is:

[0088] ;

[0089] ;

[0090] wherein, and weights for forward and backward prediction, respectively, and are the time distances to the start of forward and backward prediction, respectively.

[0091] The weighted fusion is performed in both joint angle space and coordinate space, ensuring consistency and smoothness of the results in both spaces. The fused skeletal key point data not only considers the continuity and inertia of motion, but also meets the boundary constraints of the start and end points, forming a natural and smooth transition sequence, effectively eliminating the data discontinuity problem caused by interruptions.

[0092] The measurement value of the auxiliary sensing feature vector at the interruption time is taken as a physical constraint to perform kinematic rationality verification and fine-tuning on the reconstructed skeletal key point data. Physical constraint verification is the last guarantee to ensure the quality of reconstruction, and cross-verification is performed using the complementarity of multi-modal data. The verification process first calculates kinematic features such as joint angular velocity and center of mass acceleration from the reconstructed skeletal key points; then compares them with the actual measurement values of the auxiliary sensing feature vector at the same time to calculate the degree of kinematic consistency; when the consistency is lower than the preset threshold, the system starts the fine-tuning process, which adjusts the reconstruction result through iterative optimization until the kinematic constraints are met. The fine-tuning process keeps the overall motion trend unchanged, only optimizing local details, ensuring that the reconstruction result not only meets the overall trend of time series prediction, but also meets the physical constraints of auxiliary sensing data. This multi-modal cross-verification mechanism significantly improves the physical rationality of the reconstruction result, ensuring that even in the case of complete visual data loss, the system can still generate skeletal key point data that conforms to the actual motion state of the patient based on auxiliary sensing data.

[0093] In the embodiments of the present application, based on the whole-day skeletal motion sequence, the detailed implementation steps for calculating the patient's sub-period activity index and rehabilitation action completion degree score include:

[0094] The all-day skeletal motion sequence is segmented according to a preset time window, and the cumulative motion distance and joint range of motion of the skeletal key points are calculated for each segment. Time segmentation is the basic step of activity analysis, which divides continuous data into manageable and comparable units. The segmentation process is based on a preset time window (usually 15-30 minutes, or a specific interval determined according to the rehabilitation training plan), which divides the all-day data into multiple continuous and non-overlapping time periods. In each time period, the system calculates two basic activity indicators: the cumulative motion distance is the sum of the displacement of each skeletal key point in three-dimensional space, reflecting the overall activity level; the joint range of motion is the maximum angular change range of the main joints in this period, reflecting the joint flexibility. These basic indicators directly quantify the activity level and intensity of the patient at different times, providing an objective basis for subsequent activity classification and scoring. Time segmentation analysis can reveal the time distribution pattern of the patient's all-day activity, identify high and low activity periods, and provide a reference for the time arrangement optimization of the rehabilitation plan.

[0095] The duration of the patient in different postures such as sitting, standing, walking and rehabilitation training is counted for each time period, and the posture classifier is used to identify the posture types. Posture classification is a key step in activity type analysis, which distinguishes different types of activities and quantifies their time distribution. The classification process uses a machine learning model to identify common posture types based on the spatial distribution and dynamic change characteristics of the skeletal key points. The classifier usually uses a decision tree ensemble or convolutional neural network structure, which is trained through learning a large amount of labeled data, and can accurately distinguish between static postures (such as sitting, standing, and lying in bed) and dynamic activities (such as walking, going up and down stairs, and rehabilitation training movements). For rehabilitation training movements, the system further identifies specific training types, such as upper limb training, lower limb training, and balance training, and matches them with the preset rehabilitation plan. The posture classification results provide detailed activity composition for each time period, not only quantifying the duration of each type of activity, but also revealing the transition patterns between activities, providing rich information for a comprehensive assessment of the patient's daily functional status.

[0096] Energy consumption coefficients are set for different posture types, and the cumulative motion distance, joint range of motion and energy consumption coefficients are weighted and summed to calculate the activity indicators in each time period. Activity indicator calculation is a key step in integrating activity characteristics of different dimensions into a unified evaluation standard. The calculation process first assigns energy consumption coefficients to each type of posture activity, which is determined based on exercise physiology research and reflects the energy consumption level per unit time, for example, sitting is the baseline value 1.0, standing is 1.2-1.5, walking is 2.0-3.0, and rehabilitation training is set to 2.0-5.0 according to intensity; then the cumulative motion distance and joint range of motion are standardized to eliminate dimensional differences; finally, the comprehensive activity indicator is calculated by weighted summation formula:

[0097] ;

[0098] in, For activity level indicators, for posture Duration, for posture The energy consumption coefficient, For standardized cumulative distance of movement, Standardized range of motion for joints. , , These are the weighting coefficients, and .

[0099] The weights are dynamically adjusted according to rehabilitation goals, with an emphasis on endurance training. Higher, when focusing on activity level Higher, when focusing on joint function The activity level is relatively high. The activity level index at different time periods directly reflects the intensity and quality of the patient's activities at each time period, making it easy to compare different time periods horizontally and track rehabilitation progress vertically.

[0100] Movement segments matching pre-defined rehabilitation movement templates are extracted from the 24-hour skeletal motion sequence, and the dynamic time warping distance between the joint angle trajectories of the movement segments and the template is calculated. Rehabilitation movement matching is a crucial step in assessing the quality of treatment execution, identifying the rehabilitation training actually performed by the patient and evaluating its accuracy. The matching process first searches for specific pattern features in the 24-hour sequence to initially locate possible rehabilitation movement segments; then, candidate segments are precisely compared with pre-defined rehabilitation movement templates to confirm movement types and boundaries; finally, the similarity distance between the joint angle trajectories and the template is calculated using the dynamic time warping (DTW) algorithm. The DTW algorithm can handle movement comparison at different speeds and rhythms, finding the optimal matching path through non-linear time alignment to obtain a more accurate similarity assessment. The smaller the similarity distance, the closer the movement execution is to the standard, and the higher the quality. This template-matching-based method can accurately identify all rehabilitation training movements performed by patients in their daily activities and provide an objective quality assessment, offering medical personnel a reliable basis for understanding patient treatment compliance and execution quality.

[0101] Based on the dynamic time warping distance (DTW), the duration deviation of movement completion, and the joint range of motion deviation, a rehabilitation movement completion score is calculated, and a personalized rehabilitation suggestion report is generated. The completion score is a comprehensive assessment of the rehabilitation treatment effect, quantifying the accuracy and effectiveness of the patient's execution of rehabilitation movements. The score calculation comprehensively considers three key dimensions: the accuracy of the movement trajectory (assessed via DTW distance), the rationality of the movement rhythm (assessed via duration deviation), and the adequacy of the movement range (assessed via joint range of motion deviation). The scoring formula is designed as follows:

[0102] ;

[0103] wherein, is the completion score (0-100), is the normalized DTW distance, is the duration deviation, is the standard duration, is the range of motion deviation, is the standard range of motion, , , is the weight coefficient, which is adjusted according to the focus of rehabilitation treatment.

[0104] The score results are summarized according to the action type and time period to form a completion report of the rehabilitation action; at the same time, the system generates targeted improvement suggestions based on the score results and the patient's rehabilitation stage, such as adjusting the action range, optimizing the action rhythm, and enhancing the activity of specific joints. This combination of quantitative completion score and personalized suggestions provides medical personnel with objective treatment effect evaluation and optimization direction, and provides patients with clear progress feedback and improvement goals, significantly improving the relevance and effectiveness of rehabilitation treatment.

[0105] In the embodiment of the application, the anatomical constraint verification of the key points of the skeleton in the abnormal jump frame is performed, and the detailed implementation steps of calculating whether the joint angle exceeds the physiological activity range of the human body include:

[0106] According to the coordinates of the key points of the skeleton in the abnormal jump frame, the flexion angle, the abduction angle and the rotation angle of the main joints are calculated. The joint angle calculation is the basic step of the anatomical verification, which converts the spatial coordinates into angle parameters with physiological significance. The calculation process is based on a three-dimensional anatomical coordinate system. First, a local coordinate system is established for each main joint (such as shoulder, elbow, wrist, hip, knee and ankle) to correctly reflect its anatomical direction; then the three degrees of freedom angles of the joint in the coordinate system are calculated: the flexion angle reflects the activity in the sagittal plane, the abduction angle reflects the activity in the coronal plane, and the rotation angle reflects the rotation in the transverse plane. The angle calculation uses Euler angles or quaternion method to ensure the accuracy and continuity of the numerical value. This anatomically based angle decomposition method makes the verification result directly correspond to the medical standard, which is convenient for clinical interpretation and application. The calculated joint angles comprehensively reflect the current state of the skeletal system, providing accurate input for subsequent physiological range verification.

[0107] The preset human joint range of motion database is queried to obtain the physiological upper and lower limits of each joint under different genders, ages, and rehabilitation stages. Physiological range acquisition is a key step in setting verification standards, ensuring the scientificity and individualization of the judgment basis. The database is constructed based on human kinematics research and clinical rehabilitation medicine data, recording the normal range of motion of each major joint under different conditions. The database structure includes multiple parameters: basic parameters such as joint type, movement direction, and gender differences; individualized parameters such as age, physical condition, and rehabilitation stage. The query process first filters the applicable range based on the patient's basic information (such as gender and age); then makes precise adjustments based on the rehabilitation stage and specific conditions (such as postoperative state and specific diseases); and finally obtains the individualized physiological angle limits for the current patient. This individualized range setting ensures that the verification standard not only meets medical standards but also adapts to the patient's specific situation, avoiding one-size-fits-all judgment errors. The physiological upper and lower limits provide accurate reference for subsequent out-of-range judgment and are the core basis for anatomical verification.

[0108] The flexion-extension angle, abduction-adduction angle, and rotation angle are judged to see if they are within the physiological upper and lower limits, and the number of joints exceeding the limits is counted. Out-of-range judgment is the core step of anatomical verification, identifying abnormal joint states that violate physiological laws. The judgment process compares each joint's three degrees of freedom angles with the corresponding physiological upper and lower limits to determine if they are out of range. For angles that exceed the range, the degree of exceeding (i.e., the difference from the boundary) is further calculated to assess the severity of the abnormality. Finally, the total number of joints exceeding the limits is counted to quantify the overall degree of anatomical abnormality. The judgment process takes into account the coupling relationship between angles, and some angle combinations may not conform to physiological laws even if they are within the range individually. Such cases are also marked as abnormal. The number of joints exceeding the limits is a key indicator for evaluating the anatomical reasonableness of the entire frame of data. The more the number, the higher the degree of abnormality, and the more likely it is a false identification result. This verification method based on medical knowledge can accurately identify false data that may appear reasonable in numerical terms but actually violates human physiological laws.

[0109] The length change rate of the skeleton chain of the abnormal jump frame and the front and rear frames is calculated, and when the length change rate of the skeleton chain exceeds the preset deformation threshold, it is determined that the skeleton structure is distorted. The skeleton structure verification is an important supplement to the anatomical constraint. Based on the principle that the length of the human skeleton is constant, the physically impossible deformation of the skeleton is identified. The verification process first defines a plurality of skeleton chains (such as the upper arm-forearm chain, the thigh-calf chain, etc.) according to the skeleton connection relationship; then the length of the corresponding skeleton chain of the current frame and the front and rear frames is calculated to obtain the length change rate; finally, the length change rate is compared with the preset deformation threshold (usually 1%-3%, considering the measurement error) to determine whether there is abnormal deformation. The length of the skeleton should remain relatively constant in a short period of time, and the change rate exceeding the threshold obviously violates the physical law, indicating that there is an identification error. This verification method based on rigid body constraint can effectively identify those errors that have reasonable joint angles but distorted spatial structure, and is an important part of the anatomical verification.

[0110] The frame in which the number of joints exceeding the upper and lower limits of the physiological angle is greater than the preset number threshold, or the skeleton structure is distorted, is confirmed as an error identification candidate frame. The final determination is a decision step that integrates the verification results to determine the target frame that needs to be compensated for errors. The determination adopts an "or" logic, that is, as long as any condition is met, it is determined as an error identification candidate frame: the joint angle condition is that the number of joints exceeding the physiological range is greater than the preset threshold (usually 15%-25% of the total number of joints, adjusted according to the complexity of the model); the skeleton structure condition is that there is a skeleton chain with a length change rate exceeding the threshold. This multi-condition determination strategy improves the detection rate of error identification and ensures that potential problem data is not missed. The data confirmed as an error identification candidate frame will enter the subsequent error compensation process and be corrected by the multi-modal fusion correction model to eliminate anatomical abnormalities and restore reasonable skeleton structure and joint state. This strict anatomical verification mechanism is an important part of the system data quality control, which ensures that the subsequent analysis is based on physiologically reasonable skeletal data.

[0111] In the embodiment of the present application, the detailed implementation steps of the spatial displacement correction of the skeletal key point coordinates in the visual skeletal tracking data based on the direction and amplitude of the correction vector include:

[0112] The correction vector is decomposed into three components along the anterior-posterior axis, left-right axis, and superior-inferior axis of the human anatomical coordinate system. Vector decomposition is a crucial step in adapting to the characteristics of human anatomy, making the correction process more biomechanically sound. The decomposition process first establishes a patient-centered anatomical coordinate system: the anterior-posterior axis points to the front and back of the patient, the left-right axis points to the left and right sides, and the superior-inferior axis points to the top of the head and the soles of the feet. Then, the original correction vector is transformed from the world coordinate system to this anatomical coordinate system, obtaining the component values ​​in the three directions. This anatomically based decomposition method makes the correction process more intuitive and controllable: the anterior-posterior component mainly affects sagittal plane motion, the left-right component mainly affects coronal plane motion, and the superior-inferior component mainly affects vertical motion. Decomposition in the anatomical coordinate system also facilitates the setting of directional correction constraints; for example, if the movement of certain joints is restricted in a specific direction, the corresponding components can be adjusted accordingly. Vector decomposition lays the foundation for subsequent personalized corrections, enabling the system to make precise adjustments based on the characteristics of different joints.

[0113] Based on the confidence matrix of skeletal keypoints, a correction sensitivity coefficient is assigned to each keypoint; the lower the confidence level, the higher the correction sensitivity coefficient. Sensitivity assignment is the key mechanism for achieving adaptive correction, adjusting the correction strength according to data reliability. The assignment process is based on the previously generated skeletal keypoint confidence matrix, calculating the correction sensitivity coefficient through an inverse proportional relationship:

[0114] ;

[0115] in, Key point The corrected sensitivity coefficient, Let i be the confidence value for key point i. This is the scaling factor (usually between 0.5 and 1.5). This is a smoothing factor (usually taken as 0.1-0.3 to avoid the denominator being close to zero).

[0116] This formula assigns a higher sensitivity coefficient to keypoints with lower confidence levels, resulting in a more significant correction effect; while keypoints with higher confidence levels receive smaller corrections, preserving the advantages of the original data. The sensitivity coefficient typically ranges from [0, 5], with the upper limit appropriately increased for extremely low confidence levels. This adaptive correction mechanism avoids excessive intervention in reliable data while ensuring sufficient correction of unreliable data, thus optimizing the overall correction effect.

[0117] The three components of the correction vector are multiplied by the corresponding bone key point correction sensitivity coefficient to obtain the individualized correction vector of each key point. Individualized correction is an accurate adjustment for different key point characteristics, improving the adaptability and accuracy of the correction. The calculation process multiplies the three components in the anatomical coordinate system with the correction sensitivity coefficient respectively to obtain the adjusted component values; then recombines them into a three-dimensional correction vector as the individualized correction vector of the key point. The individualized correction vector considers both the overall correction direction and the reliability characteristics of the key point itself, realizing the fine correction strategy of "same direction, different intensity". This individualized correction method avoids the discordance problem caused by simple displacement correction, making the final correction result more natural and harmonious, especially in complex poses and actions.

[0118] The bone chain constraint is applied to the individualized correction vector to ensure that the distance between the adjacent key points after correction remains within a reasonable bone length range. Bone chain constraint is a key mechanism to maintain biomechanical reasonableness, avoiding the destruction of the integrity of the bone structure during the correction process. The constraint process first defines the connection relationship between key points according to the bone model, marking each pair of adjacent key points; then predicts the position of the corrected key points, calculates the new distance between adjacent points; compares the new distance with the standard bone length to detect whether it exceeds the allowed range (usually ±3% of the standard length); for the correction vector that violates the constraint, scaling or direction adjustment is performed until the bone length constraint is met. The constraint processing uses an iterative optimization strategy to adjust from the root node to the end of the bone chain, ensuring the reasonableness of the overall structure. This bone chain-based constraint processing ensures that the correction process does not introduce new biomechanical errors, maintaining the integrity and reasonableness of the bone structure, which is an important guarantee for high-quality correction.

[0119] The individualized correction vector is superimposed on the original bone key point coordinates to complete the spatial displacement correction and update the bone key point confidence matrix. Displacement execution is the last step to complete the correction, which applies the calculated correction vector to the actual coordinates. The execution process performs vector addition of the individualized correction vector that meets the constraints and the original key point coordinates to obtain the new coordinates after correction; at the same time, the bone key point confidence matrix is updated to improve the confidence value of the corrected key points, reflecting the quality improvement brought by the correction. The confidence update formula is:

[0120] ;

[0121] wherein, is the updated confidence value, is the original confidence value, is the applied correction vector, is the maximum allowed correction amplitude, is the improvement coefficient (usually 0.3-0.6).

[0122] The confidence level is proportional to the correction range, the more obvious the correction, the more serious the problem, and the more significant the quality improvement after correction.The updated confidence matrix provides more accurate quality assessment for subsequent analysis, while recording the correction history for continuous learning and optimization of the system.The key point coordinates of the skeleton after space displacement correction not only maintain the basic characteristics of the original data, but also eliminate obvious errors and abnormalities, providing high-quality skeletal motion data for subsequent analysis.

[0123] The application constructs an accurate skeletal key point confidence evaluation mechanism and a multi-modal fusion correction model by real-time acquisition of visual skeletal tracking data, inertial measurement unit data and pressure distribution data during the patient's rehabilitation process, combines the skeletal motion pattern memory bank to intelligently compensate and reconstruct potential tracking interruptions and errors, and realizes accurate monitoring and evaluation of the patient's rehabilitation process through continuous and complete whole-day skeletal motion sequences.

[0124] The above only describes the preferred embodiments of the present application and is not intended to limit the present application.Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can modify the technical solutions described in the foregoing embodiments or replace some of the technical features with equivalent ones without departing from the spirit and principles of the present application.Any modification, equivalent replacement, improvement, etc.made within the spirit and principles of the present application shall be included in the protection scope of the present application.

[0125] It should be noted that the formulas in the specification are dimensionless numerical calculations, the formulas are obtained by software simulation of a large amount of data to obtain a formula closest to the real situation, and the preset parameters and threshold values in the formula are set by those skilled in the art according to the actual situation.

[0126] Although the embodiments of the present application have been shown and described, those skilled in the art can understand that various changes, modifications, replacements and variations can be made to these embodiments without departing from the principles and purposes of the present application, and the scope of the present application is defined by the claims and their equivalents.

Claims

1. A rehabilitation therapy monitoring system based on multi-modal sensing and AI evaluation, characterized in that, The method comprises the following steps: a data acquisition module is used to acquire visual bone tracking data, inertial measurement unit data and pressure distribution data during patient rehabilitation; the pressure distribution data is the plantar pressure distribution when the patient stands and walks; a confidence evaluation module is used to perform multi-dimensional confidence evaluation on the image features of each bone key point in the visual bone tracking data, and generate a bone key point confidence matrix; an interruption risk identification module is used to identify potential tracking interruption risk frames and error identification candidate frames based on the time sequence change rate of the bone key point confidence matrix; a multi-modal data fusion module is used to timestamp align the inertial measurement unit data and the pressure distribution data, construct auxiliary sensing feature vectors, and establish a multi-modal fusion correction model according to the kinematic consistency constraint of the auxiliary sensing feature vectors and the visual bone tracking data; a reconstruction module is used to perform error compensation on the bone key point coordinates in the error identification candidate frames based on the multi-modal fusion correction model, extract joint angle time sequence features and speed time sequence features of the visual bone tracking data, and construct a bone motion pattern memory bank; for the potential tracking interruption risk frames, the bone motion pattern memory bank and a bidirectional time sequence prediction algorithm are combined to reconstruct the missing bone key point data; the compensated bone key point coordinates and the reconstructed bone key point data are spliced to generate a continuous and complete whole-day bone motion sequence; a rehabilitation evaluation module is used to calculate the patient's time-periodic activity index and rehabilitation action completion score based on the whole-day bone motion sequence.

2. The system of claim 1, wherein, The multi-dimensional confidence evaluation on the image features of each bone key point in the visual bone tracking data generates a bone key point confidence matrix, which comprises: extracting the local gradient intensity, edge response value and texture contrast of each bone key point in the image as the image stability features of the key point; calculating the bone length ratio between the bone key point and the adjacent key point, comparing the bone length ratio with the preset human body bone anatomy ratio, and obtaining a bone structure consistency score; statistically analyzing the position jitter amplitude of the bone key point in consecutive frames, predicting the deviation between the position and the actual detection position through Kalman filtering, and calculating the time sequence stability score of the key point; estimating the occlusion probability of the pixel region around the bone key point, and obtaining the visibility score of the key point by analyzing the continuity of the color histogram and the depth gradient mutation; weighting and fusing the image stability features, the bone structure consistency score, the time sequence stability score and the visibility score to generate the bone key point confidence matrix.

3. The system of claim 1, wherein, The identification of potential tracking interruption risk frames and error identification candidate frames based on the time sequence change rate of the bone key point confidence matrix comprises: calculating the confidence decrease gradient of the bone key point confidence matrix between adjacent frames, denoted as the confidence decay rate; statistically analyzing the proportion of key points with confidence lower than the preset confidence threshold in a single frame, denoted as the intra-frame low confidence proportion. marking a frame as the potential tracking interruption risk frame when the confidence decay rate is greater than a preset decay threshold and the intra-frame low confidence proportion is greater than a preset proportion threshold; detecting a position jump distance of the skeletal key points between consecutive frames, and recording a frame as an abnormal jump frame when the position jump distance exceeds a reasonable motion range calculated based on historical motion speed; performing an anatomical constraint verification on the skeletal key points in the abnormal jump frame, calculating whether a joint angle exceeds a human physiological activity range, and marking a frame that violates the physiological activity range as the false recognition candidate frame.

4. The system of claim 1, wherein, The time stamp alignment of the inertial measurement unit data and the pressure distribution data to construct an auxiliary sensing feature vector includes: cross-correlation analysis of the sampling time stamp of the inertial measurement unit data and the frame time stamp of the visual skeletal tracking data to identify a system time delay offset; time stamp correction of the inertial measurement unit data and the pressure distribution data according to the system time delay offset to realize synchronization with the visual skeletal tracking data; extracting three-axis acceleration peak values, angular velocity change rates, and attitude quaternions from the inertial measurement unit data, and performing plantar pressure center trajectory calculation and gait phase identification on the pressure distribution data; feature splicing of the three-axis acceleration peak values, the angular velocity change rates, the attitude quaternions, the plantar pressure center trajectory, and the gait phase in a time window to generate the auxiliary sensing feature vector.

5. The system of claim 1, wherein, The multi-modal fusion correction model is established according to the kinematic consistency constraint of the auxiliary sensing feature vector and the visual skeletal tracking data, including: calculating a joint angular velocity vector and a center of mass acceleration vector from the visual skeletal tracking data as visual kinematic features; establishing a kinematic mapping relationship between the visual kinematic features and the auxiliary sensing feature vector, constructing a consistency loss function by minimizing the feature difference at the same time; collecting multi-modal data of patients under standard rehabilitation actions, training a deep neural network to learn the optimal weight parameters of the consistency loss function; using the prediction error of the deep neural network on newly collected data as a correction feedback signal, generating a correction vector when the visual skeletal tracking data and the auxiliary sensing feature vector are kinematically inconsistent; based on the direction and amplitude of the correction vector, performing spatial displacement correction on the skeletal key point coordinates in the visual skeletal tracking data to establish the multi-modal fusion correction model.

6. The system of claim 1, wherein, The joint angle time sequence feature and the speed time sequence feature of the visual skeletal tracking data are extracted to construct a skeletal motion pattern memory bank, including: calculating a three-dimensional angle sequence of main joints in the visual skeletal tracking data, performing discrete Fourier transform on the three-dimensional angle sequence to extract main frequency components and harmonic components in the frequency domain as joint angle frequency domain features; the main joints include shoulder joints, elbow joints, wrist joints, hip joints, knee joints, and ankle joints; Time-differentiate the position coordinates of the skeletal key points to obtain instantaneous speeds of the key points, and count mean, variance and peak of the instantaneous speeds as speed statistical features; input the joint angle frequency domain features and the speed statistical features into a long short-term memory network to extract time sequence encoding vectors of different rehabilitation actions; perform cluster analysis on the time sequence encoding vectors to classify similar motion patterns into the same category, and each category corresponds to a motion pattern prototype; store the motion pattern prototypes and their corresponding joint angle change curves, speed change curves and duration distributions to construct the skeletal motion pattern memory bank.

7. The system of claim 1, wherein, The method for reconstructing missing skeletal key point data in the potential tracking interruption risk frame in combination with the skeletal motion pattern memory bank and a bidirectional time sequence prediction algorithm comprises: identify effective tracking frames before and after the potential tracking interruption risk frame, and extract skeletal motion features of the effective tracking frames; perform similarity matching between the skeletal motion features and motion pattern prototypes in the skeletal motion pattern memory bank to retrieve the most similar historical motion pattern; based on a joint angle change curve of the historical motion pattern, construct a forward time sequence prediction model to predict skeletal key point positions of the interruption frame backward from the frame before interruption; at the same time, construct a backward time sequence prediction model to predict skeletal key point positions of the interruption frame forward from the frame after interruption; perform weighted fusion on prediction results of the forward time sequence prediction model and the backward time sequence prediction model, and assign weights according to the inverse ratio of the time distance from the prediction starting point to the interruption frame to generate reconstructed skeletal key point data; use the measurement value of the auxiliary sensing feature vector at the interruption moment as a physical constraint to perform kinematic rationality verification and fine tuning on the reconstructed skeletal key point data.

8. The system of claim 1, wherein, The method for calculating a patient's time-period activity index and rehabilitation action completion score based on the whole-day skeletal motion sequence comprises: segment the whole-day skeletal motion sequence according to a preset time window, and calculate cumulative motion distance and joint range of motion of skeletal key points in each segment; statistically count the duration of the patient in different postures in each time period, and identify posture types through a posture classifier; set energy consumption coefficients for different posture types, and perform weighted summation on the cumulative motion distance, the joint range of motion and the energy consumption coefficients to calculate the time-period activity index; extract an action segment matching a preset rehabilitation action standard template from the whole-day skeletal motion sequence, calculate a dynamic time warping distance between a joint angle trajectory of the action segment and the standard template, and calculate the rehabilitation action completion score according to the dynamic time warping distance, duration deviation of action completion and joint activity amplitude deviation, and generate a personalized rehabilitation recommendation report.

9. The system of claim 3, wherein, The method for performing anatomical constraint verification on the skeletal key points in the abnormal jump frame to calculate whether the joint angle exceeds the physiological activity range of the human body comprises: calculate flexion, extension, abduction and rotation angles of main joints including shoulder joint, elbow joint, wrist joint, hip joint, knee joint and ankle joint according to the skeletal key point coordinates of the abnormal jump frame; query a preset human joint range of motion database to obtain physiological upper and lower limits of each joint in different genders, ages, and rehabilitation stages; determine whether the flexion-extension angle, the abduction-adduction angle, and the rotation angle are within the physiological upper and lower limits, and count the number of joints that exceed the limits; calculate a skeletal chain length change rate of the abnormal jump frame and the frames before and after the abnormal jump frame, and determine that the skeletal structure is distorted when the skeletal chain length change rate exceeds a preset deformation threshold; frames in which the number of joints that exceed the physiological upper and lower limits is greater than a preset number threshold or the skeletal structure is distorted are confirmed as the error recognition candidate frames.

10. The system of claim 5, wherein, based on the direction and amplitude of the correction vector, the spatial displacement correction of the skeletal key point coordinates in the visual skeletal tracking data includes: decompose the correction vector into three components along the front-back axis, the left-right axis, and the up-down axis of the human anatomical coordinate system; assign a correction sensitivity coefficient to each skeletal key point according to the skeletal key point confidence matrix, and the lower the confidence of a key point, the higher the correction sensitivity coefficient of the key point; multiply the three components of the correction vector by the correction sensitivity coefficient of the corresponding skeletal key point to obtain an individualized correction vector for each key point; apply a skeletal chain constraint to the individualized correction vector to ensure that the distance between the corrected adjacent key points is within a reasonable skeletal length range; superimpose the individualized correction vector on the original skeletal key point coordinates to complete the spatial displacement correction and update the skeletal key point confidence matrix.

Citation Information

Patent Citations

  • Patient rehabilitation training data acquisition method and system based on visual identification

    CN119446391A

  • Multi-mode optical non-contact health monitoring system and method

    CN120392025A