A method and apparatus for optimizing rehabilitation training strategies based on digital twins and adaptive learning
By constructing musculoskeletal and physiological function models, generating strategy optimization agents, and synchronizing multi-source physiological data in real time, the problem of inaccurate execution of rehabilitation training strategies by patients in community hospitals and home rehabilitation environments has been solved, and personalized rehabilitation effects have been improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SECOND AFFILIATED HOSPITAL ZHEJIANG UNIV COLLEGE OF MEDICINE
- Filing Date
- 2026-03-03
- Publication Date
- 2026-08-04
AI Technical Summary
In community hospitals and home rehabilitation environments, patients find it difficult to accurately implement rehabilitation training strategies, and doctors find it difficult to assess their rehabilitation status in real time. This results in a lack of continuity and personalization in rehabilitation training strategies, which affects treatment outcomes.
By constructing musculoskeletal and physiological function models of the target object, a strategy optimization agent is generated. Multi-source physiological data is synchronized in real time to predict rehabilitation progress and optimize strategies, thereby generating optimized rehabilitation strategies.
It improved the alignment of rehabilitation training strategies with patients, enhanced rehabilitation outcomes, and enabled real-time monitoring and continuous optimization of rehabilitation progress.
Smart Images

Figure CN121789897B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of rehabilitation training management technology, and in particular to a method and apparatus for optimizing rehabilitation training strategies based on digital twins and adaptive learning. Background Technology
[0002] Rehabilitation training refers to the process of helping individuals whose physical functions have declined due to illness or injury to restore their physiological functions through systematic interventions. During rehabilitation training, targeted rehabilitation strategies are usually developed based on the patient's rehabilitation needs. In reality, patients requiring rehabilitation training often have mobility issues and primarily receive rehabilitation at community hospitals and at home. Without the supervision of a attending physician, it is difficult to ensure accurate implementation of the rehabilitation training strategies, affecting their therapeutic effectiveness. Furthermore, patient specificity exists, and the rehabilitation training process varies from patient to patient. Physicians find it difficult to conduct real-time assessments of the rehabilitation status of patients undergoing home rehabilitation to adjust the training strategies, resulting in a lack of continuous flexibility in the implementation of rehabilitation training strategies. Summary of the Invention
[0003] This application provides a method and apparatus for optimizing rehabilitation training strategies based on digital twins and adaptive learning. The method uses a multi-dimensional digital twin model to reflect the rehabilitation status of the target object in real time, and continuously optimizes the rehabilitation strategy based on the prediction results of the rehabilitation status of the digital twin model. This improves the fit between the rehabilitation training strategy and the target object, and enhances the rehabilitation effect of the target object.
[0004] To achieve the above objectives, the main technical solutions adopted in this application include: In a first aspect, embodiments of this application provide a method for optimizing rehabilitation training strategies based on digital twins and adaptive learning, characterized in that the method includes: Based on the clinical personal information of the target subjects, construct musculoskeletal models and physiological function models, and generate a strategy optimization agent for the target subjects; The target object's rehabilitation training process involves acquiring multi-source physiological data, inputting the multi-source physiological data into the musculoskeletal model and the physiological function model, and performing coupled simulation prediction through the musculoskeletal model and the physiological function model to obtain the rehabilitation progress of the target object. When the rehabilitation progress meets the preset risk conditions, the rehabilitation strategy is predicted and optimized based on the multi-source physiological data by using the strategy optimization agent, the musculoskeletal model, and the physiological function model to generate an optimized rehabilitation strategy for the target object.
[0005] The rehabilitation training strategy optimization method based on digital twins and adaptive learning proposed in this application constructs a musculoskeletal model and a physiological function model for the target object, and generates a strategy optimization agent related to the target object's rehabilitation needs. Multi-source physiological data of the target object during rehabilitation training is synchronized in real time to the musculoskeletal model and the physiological function model, and the rehabilitation progress of the target object at the current stage is predicted to obtain the target object's rehabilitation progress status. When it is determined that the target object has health risks based on the rehabilitation progress status, the rehabilitation strategy is predicted and optimized through the strategy optimization agent, the musculoskeletal model, and the physiological function model to generate an optimized rehabilitation strategy. Compared with related technologies, this application constructs a multi-dimensional digital twin agent for the target object, inputting multi-source physiological data of the target object during rehabilitation training into the multi-dimensional digital twin agent, thereby enabling real-time synchronization of the target object's rehabilitation status from multiple perspectives, so that the attending physician can promptly grasp the target object's rehabilitation training progress. Furthermore, this application also continuously optimizes the rehabilitation training strategy based on multi-source physiological data of the target object through the strategy optimization agent and the rehabilitation effect prediction of the multi-dimensional digital twin agent, thereby improving the fit between the rehabilitation training strategy and the target object and thus improving the rehabilitation effect of the target object.
[0006] Optionally, the rehabilitation training process includes multiple rehabilitation cycles, the multi-source physiological data includes movement data and physiological data, and the acquisition cycle of the physiological data is longer than that of the movement data; the step of obtaining the rehabilitation progress of the target object through coupled simulation prediction using the musculoskeletal model and the physiological function model includes: In any rehabilitation cycle other than the first rehabilitation cycle among the multiple rehabilitation cycles, the rehabilitation effect is predicted by the musculoskeletal model based on the movement data of the rehabilitation cycle and the prediction results of the physiological function model in the previous rehabilitation cycle of the rehabilitation cycle, so as to obtain the rehabilitation training effect of the target object. Based on the movement data of any rehabilitation cycle and the prediction results of the physiological function model in the previous rehabilitation cycle of any rehabilitation cycle, the rehabilitation burden is predicted through the physiological function model to obtain the rehabilitation burden data of the target object. The rehabilitation progress of the target subject in any given rehabilitation cycle is obtained based on the rehabilitation training effect and the rehabilitation burden data.
[0007] Optionally, the step of optimizing the intelligent agent, the musculoskeletal model, and the physiological function model through the strategy, and predicting and optimizing the rehabilitation strategy based on the multi-source physiological data to generate an optimized rehabilitation strategy for the target object includes: A proposed rehabilitation strategy is generated based on the current rehabilitation strategy of the target object. The proposed rehabilitation strategy is input into the musculoskeletal model and the physiological function model. The proposed rehabilitation strategy is then subjected to coupled simulation prediction using the musculoskeletal model and the physiological function model to obtain the predicted rehabilitation status of the target object under the proposed rehabilitation strategy. The reward is calculated based on the predicted recovery status and the reward function of the strategy optimization agent to obtain the reward function value of the proposed recovery strategy; Based on the reward function values of multiple proposed rehabilitation strategies, the parameters of the strategy optimization agent are optimized to obtain the optimal model parameters of the strategy optimization agent. Based on the optimal model parameters, the optimized rehabilitation strategy is generated by the strategy optimization agent.
[0008] Optionally, the rehabilitation training process includes multiple rehabilitation cycles, each corresponding to its own predicted rehabilitation status; the step of calculating the reward based on the predicted rehabilitation status and the reward function of the strategy optimization agent to obtain the reward function value of the proposed rehabilitation strategy includes: Risk features are extracted based on the predicted rehabilitation status in each rehabilitation cycle to obtain the rehabilitation risk features of the proposed rehabilitation strategy in the multiple rehabilitation cycles. Based on the predicted rehabilitation status and cycle length of each of the multiple rehabilitation cycles, the rehabilitation efficiency is evaluated to obtain the rehabilitation rate data of the proposed rehabilitation strategy in the multiple rehabilitation cycles. The reward is calculated based on the predicted recovery status, the recovery risk characteristics, the recovery rate data, and the reward function to obtain the reward function value.
[0009] Optionally, the step of calculating the reward based on the predicted rehabilitation status, the rehabilitation risk characteristics, the rehabilitation rate data, and the reward function to obtain the reward function value includes: Obtain the target subject's subjective feelings about rehabilitation; conduct a graded assessment of the subjective feelings about rehabilitation to obtain the target subject's subjective feeling score; The reward is calculated based on the predicted rehabilitation status, the rehabilitation risk characteristics, the rehabilitation rate data, the subjective feeling score, and the reward function to obtain the reward function value.
[0010] Optionally, generating the optimized rehabilitation strategy through the strategy optimization agent based on the optimal model parameters includes: Based on the rehabilitation knowledge graph, a fusion correlation analysis is performed between the current rehabilitation strategy and the multi-source physiological data and rehabilitation progress of the target object to obtain the target correlation features of the current rehabilitation strategy. Generate a rehabilitation adjustment strategy for the target object based on the target association features; The intermediate rehabilitation strategy is generated by the agent using the strategy optimization of the optimal model parameters, and the intermediate rehabilitation strategy and the rehabilitation adjustment strategy are fused to obtain the optimized rehabilitation strategy.
[0011] Optionally, the policy optimization agent is trained based on the historical rehabilitation data of the sample objects, including the historical initial situation, historical rehabilitation status, historical progress data, and historical risk characteristics of the sample objects; the reward function of the policy optimization agent is obtained through the following method: Based on the historical rehabilitation situation and the historical initial situation, the rehabilitation benefits are calculated to obtain a rehabilitation benefit reward item used to incentivize the strategy optimization agent to improve the rehabilitation effect of the target object; The musculoskeletal model is used to reenact the physiological functions of the historical process data in order to evaluate the rehabilitation efficiency of the historical process data and obtain rehabilitation efficiency reward items to incentivize the strategy optimization agent to improve the rehabilitation efficiency of the target object. The physiological load is replayed on the historical process data by the physiological function model to assess the rehabilitation load of the historical process data and obtain a rehabilitation load penalty term for suppressing the physiological load caused by the policy optimization agent to the target object. Based on the historical risk characteristics, the risk index of the sample object during the rehabilitation process is analyzed to obtain a rehabilitation risk penalty term used to suppress the rehabilitation risk caused by the strategy optimization agent to the target object; The reward function is obtained based on the rehabilitation benefit reward item, the rehabilitation efficiency reward item, the rehabilitation load penalty item, and the rehabilitation risk penalty item.
[0012] Optionally, the rehabilitation training process includes multiple rehabilitation cycles, and the preset risk conditions include a normal threshold range and preset risk characteristics; the progress of rehabilitation is determined to meet the preset risk conditions in the following ways: A rehabilitation risk assessment is performed on the rehabilitation progress of each of the multiple rehabilitation cycles to obtain the rehabilitation risk index for each of the multiple rehabilitation cycles. Based on the rehabilitation risk index of the multiple rehabilitation cycles, temporal features are extracted to obtain the risk temporal features of the target object in the multiple rehabilitation cycles; If the rehabilitation risk index exceeds the normal threshold range and there is similarity between the risk time series characteristics and the preset risk characteristics, the rehabilitation progress is determined to meet the preset risk conditions.
[0013] Optionally, the method further includes: The optimized rehabilitation strategy is sent to the attending physician of the target patient, who then reviews the optimized rehabilitation strategy. Once the optimized rehabilitation strategy is approved, it will be shared with the target individual and the community hospital in their place of residence to achieve three-level linkage management of the rehabilitation process for the target individual.
[0014] Secondly, embodiments of this application provide a rehabilitation training strategy optimization device based on digital twins and adaptive learning, characterized in that the device comprises: The twin model construction module is used to construct a musculoskeletal model and a physiological function model based on the clinical personal information of the target object, and generate a strategy optimization agent for the target object; The rehabilitation progress prediction module is used to acquire multi-source physiological data of the target object during the rehabilitation training process, input the multi-source physiological data into the musculoskeletal model and the physiological function model, and perform coupled simulation prediction through the musculoskeletal model and the physiological function model to obtain the rehabilitation progress of the target object. The rehabilitation strategy optimization module is used to generate an optimized rehabilitation strategy for the target object by using the strategy optimization agent, the musculoskeletal model, and the physiological function model to predict and optimize the rehabilitation strategy based on the multi-source physiological data when the rehabilitation progress meets the preset risk conditions.
[0015] Thirdly, embodiments of this application provide a rehabilitation training linkage management platform for performing the method described in any of the above embodiments.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer instructions, which are used to cause a computer to perform the method described in any one of the above embodiments.
[0017] Fifthly, embodiments of this application provide a computer program product, including computer instructions, which are used to cause a computer to perform the method described in any of the above embodiments. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating the steps of the rehabilitation training strategy optimization method based on digital twin and adaptive learning provided in this application embodiment; Figure 2 This is a flowchart illustrating the steps of coupled simulation prediction in the embodiments of this application; Figure 3 This is a flowchart illustrating the collaboration between the musculoskeletal model and the physiological function model in the embodiments of this application. Figure 4 This is a flowchart illustrating the steps of rehabilitation strategy prediction and optimization in the embodiments of this application; Figure 5 This is a flowchart illustrating the reward calculation steps in an embodiment of this application; Figure 6 This is a diagram illustrating the reward calculation steps that take into account the subjective feelings of rehabilitation in the embodiments of this application; Figure 7 This is a flowchart illustrating the steps involved in generating an optimized rehabilitation strategy in an embodiment of this application. Figure 8 This is a flowchart illustrating the steps involved in obtaining the reward function in an embodiment of this application. Figure 9 This is a flowchart illustrating the steps for determining preset risk conditions in an embodiment of this application; Figure 10 This is a flowchart illustrating the steps of the three-level linkage management in this application embodiment; Figure 11 This is a schematic diagram of the three-level linkage management in the embodiments of this application; Figure 12 A block diagram of a rehabilitation training strategy optimization device based on digital twin and adaptive learning provided in an embodiment of this application; Figure 13 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] Rehabilitation training refers to the process of helping individuals whose physical functions have declined due to illness or injury to restore their physiological functions through systematic interventions. During rehabilitation training, targeted rehabilitation strategies are usually developed based on the patient's rehabilitation needs. In reality, patients requiring rehabilitation training often have mobility issues and primarily receive rehabilitation at community hospitals and at home. Without the supervision of a attending physician, it is difficult to ensure accurate implementation of the rehabilitation training strategies, affecting their therapeutic effectiveness. Furthermore, patient specificity exists, and the rehabilitation training process varies from patient to patient. Physicians find it difficult to conduct real-time assessments of the rehabilitation status of patients undergoing home rehabilitation to adjust the training strategies, resulting in a lack of continuous flexibility in the implementation of rehabilitation training strategies.
[0022] To address the aforementioned issues, this application provides a method and apparatus for optimizing rehabilitation training strategies based on digital twins and adaptive learning. The method constructs a musculoskeletal model and a physiological function model based on the clinical personal information of the target subject, and generates a strategy optimization agent for the target subject. It acquires multi-source physiological data of the target subject during rehabilitation training, inputs this data into the musculoskeletal model and the physiological function model, and performs coupled simulation prediction using these models to obtain the target subject's rehabilitation progress. When the rehabilitation progress meets preset risk conditions, the strategy optimization agent, the musculoskeletal model, and the physiological function model predict and optimize the rehabilitation strategy based on the multi-source physiological data, generating an optimized rehabilitation strategy for the target subject.
[0023] The rehabilitation training strategy optimization method based on digital twins and adaptive learning provided in this application constructs a musculoskeletal model and a physiological function model for the target object, and generates a strategy optimization agent related to the rehabilitation needs of the target object; it synchronizes multi-source physiological data of the target object in the rehabilitation training process to the musculoskeletal model and the physiological function model in real time, and predicts the rehabilitation progress of the target object at the current stage to obtain the rehabilitation progress status of the target object; when it is determined that the target object has health risks based on the rehabilitation progress status, the rehabilitation strategy is predicted and optimized through the strategy optimization agent, the musculoskeletal model and the physiological function model to generate an optimized rehabilitation strategy.
[0024] Compared with related technologies, this application constructs a multi-dimensional digital twin intelligent agent for the target subject, inputting multi-source physiological data of the target subject during rehabilitation training into the multi-dimensional digital twin intelligent agent. This allows for real-time synchronization of the target subject's rehabilitation status from multiple perspectives, enabling the attending physician to promptly grasp the target subject's rehabilitation training progress. Furthermore, based on the target subject's multi-source physiological data, this application continuously optimizes the rehabilitation training strategy through a strategy optimization intelligent agent and the rehabilitation effect prediction of the multi-dimensional digital twin intelligent agent, improving the alignment between the rehabilitation training strategy and the target subject, thereby enhancing the target subject's rehabilitation outcome.
[0025] According to an embodiment of this application, an embodiment of a rehabilitation training strategy optimization method based on digital twin and adaptive learning is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0026] Reference Figure 1 As shown in this embodiment, a method for optimizing rehabilitation training strategies based on digital twins and adaptive learning is provided. This method includes: S100. Construct a musculoskeletal model and a physiological function model based on the target subject's clinical personal information, and generate a strategy optimization agent for the target subject.
[0027] S200. Obtain multi-source physiological data of the target object during the rehabilitation training process, input the multi-source physiological data into the musculoskeletal model and the physiological function model, and perform coupled simulation prediction through the musculoskeletal model and the physiological function model to obtain the rehabilitation progress of the target object.
[0028] S300. When the rehabilitation progress meets the preset risk conditions, the system optimizes the rehabilitation strategy based on multi-source physiological data by using the strategy optimization agent, musculoskeletal model and physiological function model to predict and optimize the rehabilitation strategy, and generates an optimized rehabilitation strategy for the target object.
[0029] The target subject's clinical personal information may include individual data, anatomical data, biomechanical data, and physiological function data. Anatomical data can be imaging data obtained through three-dimensional medical imaging technology to describe the target subject's three-dimensional anatomical structure, including but not limited to CT (Computed Tomography) images and MRI images. Biomechanical data describes the target subject's motor capabilities and may include data such as muscle strength, joint range of motion, and pain scores. Physiological function data describes the target subject's physiological state and may include static and dynamic physiological data. Static physiological data includes data on the target subject in a static state, such as resting heart rate, resting respiratory rate, and resting blood pressure; dynamic physiological data includes data on the target subject during exercise, such as heart rate change curves, respiratory rate change curves, blood pressure change curves, and normal parameter ranges.
[0030] A musculoskeletal model is a digital twin built based on the individual data, anatomical structure data, and biomechanical data of the target object. By simplifying the target object into a mechanical system containing physiological structures such as bones, muscles, and joints, it simulates the motion state of each physiological structure during rehabilitation training and can analyze specific mechanical changes. A physiological function model is a digital twin built based on the individual data and physiological function data of the target object, used to simulate changes in physiological variables within the target object under different exercise conditions. It can be understood that the physiological function model can be built on top of a pre-constructed general model. By inputting the individual data and physiological function data of the target object, a personalized model can be created to realistically reflect the actual physiological state of the target object.
[0031] Multi-source physiological data can be physiological data generated by the target subject during home rehabilitation, which can be acquired through home smart terminals. This includes motion state data, physiological state data, and environmental context data. Physiological state data describes the target subject's physiological state, including heart rate, respiratory rate, and blood pressure during rehabilitation training. Environmental context data describes the interaction between the target subject and their environment, including but not limited to the target subject's activity trajectory, tool usage, and physical environmental parameters.
[0032] Motion state data can be used to describe the motion state of the target object, including limb motion data, muscle biomechanical signals, and dynamic balance data during rehabilitation movements. Limb motion data can include limb movement speed and amplitude, muscle biomechanical signals can include electromyography (EMG) and other muscle biomechanical signals, and dynamic balance data can include support surface pressure data and center of gravity movement data. For example, limb motion data can be acquired by using an inertial measurement unit (IMU) and a camera to collect data from the target object. Muscle biomechanical signals can be obtained by using flexible epidermal strain sensors attached to the skin surface of key muscle groups, non-invasively acquiring the target object's electromyography (EMG) and mechanomyography (MMG) signals in real time. Dynamic balance data can be obtained by quantitatively analyzing the pressure center trajectory of the target object using embedded force-sensitive resistors integrated into yoga mats or insoles.
[0033] Specifically, the process involves acquiring the target subject's clinical personal information, constructing a digital twin based on this information, and obtaining a musculoskeletal model and a physiological function model. The musculoskeletal model construction process may include: performing 3D modeling based on the target subject's anatomical data to create a limb mesh model, including multiple limb units and connecting joints between these units, allowing relative rotation between limb units based on these joints; inputting the target subject's individual data into the limb mesh model, defining parameters such as the mass and moment of inertia of each limb unit, and setting the rotation range of the connecting joints to obtain the musculoskeletal model. The physiological function model construction process may include: performing general model matching based on the target subject's individual data to obtain a matched physiological model; inputting the target subject's physiological function data into the matched physiological model to adjust its parameters, thus obtaining the physiological function model.
[0034] Furthermore, during the rehabilitation training of the target subject, multi-source physiological data of the target subject is acquired in real time through a home smart terminal, and the multi-source physiological data is input into the musculoskeletal model and the physiological function model. The musculoskeletal model simulates the movement state of various physiological structures of the target subject during rehabilitation training, and the physiological function model simulates the changes in physiological variables in the target subject under different movement conditions.
[0035] In some embodiments, the simulation process in the musculoskeletal model may include: during a rehabilitation training process comprising multiple rehabilitation cycles, acquiring, within any given rehabilitation cycle, the predicted state data of the target object from the musculoskeletal model in the previous rehabilitation cycle, and multi-source physiological data of the target object in that rehabilitation cycle; calculating the error between the predicted state data and the multi-source physiological data to obtain the predicted data error for each physiological data item; and fusing the multi-source physiological data with the predicted state data based on the predicted data errors and uncertainty weights for each physiological data item to obtain the biomechanical synchronization data of the target object. Similarly, the simulation process of the physiological function model can be the same as that of the musculoskeletal model, obtaining the physiological synchronization data of the target object, and combining the biomechanical synchronization data and the physiological synchronization data to obtain the synchronized state data of the target object.
[0036] Understandably, by fusing predicted state data and multi-source physiological data to obtain synchronized state data, and then combining this with predicted data based on physical laws and actual acquired detection data, virtual observation of the rehabilitation process within the target subject can be achieved, reducing the uncertainty in understanding the complex human body system. Furthermore, by fusing data based on the uncertainties of both the prediction and observation methods, the data noise of both the predicted state data and the multi-source physiological data is effectively reduced, improving the credibility and reliability of the synchronized state data.
[0037] Furthermore, based on the target object's synchronous state data and rehabilitation training strategy in any given rehabilitation cycle, the musculoskeletal model predicts the target object's biomechanical state after implementing the rehabilitation training strategy in the next rehabilitation cycle, and the physiological function model predicts the target object's physiological function state after implementing the rehabilitation training strategy in the next rehabilitation cycle, thus obtaining the target object's rehabilitation progress. It can be understood that the musculoskeletal model and the physiological function model can predict the target object's state from different dimensions. The prediction results output by the musculoskeletal model can be used to adjust the prediction method of the physiological function model, correcting the prediction path between the rehabilitation training strategy and the output results; while the prediction results output by the physiological function model can describe the changes in the target object's physiological function conditions during rehabilitation training, serving as constraints for the musculoskeletal model and improving its prediction accuracy.
[0038] Furthermore, risk identification is conducted based on the target individual's rehabilitation progress. This involves analyzing whether the current rehabilitation training strategy poses a risk of harm or prevents the target individual from completing the strategy, thus determining if the rehabilitation progress meets pre-set risk conditions. If the rehabilitation progress does not meet the pre-set risk conditions, it indicates that the risk level of the rehabilitation training strategy is within an acceptable range and can help the target individual restore physiological function; in this case, the current rehabilitation training strategy can be maintained. If the rehabilitation progress meets the pre-set risk conditions, then the rehabilitation training strategy needs to be optimized and adjusted to reduce the risk of harm to the target individual.
[0039] In some embodiments, the process of rehabilitation strategy prediction and optimization may include: fine-tuning the rehabilitation training strategy through a strategy optimization agent, and predicting the rehabilitation effect based on the fine-tuned rehabilitation training strategy using a musculoskeletal model and a physiological function model, so as to select the rehabilitation training strategy with the best rehabilitation effect for the target subject. The process of rehabilitation strategy prediction and optimization may also include: generating a long-term macro-strategy through a strategy optimization agent, and periodically filling the specific training parameters in the long-term macro-strategy using a musculoskeletal model and a physiological function model, so as to adjust the rehabilitation training strategy in real time during each rehabilitation cycle.
[0040] The rehabilitation training strategy optimization method based on digital twins and adaptive learning provided in this embodiment constructs a musculoskeletal model and a physiological function model for the target object, and generates a strategy optimization agent related to the rehabilitation needs of the target object; it synchronizes multi-source physiological data of the target object during the rehabilitation training process to the musculoskeletal model and the physiological function model in real time, and predicts the rehabilitation progress of the target object at the current stage to obtain the rehabilitation progress status of the target object; when it is determined that the target object has health risks based on the rehabilitation progress status, the rehabilitation strategy is predicted and optimized through the strategy optimization agent, the musculoskeletal model and the physiological function model to generate an optimized rehabilitation strategy.
[0041] Compared with related technologies, this application constructs a multi-dimensional digital twin intelligent agent for the target subject, inputting multi-source physiological data of the target subject during rehabilitation training into the multi-dimensional digital twin intelligent agent. This allows for real-time synchronization of the target subject's rehabilitation status from multiple perspectives, enabling the attending physician to promptly grasp the target subject's rehabilitation training progress. Furthermore, based on the target subject's multi-source physiological data, this application continuously optimizes the rehabilitation training strategy through a strategy optimization intelligent agent and the rehabilitation effect prediction of the multi-dimensional digital twin intelligent agent, improving the alignment between the rehabilitation training strategy and the target subject, thereby enhancing the target subject's rehabilitation outcome.
[0042] Reference Figure 2As shown in this embodiment of the application, the rehabilitation training process includes multiple rehabilitation cycles, and the multi-source physiological data includes movement data and physiological data, with the acquisition cycle for physiological data being longer than that for movement data. Coupled simulation prediction is performed using a musculoskeletal model and a physiological function model to obtain the rehabilitation progress of the target object, including: S210. In any rehabilitation cycle other than the first rehabilitation cycle, based on the movement data of any rehabilitation cycle and the prediction results of the physiological function model in the previous rehabilitation cycle, the rehabilitation effect is predicted by the musculoskeletal model to obtain the rehabilitation training effect of the target object.
[0043] S220. Based on the movement data of any rehabilitation cycle and the prediction results of the physiological function model in the previous rehabilitation cycle, the rehabilitation burden is predicted through the physiological function model to obtain the rehabilitation burden data of the target object.
[0044] S230. Obtain the rehabilitation progress of the target subject within any rehabilitation cycle based on the rehabilitation training effect and rehabilitation burden data.
[0045] Specifically, related technologies typically employ a single, fully integrated digital twin to synchronize the state of a target object and simultaneously predict its biomechanical and physiological functional states. In practice, the synchronization and prediction of biomechanical states require calculations based on mechanical equations, while the synchronization and prediction of physiological functional states require calculations based on energy balance equations. Forcibly coupling two different types of equations within the same digital twin means the calculation step size can only be set based on the shorter step size. Furthermore, the multi-source physiological data of the target object are not all acquired at the same time intervals, failing to meet the requirements of high-frequency calculations. In addition, biomechanical states are calculated based on the target object's anatomical structure data, while physiological functional states are calculated based on the target object's physiological function data; calibration and verification of these different calculation processes must be performed separately and cannot be done simultaneously.
[0046] Considering the above issues, this application constructs a skeletal muscle model and a physiological function model for the target object, respectively, and performs coupled simulation prediction of the target object through the collaborative work of the skeletal muscle model and the physiological function model. (Refer to...) Figure 3As shown, the musculoskeletal model and the physiological function model perform simulation predictions independently. The musculoskeletal model predicts the rehabilitation effect based on the motion state data collected in cycle T1, predicting the degree of physical recovery of the target subject after implementing the rehabilitation training strategy of the current rehabilitation cycle. The physiological function model predicts the rehabilitation burden based on the physiological state data collected in cycle T2, predicting the physical burden caused by the rehabilitation training strategy of the current rehabilitation cycle to the target subject. For example, the duration of cycle T2 is longer than the duration of cycle T1. The duration of cycle T1 can be the same as the duration of one rehabilitation cycle, and the duration of cycle T2 can be the same as the duration of four rehabilitation cycles.
[0047] In the first rehabilitation cycle, the musculoskeletal model predicts the rehabilitation effect based on collected motion and physiological data, obtaining the rehabilitation training effect for the first cycle. The physiological function model predicts the rehabilitation burden based on the collected motion and physiological data, obtaining the rehabilitation burden effect for the first cycle. After the first rehabilitation cycle, the musculoskeletal model shares the rehabilitation training effect with the physiological function model, and the physiological function model shares the rehabilitation burden effect with the musculoskeletal model. Understandably, since no physiological data is collected in the second rehabilitation cycle, the musculoskeletal model and physiological function model cannot obtain the physiological data for calculation. In this case, the rehabilitation burden effect of the first rehabilitation cycle can be input as physiological data into the musculoskeletal model and physiological function model. The musculoskeletal model predicts the rehabilitation effect based on the motion data and the rehabilitation burden effect of the first cycle, obtaining the rehabilitation training effect for the second cycle; the physiological function model predicts the rehabilitation burden based on the motion data and the rehabilitation burden effect of the first cycle, obtaining the rehabilitation burden effect for the second cycle. Similarly, the third and fourth rehabilitation cycles can obtain the rehabilitation training effect and rehabilitation burden effect in the same way as the second rehabilitation cycle, until new physiological data is collected in the fifth rehabilitation cycle.
[0048] Reference Figure 4 As shown, in one embodiment of this application, a rehabilitation strategy is generated by optimizing a strategy agent, a musculoskeletal model, and a physiological function model based on multi-source physiological data to predict and optimize the rehabilitation strategy for the target object, including: S310. Generate a proposed rehabilitation strategy based on the target object's current rehabilitation strategy, input the proposed rehabilitation strategy into the musculoskeletal model and the physiological function model, and perform coupled simulation prediction on the proposed rehabilitation strategy through the musculoskeletal model and the physiological function model to obtain the predicted rehabilitation status of the target object under the proposed rehabilitation strategy.
[0049] S320. Calculate the reward based on the predicted recovery status and the reward function of the strategy optimization agent to obtain the reward function value of the proposed recovery strategy.
[0050] S330. Based on the reward function values of multiple proposed rehabilitation strategies, optimize the parameters of the strategy optimization agent to obtain the optimal model parameters of the strategy optimization agent.
[0051] S340. Based on the optimal model parameters, optimize the rehabilitation strategy by generating an intelligent agent through strategy optimization.
[0052] Specifically, the strategy optimization agent can be an agent built based on a reinforcement learning network, possessing a strategy action space. This strategy action space, derived from medical knowledge, ergonomic constraints, and individual and biomechanical data of the target object, represents the range of rehabilitation movements the target object can perform. Based on the target object's multi-source physiological data from the previous rehabilitation cycle, the strategy optimization agent searches for optimized strategies within the strategy action space, outputting proposed rehabilitation strategies. It is understood that the optimized strategy search can involve adjusting one or more parameters of the rehabilitation training strategy, such as training parameters, movement type, movement frequency, and movement duration.
[0053] Furthermore, using musculoskeletal and physiological function models, and based on multi-source physiological data of the target object, the rehabilitation training effect and rehabilitation burden data of the target object after implementing the proposed rehabilitation strategy are predicted, thus obtaining the predicted rehabilitation status of the target object under the proposed rehabilitation strategy. The reward function of the strategy optimization agent is calculated based on the predicted rehabilitation status, yielding the reward function value of the proposed rehabilitation strategy, which is used to determine the rehabilitation benefits that the proposed rehabilitation strategy can generate for the target object.
[0054] Furthermore, based on multiple proposed rehabilitation strategies and their corresponding reward function values, the parameters of the strategy optimization agent are optimized until the parameter optimization termination condition is met, thus obtaining the optimal model parameters of the strategy optimization agent. In some embodiments, the parameter optimization process may include: calculating the policy gradient among multiple proposed rehabilitation strategies to determine the policy adjustment gradient; calculating the reward gradient for the reward function values corresponding to multiple proposed rehabilitation strategies to obtain the reward change gradient among multiple proposed rehabilitation strategies; adjusting the parameters of the strategy optimization agent according to the policy adjustment gradient and the reward change gradient, so that the strategy optimization agent tends to output proposed rehabilitation strategies with high rehabilitation benefits. The parameter optimization process may also include: dividing the strategy optimization agent into an actor network and a critic network; learning the proposal method for generating proposed rehabilitation strategies through the actor network; defining an evaluation system for the critic network to evaluate the proposal results output by the actor network; and adjusting the network parameters of the actor network according to the evaluation results of the actor network to obtain the optimal model parameters of the strategy optimization agent.
[0055] Furthermore, upon meeting the termination conditions for parameter optimization, parameter optimization is stopped, and the optimal model parameters are used. Based on the optimal model parameters, an optimized rehabilitation strategy is generated by a policy optimization agent to assist the target object's rehabilitation process. For example, the termination conditions may include the policy optimization agent's performance no longer changing, the policy optimization agent's performance level reaching a preset level, or the optimization count reaching its maximum limit.
[0056] Reference Figure 5 As shown, in one embodiment of this application, the rehabilitation training process includes multiple rehabilitation cycles, each corresponding to its own predicted rehabilitation status; rewards are calculated based on the predicted rehabilitation status and the reward function of the strategy optimization agent to obtain the reward function value of the proposed rehabilitation strategy, including: S322. Based on the predicted rehabilitation status in each rehabilitation cycle, risk features are extracted to obtain the rehabilitation risk features of the proposed rehabilitation strategy in multiple rehabilitation cycles.
[0057] S324. Evaluate rehabilitation efficiency based on the predicted rehabilitation status and cycle length of each of the multiple rehabilitation cycles to obtain rehabilitation rate data of the proposed rehabilitation strategy in multiple rehabilitation cycles.
[0058] S326. Calculate the reward based on the predicted rehabilitation status, rehabilitation risk characteristics, rehabilitation rate data, and reward function to obtain the reward function value.
[0059] Specifically, a machine learning model extracts risk features based on predicted recovery outcomes to obtain rehabilitation risk features for the proposed rehabilitation strategy, representing the potential health risks faced by the target individual when implementing the proposed strategy. For example, the machine learning model used for risk feature extraction could be a survival analysis model or a time series prediction model, etc.
[0060] Furthermore, rehabilitation efficiency is assessed based on the predicted rehabilitation status and cycle length of each of the multiple rehabilitation cycles. The time required for the target subject to reach the predicted rehabilitation status is calculated to determine the rehabilitation efficiency of the target subject when implementing the proposed rehabilitation strategy, thereby obtaining the rehabilitation rate data of the proposed rehabilitation strategy.
[0061] Furthermore, based on the predicted rehabilitation status, rehabilitation risk characteristics, rehabilitation rate data, and reward function, the reward term and penalty term in the reward function are calculated respectively to obtain the reward value and penalty value, and then the reward function value is obtained.
[0062] Reference Figure 6 As shown in one embodiment of this application, a reward function value is obtained by calculating the reward based on the predicted rehabilitation status, rehabilitation risk characteristics, rehabilitation rate data, and a reward function, including: S328. Obtain the target subject's subjective feelings about rehabilitation; conduct a graded assessment of the subjective feelings about rehabilitation to obtain the target subject's subjective feeling score.
[0063] S329. Calculate the reward based on the predicted rehabilitation status, rehabilitation risk characteristics, rehabilitation rate data, subjective feeling score, and reward function to obtain the reward function value.
[0064] Specifically, the subjective feelings of the target subject during the training process are obtained, including data such as pain feedback scores, sleep quality and perceived fatigue. Each subjective feeling of the target subject is graded and evaluated to obtain a subjective feeling score.
[0065] Furthermore, the reward and penalty terms in the reward function are calculated based on the subjective feeling score, and the reward is calculated based on the predicted rehabilitation status, rehabilitation risk characteristics, and rehabilitation rate data to obtain the reward function value. For example, a penalty value can be calculated based on the target subject's pain feedback score to reduce the probability of the proposed rehabilitation strategy causing pain or other discomfort. A reward value can be calculated based on the target subject's sleep quality to improve sleep quality when implementing the proposed rehabilitation strategy. A penalty value can be calculated based on the target subject's perceived fatigue level to reduce fatigue when implementing the proposed rehabilitation strategy.
[0066] Reference Figure 7As shown, in one embodiment of this application, an optimized rehabilitation strategy is generated through a strategy optimization agent based on optimal model parameters, including: S342. Based on the rehabilitation knowledge graph, perform a fusion correlation analysis between the current rehabilitation strategy and the multi-source physiological data and rehabilitation progress of the target object to obtain the target correlation features of the current rehabilitation strategy.
[0067] S344. Generate rehabilitation adjustment strategies for the target object based on the target association characteristics.
[0068] S346. By using the strategy of optimizing the agent with the optimal model parameters to generate an intermediate rehabilitation strategy, the intermediate rehabilitation strategy and the rehabilitation adjustment strategy are fused to obtain the optimized rehabilitation strategy.
[0069] Specifically, feature extraction is performed on multi-source physiological data of the target object to obtain physiological feature data in natural language form. In some embodiments, judgment conditions or judgment thresholds can be set for the multi-source physiological data to convert it into natural language form or structured parameter form. After the multi-source physiological data is converted, causal reasoning is performed on the multi-source physiological data and rehabilitation progress through a rehabilitation knowledge graph to determine the causal relationship chain causing the rehabilitation progress and obtain the target-related features of the current rehabilitation strategy.
[0070] For example, multi-source physiological data may include the electromyographic amplitude of the middle deltoid muscle, the shoulder flexion angle, and the target subject's subjective feeling score. The subjective feeling score may include pain feedback score and sleep duration. Rehabilitation progress may include the target subject's failure to achieve expected progress in the current rehabilitation cycle, specifically manifested as an inability to increase range of motion. The correlation between the above data could be that sleep duration indicates poor sleep quality, leading to an increased pain feedback score and consequently a decrease in joint mobility; or it could be that the target subject's supraspinatus muscle exhibits delayed activation, causing compensatory overactivation of the deltoid muscle, thus increasing the risk of subacromial impingement. Correlation analysis based on the above data correlations yields target correlation characteristics. These characteristics could be the main reason why the target subject's rehabilitation progress in the current cycle is not meeting expectations, stemming from pain experienced by the target subject, with the increased pain attributed to poor sleep quality. Another possible characteristic is the inability to increase range of motion, primarily due to excessive tension in the hamstrings (antagonist muscles) and insufficient strength in the quadriceps (non-agonist muscles).
[0071] Furthermore, after determining the target-related features, rehabilitation adjustment suggestions with medical interpretation are generated for the target object based on these features, forming an intermediate rehabilitation strategy. It is understood that the intermediate rehabilitation strategy optimizes and adjusts the rehabilitation training strategy from the perspective of causal reasoning, and is more targeted and interpretable compared to strategy search solely through a strategy optimization agent. Through strategy fusion between the intermediate rehabilitation strategy and the rehabilitation adjustment strategy, the intelligence level of rehabilitation strategy prediction and optimization can be effectively improved, thereby further improving the rehabilitation progress of the target object. In some embodiments, the intermediate rehabilitation strategy can be generated by a strategy optimization agent using optimal model parameters, or it can be formulated by the target object's attending physician based on the target-related features. For example, for the aforementioned target-related features, the intermediate rehabilitation strategy could be to add 5 minutes of static hamstring stretching to the target object before training in each rehabilitation cycle. The intermediate rehabilitation strategy could also be to extend the end-of-movement hold time in leg flexion and extension training from 2 seconds to 3 seconds to strengthen the target object's eccentric control ability.
[0072] Reference Figure 8 As shown in one embodiment of this application, the policy optimization agent is trained based on the historical rehabilitation data of the sample objects. The historical rehabilitation data includes the sample objects' historical initial conditions, historical rehabilitation status, historical progress data, and historical risk characteristics. The reward function of the policy optimization agent is obtained in the following manner: S410. Calculate rehabilitation benefits based on historical rehabilitation status and historical initial status to obtain rehabilitation benefit reward items used to incentivize the agent to optimize the strategy and improve the rehabilitation effect of the target object.
[0073] S420. Physiological function reenactment is performed on historical process data using a musculoskeletal model to evaluate rehabilitation efficiency based on historical process data, and rehabilitation efficiency reward items are obtained to incentivize the agent to improve the rehabilitation efficiency of the target object.
[0074] S430. Physiological load is replayed on historical process data through a physiological function model to assess the rehabilitation load on historical process data and obtain a rehabilitation load penalty term used to suppress the physiological load caused to the target object by the strategy optimization agent.
[0075] S440. Analyze the risk index of the sample objects during the rehabilitation process based on historical risk characteristics to obtain a rehabilitation risk penalty term used to suppress the rehabilitation risk caused by the strategy optimization agent to the target object.
[0076] S450. Obtain the reward function based on the rehabilitation benefit reward, rehabilitation energy efficiency reward, rehabilitation load penalty, and rehabilitation risk penalty.
[0077] Specifically, the historical initial situation represents the initial state of the sample object before undergoing rehabilitation training, while the historical rehabilitation situation represents the rehabilitation state of the sample object after undergoing rehabilitation training. By calculating the difference between the two, the overall rehabilitation degree of the sample object after rehabilitation training can be obtained. Based on the overall rehabilitation degree and the total duration of the rehabilitation training process, the rehabilitation efficiency of the rehabilitation training process is calculated to obtain the overall rehabilitation efficiency. Based on the overall rehabilitation degree and overall rehabilitation efficiency, a rehabilitation benefit reward can be obtained to incentivize the strategy optimization agent to improve the rehabilitation degree and efficiency of the target object.
[0078] Furthermore, historical progress data can include multi-source physiological data of the sample subjects in each historical rehabilitation cycle, as well as the rehabilitation training strategies adopted, to describe the rehabilitation training process of the sample subjects. By replaying physiological functions based on historical progress data using a musculoskeletal model, rehabilitation efficiency is assessed for each historical rehabilitation cycle experienced by the sample subjects, resulting in rehabilitation efficiency reward items. These rewards are used to suppress the physiological load imposed on the target subject by the strategy optimization agent.
[0079] In some embodiments, the rehabilitation efficiency reward item may include an efficient training item, a metabolic cost item, and a rehabilitation stimulation item. The efficient training item incentivizes the strategy optimization agent to generate rehabilitation training strategies that require less training and result in greater rehabilitation progress for the target subject. This item can be obtained based on the rehabilitation progress and training volume of the sample subject in each historical rehabilitation cycle; the greater the rehabilitation progress and the less training volume in any historical rehabilitation cycle, the higher the reward value. The metabolic cost item incentivizes the strategy optimization agent to generate rehabilitation training strategies that match the target subject's metabolic capacity. This item can be obtained based on the actual metabolic cost of the sample subject during historical rehabilitation training and the designed metabolic cost of the corresponding rehabilitation training strategy; the closer the actual metabolic cost is to the designed metabolic cost, the higher the reward value. The rehabilitation stimulation item incentivizes the strategy optimization agent to generate rehabilitation training strategies that accurately stimulate the target subject's muscles. This item can be obtained based on the stress stimulation amount and rehabilitation coefficient of the target subject's muscles; for any target muscle, the higher the stress stimulation amount or the higher the rehabilitation coefficient, the higher the reward value.
[0080] Furthermore, by replaying the physiological load using a physiological function model based on historical progress data, the rehabilitation load is assessed for each historical rehabilitation cycle experienced by the sample object, resulting in a rehabilitation load penalty term. This term is used to suppress the physiological load caused to the target object by the strategy optimization agent. In some embodiments, the rehabilitation load penalty term may include a joint contact term and a load recovery term. The joint contact term is used to suppress situations where the rehabilitation training strategies generated by the strategy optimization agent cause joint and muscle damage to the target object. It can be obtained based on the peak joint contact force and safety threshold of the sample object, and a severe penalty is applied when the peak joint contact force exceeds the safety threshold. The load recovery term is used to suppress the strategy optimization agent from generating rehabilitation training strategies that exceed the target object's recovery ability. It can be obtained based on the training volume and recovery ability of the sample object in each historical rehabilitation cycle. Recovery ability may include muscle fatigue recovery ability and autonomic nervous system recovery ability, etc. A penalty value is applied when the training volume exceeds the sample object's recovery ability in any historical rehabilitation cycle.
[0081] Furthermore, historical risk characteristics include the risk index for each historical rehabilitation cycle and the risk change pattern between historical rehabilitation cycles. The risk change pattern represents the fluctuation of the risk index across multiple consecutive historical rehabilitation cycles. Understandably, the rehabilitation risk penalty term includes single-cycle risk terms and cross-cycle risk terms. The single-cycle risk term is used to suppress excessive risk to the target subject within a single rehabilitation cycle and can be obtained from the maximum value of the risk index during historical rehabilitation training; the larger the maximum value of the risk index, the larger the penalty value. The cross-cycle risk term is used to suppress the gradual increase in the level of risk faced by the target subject during rehabilitation training and can be obtained from the fluctuation of the risk index during historical rehabilitation training.
[0082] For example, the reward function can be expressed as: in, For the reward function; Awards for rehabilitation benefits; This is a rehabilitation efficiency award item; This is a penalty item for rehabilitation workload; This refers to the penalty for rehabilitation risks. The reward for rehabilitation benefits can be represented as: in, Regarding the historical recovery situation, This represents the initial historical situation. This refers to the total duration of the historical rehabilitation training process; and As a reward weight, the rehabilitation efficiency reward item can be represented as: in, This represents the rehabilitation progress during the i-th historical rehabilitation cycle. Let i be the training volume for the i-th historical rehabilitation cycle. This represents the total number of historical rehabilitation cycles. For actual metabolic costs, To design metabolic costs; Let J be the stress stimulation amount for the muscles of the j-th rehabilitation subject. Let be the rehabilitation coefficient of the muscles of the j-th rehabilitation subject; The total number of muscles in the rehabilitation subject; , and The reward weight is used. The rehabilitation load penalty term can be represented as: in, The peak value of the joint contact force of the sample object. The safe contact force threshold; The recovery ability of the sample subjects in the k-th historical rehabilitation cycle; and As a reward weight, the rehabilitation risk penalty item can be represented as: in, This refers to the risk index during the historical rehabilitation training process; Indicates the calculation of the degree of fluctuation; and For reward weighting.
[0083] Reference Figure 9 As shown in this embodiment of the application, the rehabilitation training process includes multiple rehabilitation cycles, and the preset risk conditions include a normal threshold range and preset risk characteristics; the progress of rehabilitation is determined to meet the preset risk conditions in the following ways: S302. Conduct a rehabilitation risk assessment on the rehabilitation progress of each of the multiple rehabilitation cycles to obtain the rehabilitation risk index for each of the multiple rehabilitation cycles.
[0084] S304. Extract time-series features based on the rehabilitation risk index of multiple rehabilitation cycles to obtain the risk time-series features of the target object in multiple rehabilitation cycles.
[0085] S306. If the rehabilitation risk index exceeds the normal threshold range and there is similarity between the risk time sequence characteristics and the preset risk characteristics, the rehabilitation progress is determined to meet the preset risk conditions.
[0086] Specifically, based on the rehabilitation progress of the target subject in each rehabilitation cycle, a risk assessment is conducted to determine the risks faced by the target subject in the corresponding rehabilitation cycle, resulting in a rehabilitation risk index. For example, the rehabilitation risk index may include structural injury risk, systemic decompensation risk, and behavioral and functional dysregulation risk. Structural injury risk can be used to assess the mechanical overload and cumulative fatigue damage to tissues such as joints, ligaments, and cartilage. Systemic decompensation risk can be used to assess the depletion of functional reserves and regulatory failure of physiological systems such as the cardiovascular, respiratory, and metabolic systems under rehabilitation load. Behavioral and functional dysregulation risk can be used to assess the risk of reduced efficacy or secondary injury due to worsening movement patterns, pain avoidance behaviors, and decreased adherence by the target subject.
[0087] Furthermore, within multiple consecutive rehabilitation cycles, temporal features are extracted from the rehabilitation risk index to obtain the risk temporal characteristics of the target subject. In some embodiments, temporal feature extraction may include trend feature extraction and fluctuation feature extraction, where trend feature extraction may extract the changing trend of the rehabilitation risk index across multiple rehabilitation cycles, indicating whether the rehabilitation training strategy causes cumulative harm to the target subject, and can be used to determine whether the rehabilitation training strategy is suitable for the target subject. Fluctuation feature extraction may extract the fluctuation characteristics of the rehabilitation risk index across multiple rehabilitation cycles, representing the state stability of the target subject during the rehabilitation training process.
[0088] In some embodiments, rehabilitation risk assessment and time-series feature extraction can be performed using a machine learning model to extract risk features based on predicted rehabilitation outcomes, thereby obtaining rehabilitation risk features of the proposed rehabilitation strategy to represent the health risks that the target individual may face when implementing the proposed rehabilitation strategy. For example, the machine learning model used for risk feature extraction may be a survival analysis model or a time-series prediction model, etc.
[0089] Furthermore, the rehabilitation risk index for each rehabilitation cycle is compared with the normal threshold range, and similarity calculations are performed on the risk time series characteristics and preset risk characteristics to obtain risk assessment results. If the risk assessment results indicate that the rehabilitation risk index for multiple rehabilitation cycles exceeds the normal threshold range, and there is similarity between the risk time series characteristics and some preset risk characteristics, it is determined that the rehabilitation progress meets the preset risk conditions, and rehabilitation strategy prediction and optimization are required.
[0090] Reference Figure 10 As shown in one embodiment of this application, the method further includes: S510. Send the optimized rehabilitation strategy to the attending physician of the target patient, and have the attending physician review the optimized rehabilitation strategy.
[0091] S520. Once the optimized rehabilitation strategy has been approved, the optimized rehabilitation strategy will be shared with the target individuals and their local community hospitals to achieve three-level linkage management of the rehabilitation process for the target individuals.
[0092] Reference Figure 11 As shown, in this embodiment, the three-level linkage management platform includes the attending physician's end, the community hospital's end, and the target object's end. The attending physician's end is connected to the community hospital's end and the target object's end, and is used to issue optimized rehabilitation strategies after review. The community hospital's end is also connected to the target object's end, and is used to supervise the target object's rehabilitation training process and provide medical guidance to the target object when necessary.
[0093] Specifically, after the strategy optimization agent generates an optimized rehabilitation strategy, the strategy is first sent to the target patient's attending physician. The physician reviews the strategy to determine its clinical rationality and medical logic, ensuring the target patient can safely undergo rehabilitation training. If the strategy passes review, it can be adjusted by the attending physician or the strategy optimization agent and re-reviewed until it is approved. After approval, the optimized rehabilitation strategy is shared with the target patient and their local community hospital, enabling the target patient to understand and implement it. Simultaneously, the strategy ensures the community hospital also has access to the target patient's optimized rehabilitation strategy, allowing them to provide rehabilitation guidance and facilitate professional medical support during home rehabilitation, thus improving the effectiveness of home rehabilitation.
[0094] Accordingly, please refer to Figure 12 This application provides a rehabilitation training strategy optimization device based on digital twin and adaptive learning, the device comprising: The twin model construction module 1210 is used to construct a musculoskeletal model and a physiological function model based on the clinical personal information of the target object, and generate a strategy optimization agent for the target object.
[0095] The rehabilitation progress prediction module 1220 is used to acquire multi-source physiological data of the target object during the rehabilitation training process, input the multi-source physiological data into the musculoskeletal model and the physiological function model, and perform coupled simulation prediction through the musculoskeletal model and the physiological function model to obtain the rehabilitation progress of the target object.
[0096] The rehabilitation strategy optimization module 1230 is used to predict and optimize rehabilitation strategies based on multi-source physiological data when the rehabilitation progress meets preset risk conditions. This is achieved through a strategy optimization agent, a musculoskeletal model, and a physiological function model. The optimized rehabilitation strategy is generated for the target object.
[0097] In some alternative implementations, the rehabilitation progress prediction module 1220 includes: The rehabilitation effect prediction unit is used to predict the rehabilitation effect of the target object in any rehabilitation cycle other than the first rehabilitation cycle, based on the movement data of any rehabilitation cycle and the prediction results of the physiological function model in the previous rehabilitation cycle, through the musculoskeletal model.
[0098] The rehabilitation burden prediction unit is used to predict the rehabilitation burden of the target subject based on the movement data of any rehabilitation cycle and the prediction results of the physiological function model in the previous rehabilitation cycle.
[0099] The progress acquisition unit is used to obtain the rehabilitation progress of the target subject within any rehabilitation cycle based on the rehabilitation training effect and rehabilitation burden data.
[0100] In some optional implementations, the rehabilitation strategy optimization module 1230 includes: The coupled simulation prediction unit is used to generate a proposed rehabilitation strategy based on the target object's current rehabilitation strategy. The proposed rehabilitation strategy is input into the musculoskeletal model and the physiological function model. The proposed rehabilitation strategy is then subjected to coupled simulation prediction using the musculoskeletal model and the physiological function model to obtain the predicted rehabilitation status of the target object under the proposed rehabilitation strategy.
[0101] The reward function calculation unit is used to calculate the reward based on the predicted rehabilitation status and the reward function of the strategy optimization agent, and obtain the reward function value of the proposed rehabilitation strategy.
[0102] The agent parameter optimization unit is used to optimize the parameters of the policy optimization agent based on the reward function values of multiple proposed rehabilitation strategies, so as to obtain the optimal model parameters of the policy optimization agent.
[0103] The rehabilitation strategy generation unit is used to generate optimized rehabilitation strategies based on the optimal model parameters through a strategy optimization agent.
[0104] In some optional implementations, the reward function calculation unit includes: The risk feature extraction subunit is used to extract risk features based on the predicted rehabilitation status in each rehabilitation cycle, so as to obtain the rehabilitation risk features of the proposed rehabilitation strategy in multiple rehabilitation cycles.
[0105] The rehabilitation efficiency assessment subunit is used to assess rehabilitation efficiency based on the predicted rehabilitation status and cycle length of multiple rehabilitation cycles, and to obtain rehabilitation rate data of the proposed rehabilitation strategy in multiple rehabilitation cycles.
[0106] The reward calculation subunit is used to calculate the reward based on the predicted rehabilitation status, rehabilitation risk characteristics, rehabilitation rate data, and reward function, and obtain the reward function value.
[0107] In some optional implementations, the reward function calculation unit further includes: The subjective grading assessment subunit is used to obtain the target subject's subjective feelings about rehabilitation; the subjective feelings about rehabilitation are graded and assessed to obtain the target subject's subjective feeling score.
[0108] The reward calculation subunit is used to calculate the reward based on the predicted rehabilitation status, rehabilitation risk characteristics, rehabilitation rate data, subjective feeling score, and reward function, and obtain the reward function value.
[0109] In some optional implementations, the rehabilitation strategy generation unit includes: The fusion association analysis subunit is used to perform fusion association analysis between the current rehabilitation strategy and the target object's multi-source physiological data and rehabilitation progress based on the rehabilitation knowledge graph, so as to obtain the target association features of the current rehabilitation strategy.
[0110] The adjustment strategy generation sub-unit is used to generate rehabilitation adjustment strategies for the target object based on the target association characteristics.
[0111] The strategy fusion subunit is used to generate intermediate rehabilitation strategies by using the strategy optimization agent with optimal model parameters, and to fuse the intermediate rehabilitation strategies and rehabilitation adjustment strategies to obtain the optimized rehabilitation strategy.
[0112] In some alternative implementations, the device further includes a reward function acquisition module, comprising: The rehabilitation benefit calculation unit is used to calculate rehabilitation benefits based on historical rehabilitation status and historical initial status, and to obtain rehabilitation benefit reward items for incentivizing the intelligent agent to improve the rehabilitation effect of the target object.
[0113] The rehabilitation efficiency assessment unit is used to reenact physiological functions from historical process data using a musculoskeletal model, thereby assessing the rehabilitation efficiency of the historical process data and obtaining rehabilitation efficiency reward items to incentivize the agent to optimize the strategy and improve the rehabilitation efficiency of the target object.
[0114] The rehabilitation load assessment unit is used to replay the physiological load on historical process data through a physiological function model, so as to assess the rehabilitation load on the historical process data and obtain a rehabilitation load penalty term for suppressing the physiological load caused by the strategy optimization agent to the target object.
[0115] The risk feature analysis unit is used to analyze the risk index of the sample object during the rehabilitation process based on historical risk features, so as to obtain a rehabilitation risk penalty term used to suppress the rehabilitation risk caused by the strategy optimization agent to the target object.
[0116] The reward function acquisition unit is used to obtain the reward function based on the rehabilitation benefit reward item, rehabilitation energy efficiency reward item, rehabilitation load penalty item, and rehabilitation risk penalty item.
[0117] In some optional implementations, the rehabilitation strategy optimization module 1230 includes a risk condition assessment unit, comprising: The rehabilitation risk assessment subunit is used to assess the rehabilitation progress of multiple rehabilitation cycles and obtain the rehabilitation risk index for each cycle.
[0118] The temporal feature extraction subunit is used to extract temporal features based on the rehabilitation risk index of multiple rehabilitation cycles, so as to obtain the risk temporal features of the target object in multiple rehabilitation cycles.
[0119] The risk feature comparison subunit is used to determine whether the rehabilitation progress meets the preset risk conditions when the rehabilitation risk index exceeds the normal threshold range and there is similarity between the risk time sequence characteristics and the preset risk characteristics.
[0120] In some optional implementations, the device further includes a three-level linkage management module, comprising: The rehabilitation strategy review unit is used to send optimized rehabilitation strategies to the attending physicians of the target patients, who then review the optimized rehabilitation strategies.
[0121] The rehabilitation strategy linkage unit is used to share the optimized rehabilitation strategy with the target individuals and their local community hospitals, once the optimized rehabilitation strategy has been approved, in order to achieve three-level linkage management of the rehabilitation process for the target individuals.
[0122] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0123] The rehabilitation training strategy optimization device based on digital twin and adaptive learning in this embodiment is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0124] Please see Figure 13 , Figure 13 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application, such as... Figure 13As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 13 Take a processor 10 as an example.
[0125] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.
[0126] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.
[0127] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0128] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0129] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.
[0130] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the methods shown in the above embodiments are implemented.
[0131] This application provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method of any embodiment of this application.
[0132] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and all such modifications and variations fall within the scope defined by the appended claims.
[0133] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0134] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0135] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0136] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0137] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0138] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0139] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0140] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0141] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
[0142] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for optimizing rehabilitation training strategies based on digital twins and adaptive learning, characterized in that, The method includes: Based on the clinical personal information of the target subject, a musculoskeletal model and a physiological function model are constructed, and a strategy optimization agent is generated for the target subject; wherein, the physiological function model is used to simulate the changes of physiological variables in the target subject under different exercise conditions; Multi-source physiological data of the target object during rehabilitation training is acquired, and the multi-source physiological data is input into the musculoskeletal model and the physiological function model. Coupled simulation prediction is then performed using the musculoskeletal model and the physiological function model to obtain the rehabilitation progress of the target object. The multi-source physiological data refers to the physiological data generated by the target object during home rehabilitation, including physiological state data, which describes the physiological state of the target object. The rehabilitation progress includes rehabilitation burden data, which is obtained by predicting the rehabilitation burden based on the physiological state data using the physiological function model, representing the physical burden caused by the rehabilitation training strategy to the target object. When the rehabilitation progress meets preset risk conditions, based on the target object's multi-source physiological data from the previous rehabilitation cycle, the strategy optimization agent searches for optimized strategies in the strategy action space and outputs proposed rehabilitation strategies. Using the musculoskeletal model and the physiological function model, based on the target object's multi-source physiological data from the previous rehabilitation cycle, the agent predicts the rehabilitation training effect and rehabilitation burden data of the target object after implementing the proposed rehabilitation strategies, obtaining the predicted rehabilitation status of the target object under the proposed rehabilitation strategies. The agent's reward function is calculated based on the predicted rehabilitation status, obtaining the reward function value of the proposed rehabilitation strategies. Based on multiple proposed rehabilitation strategies and their corresponding reward function values, the agent's parameters are optimized until the parameter optimization termination condition is met, obtaining the optimal model parameters of the agent. Based on the optimal model parameters, the agent generates an optimized rehabilitation strategy for the target object.
2. The method according to claim 1, characterized in that, The rehabilitation training process includes multiple rehabilitation cycles, and the multi-source physiological data includes movement data and physiological data. The acquisition cycle of the physiological data is longer than that of the movement data. The process of obtaining the rehabilitation progress of the target object through coupled simulation prediction using the musculoskeletal model and the physiological function model includes: In any rehabilitation cycle other than the first rehabilitation cycle among the multiple rehabilitation cycles, the rehabilitation effect is predicted by the musculoskeletal model based on the movement data of the rehabilitation cycle and the prediction results of the physiological function model in the previous rehabilitation cycle of the rehabilitation cycle, so as to obtain the rehabilitation training effect of the target object. Based on the movement data of any rehabilitation cycle and the prediction results of the physiological function model in the previous rehabilitation cycle of any rehabilitation cycle, the rehabilitation burden is predicted through the physiological function model to obtain the rehabilitation burden data of the target object. The rehabilitation progress of the target subject in any given rehabilitation cycle is obtained based on the rehabilitation training effect and the rehabilitation burden data.
3. The method according to claim 1, characterized in that, The rehabilitation training process includes multiple rehabilitation cycles, each corresponding to a predicted rehabilitation outcome. The step of calculating the reward function value of the proposed rehabilitation strategy based on the predicted rehabilitation outcome and the reward function of the strategy optimization agent includes: Risk features are extracted based on the predicted rehabilitation status in each rehabilitation cycle to obtain the rehabilitation risk features of the proposed rehabilitation strategy in the multiple rehabilitation cycles. Based on the predicted rehabilitation status and cycle length of each of the multiple rehabilitation cycles, the rehabilitation efficiency is evaluated to obtain the rehabilitation rate data of the proposed rehabilitation strategy in the multiple rehabilitation cycles. The reward is calculated based on the predicted recovery status, the recovery risk characteristics, the recovery rate data, and the reward function to obtain the reward function value.
4. The method according to claim 3, characterized in that, The step of calculating the reward based on the predicted recovery status, the recovery risk characteristics, the recovery rate data, and the reward function to obtain the reward function value includes: Obtain the target subject's subjective feelings about rehabilitation; conduct a graded assessment of the subjective feelings about rehabilitation to obtain the target subject's subjective feeling score; The reward is calculated based on the predicted rehabilitation status, the rehabilitation risk characteristics, the rehabilitation rate data, the subjective feeling score, and the reward function to obtain the reward function value.
5. The method according to claim 1, characterized in that, The step of generating the optimized rehabilitation strategy based on the optimal model parameters and through the strategy optimization agent includes: Based on the rehabilitation knowledge graph, a fusion correlation analysis is performed between the current rehabilitation strategy and the multi-source physiological data and rehabilitation progress of the target object to obtain the target correlation features of the current rehabilitation strategy. Generate a rehabilitation adjustment strategy for the target object based on the target association features; The intermediate rehabilitation strategy is generated by the agent using the strategy optimization of the optimal model parameters, and the intermediate rehabilitation strategy and the rehabilitation adjustment strategy are fused to obtain the optimized rehabilitation strategy.
6. The method according to claim 1, characterized in that, The strategy optimization agent is trained based on the historical rehabilitation data of the sample objects. The historical rehabilitation data includes the historical initial situation, historical rehabilitation status, historical progress data, and historical risk characteristics of the sample objects. The reward function of the policy-optimizing agent is obtained in the following way: Based on the historical rehabilitation situation and the historical initial situation, the rehabilitation benefits are calculated to obtain a rehabilitation benefit reward item used to incentivize the strategy optimization agent to improve the rehabilitation effect of the target object; The musculoskeletal model is used to reenact the physiological functions of the historical process data in order to evaluate the rehabilitation efficiency of the historical process data and obtain rehabilitation efficiency reward items to incentivize the strategy optimization agent to improve the rehabilitation efficiency of the target object. The physiological load is replayed on the historical process data by the physiological function model to assess the rehabilitation load of the historical process data and obtain a rehabilitation load penalty term for suppressing the physiological load caused by the policy optimization agent to the target object. Based on the historical risk characteristics, the risk index of the sample object during the rehabilitation process is analyzed to obtain a rehabilitation risk penalty term used to suppress the rehabilitation risk caused by the strategy optimization agent to the target object; The reward function is obtained based on the rehabilitation benefit reward item, the rehabilitation efficiency reward item, the rehabilitation load penalty item, and the rehabilitation risk penalty item.
7. The method according to claim 1, characterized in that, The rehabilitation training process includes multiple rehabilitation cycles, and the preset risk conditions include normal threshold ranges and preset risk characteristics. The following methods are used to determine whether the recovery progress meets the preset risk conditions: A rehabilitation risk assessment is performed on the rehabilitation progress of each of the multiple rehabilitation cycles to obtain the rehabilitation risk index for each of the multiple rehabilitation cycles. Based on the rehabilitation risk index of the multiple rehabilitation cycles, temporal features are extracted to obtain the risk temporal features of the target object in the multiple rehabilitation cycles; If the rehabilitation risk index exceeds the normal threshold range and there is similarity between the risk time series characteristics and the preset risk characteristics, the rehabilitation progress is determined to meet the preset risk conditions.
8. The method according to claim 1, characterized in that, The method further includes: The optimized rehabilitation strategy is sent to the attending physician of the target patient, who then reviews the optimized rehabilitation strategy. Once the optimized rehabilitation strategy is approved, it will be shared with the target individual and the community hospital in their place of residence to achieve three-level linkage management of the rehabilitation process for the target individual.
9. A rehabilitation training strategy optimization device based on digital twin and adaptive learning, characterized in that, The device includes: The twin model construction module is used to construct a musculoskeletal model and a physiological function model based on the clinical personal information of the target object, and generate a strategy optimization agent for the target object; wherein, the physiological function model is used to simulate the changes of physiological variables in the target object under different exercise conditions; The rehabilitation progress prediction module is used to acquire multi-source physiological data of the target object during rehabilitation training, input the multi-source physiological data into the musculoskeletal model and the physiological function model, and perform coupled simulation prediction through the musculoskeletal model and the physiological function model to obtain the rehabilitation progress of the target object; wherein, the multi-source physiological data is physiological data generated by the target object during home rehabilitation, including physiological state data, which is used to describe the physiological state of the target object; the rehabilitation progress includes rehabilitation burden data, which is obtained by the physiological function model based on the physiological state data to predict the rehabilitation burden, representing the physical burden caused by the rehabilitation training strategy to the target object; The rehabilitation strategy optimization module is used to, when the rehabilitation progress meets preset risk conditions, search for optimized strategies in the strategy action space through the strategy optimization agent based on the target object's multi-source physiological data from the previous rehabilitation cycle, and output proposed rehabilitation strategies; using the musculoskeletal model and the physiological function model, based on the target object's multi-source physiological data from the previous rehabilitation cycle, predict the rehabilitation training effect and rehabilitation burden data of the target object after implementing the proposed rehabilitation strategy, and obtain the predicted rehabilitation status of the target object under the proposed rehabilitation strategy; calculate the reward function of the strategy optimization agent based on the predicted rehabilitation status, and obtain the reward function value of the proposed rehabilitation strategy; optimize the parameters of the strategy optimization agent based on multiple proposed rehabilitation strategies and their corresponding reward function values until the parameter optimization termination condition is met, and obtain the optimal model parameters of the strategy optimization agent; and generate an optimized rehabilitation strategy for the target object through the strategy optimization agent based on the optimal model parameters.