Multimodal rehabilitation decision dynamic adjustment method and device based on reinforcement learning, equipment and medium

Through a multimodal rehabilitation decision-making dynamic adjustment method based on reinforcement learning, the rehabilitation treatment plan is optimized by combining physiological, emotional and environmental information, which solves the problem that individual differences are not taken into account in existing technologies and realizes personalized and efficient rehabilitation treatment.

CN120708913APending Publication Date: 2025-09-26AFFILIATED HOSPITAL OF NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510916775.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing rehabilitation decision support methods fail to effectively consider individual differences among patients, resulting in a lack of personalization and dynamic adjustment of rehabilitation programs.

Method used

A multimodal rehabilitation decision-making dynamic adjustment method based on reinforcement learning is adopted. The current state is obtained through physiological information, emotional feedback and environmental information. The greedy strategy and Q-learning algorithm are used to select the treatment plan, and the Q value is updated according to patient feedback to optimize the treatment path.

Benefits of technology

It achieves personalized and adaptive rehabilitation treatment, improves treatment efficiency and patient satisfaction, and ensures that patients always receive the most effective treatment plan.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708913A_ABST
    Figure CN120708913A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode rehabilitation decision dynamic adjustment method and device based on reinforcement learning, equipment and a medium, and the method comprises the steps: obtaining a current state of a first target according to the physiological information, emotion subjective feedback information and environment information of the first target; based on a greedy strategy, an action corresponding to the current state is determined from a preset treatment scheme knowledge base, and the preset treatment scheme knowledge base comprises a plurality of rehabilitation treatment schemes; determining a reward corresponding to the current state according to the current state of the first target and the corresponding action; and based on a Q-learning algorithm and according to the reward corresponding to the current state, updating the Q value, and determining an expected effect of the action corresponding to the current state. The invention belongs to the field of strategy optimization. The treatment action is selected from the preset treatment scheme knowledge base by using the greedy strategy, and the calculation reward is fed back to update the Q value, so that the optimal treatment path is gradually converged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of strategy optimization, and in particular to a method, device, equipment and medium for dynamic adjustment of multimodal rehabilitation decision-making based on reinforcement learning. Background Art

[0002] Rehabilitation decision support refers to the use of information technology, data analysis, and clinical knowledge to assist medical professionals in making more informed rehabilitation treatment decisions.

[0003] Currently, several rehabilitation decision support methods based on data analysis exist. These methods typically rely on standardized processes to generate rehabilitation plans, rarely considering individual patient differences and failing to dynamically adjust plans. The application of reinforcement learning in rehabilitation is still in its early stages. While some studies have attempted to apply reinforcement learning to disease treatment selection and medical diagnosis, there are still relatively few reinforcement learning programs specifically designed for personalized rehabilitation. Therefore, how to evaluate targeted plans based on individual differences and select the most appropriate ones is an urgent issue. Summary of the Invention

[0004] The present invention solves the technical problem of how to provide more suitable rehabilitation plans based on the individual differences of targets in the prior art by providing a multimodal rehabilitation decision-making dynamic adjustment method, device, equipment and medium based on reinforcement learning, and achieves the technical effect of providing more suitable rehabilitation plans based on the individual differences of targets.

[0005] In a first aspect, the present invention provides a multimodal rehabilitation decision-making dynamic adjustment method based on reinforcement learning, the method comprising: Obtaining a current state of the first target based on the first target's physiological information, emotional subjective feedback information, and environmental information; Based on the greedy strategy, the action corresponding to the current state is determined from the preset treatment plan knowledge base, which includes several rehabilitation treatment plans; Determine a reward corresponding to the current state based on the current state of the first target and the corresponding action; Based on the Q-learning algorithm and according to the reward corresponding to the current state, the Q value is updated and the expected effect of the action corresponding to the current state is determined.

[0006] Furthermore, based on the greedy strategy, the action corresponding to the current state is determined from the preset treatment plan knowledge base, including:

[0007] in, For the The actions in the steps, is the greedy probability, For the The status of the step, For the Execute all actions in the step and state The maximum value obtained when value, To select the maximum The action corresponding to the value is the Actions in steps.

[0008] Furthermore, according to the current state of the first target and the corresponding action, a reward corresponding to the current state is determined, including:

[0009] in, For the The reward in the step, i.e. the state Execute an action Rewards, to All are preset weights. for The heart rate difference between the first step and the last step, for The difference in muscle strength between the first step and the last step, for The difference in pain between the first step and the last step, for The difference between the blood oxygen saturation of the step and the previous step, for The weight difference between the step and the previous step, for The temperature difference between the first step and the last step, for The blood pressure difference between the step and the previous step, for The difference in respiratory rate between the first step and the last step, for The difference between the emotional state of the step and the previous step, for The difference between the ambient temperature of the next step and the previous step, for The difference in range of motion between the next step and the previous step, for The difference in fatigue between the first step and the last step, for The difference in ambient humidity between the first step and the previous step, for The height difference between the next step and the previous step.

[0010] Furthermore, based on the Q-learning algorithm and the reward corresponding to the current state, the Q value is updated, including:

[0011] in, for Steps value, for The learning rate of the step, for Automatic discount factor for steps, for Execute all actions in the step The maximum after value.

[0012] Furthermore, the learning rate includes:

[0013] in, for The learning rate of the step, is the first preset adjustment coefficient, for The error between the step and the previous step.

[0014] Furthermore, errors include:

[0015] in, Status Execute an action rewards.

[0016] Furthermore, ,include:

[0017] in, for The long-term rewards of steps is the second preset adjustment coefficient; in, include:

[0018] in, For the Rewards in steps.

[0019] In a second aspect, the present invention provides a multimodal rehabilitation decision-making dynamic adjustment device based on reinforcement learning, the device comprising: an acquisition module, configured to obtain a current state of the first target based on the first target's physiological information, emotional subjective feedback information, and environmental information; The action module is used to determine the action corresponding to the current state from the preset treatment plan knowledge base based on a greedy strategy. The preset treatment plan knowledge base includes several rehabilitation treatment plans; A reward module, configured to determine a reward corresponding to the current state based on the current state of the first target and the corresponding action; The expected estimation module is used to update the Q value based on the Q-learning algorithm and the reward corresponding to the current state, and determine the expected effect of the action corresponding to the current state.

[0020] In a third aspect, the present invention provides an electronic device, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to execute to implement the multimodal rehabilitation decision dynamic adjustment method based on reinforcement learning as provided in the first aspect.

[0021] In a fourth aspect, the present invention provides a non-temporary computer-readable storage medium, which, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the multimodal rehabilitation decision dynamic adjustment method based on reinforcement learning as provided in the first aspect.

[0022] One or more technical solutions provided in the present invention have at least the following technical effects or advantages: By introducing comprehensive physiological data, multi-task reward functions, automated learning rates, and discount factor adjustment mechanisms, the present invention can more accurately evaluate the effectiveness of each treatment plan, thereby optimizing the rehabilitation treatment plan for the first goal and improving rehabilitation efficiency and first goal satisfaction.

[0023] This invention is based on the Q-learning algorithm and uses a greedy strategy to select treatment actions from a knowledge base of pre-set treatment plans. It then calculates rewards based on patient feedback to update the Q value, gradually converging to the optimal treatment path. The adjustment and optimization mechanism provided by this invention not only improves treatment efficiency and effectiveness, but also reduces risks, enabling personalized, adaptive rehabilitation treatment and ensuring that patients always receive the most effective treatment plan. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0025] Figure 1 A schematic diagram of the process of the multimodal rehabilitation decision-making dynamic adjustment method based on reinforcement learning provided by the present invention; Figure 2 This is a structural diagram of the multimodal rehabilitation decision-making dynamic adjustment device based on reinforcement learning provided by the present invention. DETAILED DESCRIPTION

[0026] The embodiments of the present invention solve the technical problem of how to provide more suitable rehabilitation plans for individual differences of targets in the prior art by providing a multimodal rehabilitation decision-making dynamic adjustment method based on reinforcement learning.

[0027] The technical solution of the present invention is to solve the above technical problems, and the overall idea is as follows: A multimodal rehabilitation decision-making dynamic adjustment method based on reinforcement learning includes: obtaining the current state of the first target based on the first target's physiological information, emotional subjective feedback information, and environmental information; determining the action corresponding to the current state from a preset treatment plan knowledge base based on a greedy strategy, wherein the preset treatment plan knowledge base includes several rehabilitation treatment plans; determining the reward corresponding to the current state based on the current state of the first target and the corresponding action; and updating the Q value based on the Q-learning algorithm and the reward corresponding to the current state, and determining the expected effect of the action corresponding to the current state.

[0028] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0029] First, the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. Furthermore, the character " / " in this document generally indicates an "or" relationship between the associated objects.

[0030] The present invention provides Figure 1 The multimodal rehabilitation decision dynamic adjustment method based on reinforcement learning shown includes steps S11-S14: Regarding step S11, the current state of the first target is obtained according to the physiological information, emotional subjective feedback information and environmental information of the first target.

[0031] In order to accurately reflect the recovery status of the first target, the first target's physiological information, emotional subjective feedback information, and environmental information can be collected as comprehensively as possible. The physiological information specifically includes: Heart rate: the number of heartbeats per minute, reflecting the cardiac load and exercise intensity of the first target; Blood pressure: including systolic and diastolic blood pressure, reflecting heart function and vascular health; Body temperature: reflecting the changes in the first target's body temperature, which may be related to health conditions such as infection; Respiratory rate: the number of breaths per minute, showing the first target's respiratory condition; Blood oxygen saturation: the proportion of oxygen content in the blood, reflecting lung function; Muscle strength: assessing muscle recovery through muscle strength tests (such as grip strength tests); Range of motion: joint mobility, especially the first target during the rehabilitation period, which often needs to be measured; Weight: affects exercise intensity and rehabilitation progress; Height: used to calculate health indicators such as body mass index; Emotional subjective feedback information includes: Pain level: the first target's self-rating of the current pain; Fatigue: the first target's subjective sense of fatigue; Emotional Status: the first target's emotional state (such as anxiety, depression, etc.) is obtained through questionnaires or facial expression recognition.

[0032] Environmental information includes: Indoor Temperature: The temperature of the indoor environment may affect the rehabilitation effect; Indoor Humidity: Indoor humidity may affect breathing and comfort.

[0033] The collected multimodal data can be preprocessed by denoising, standardizing, filling missing values, and normalizing. Regarding step S12, based on a greedy strategy, an action corresponding to the current state is determined from a preset treatment plan knowledge base, where the preset treatment plan knowledge base includes several rehabilitation treatment plans.

[0034] The greedy strategy makes the currently optimal choice in each step, hoping that such a choice can lead to the global optimal solution. The greedy strategy is suitable for problems where the global optimal solution can be constructed from the local optimal solution.

[0035] Based on the greedy strategy, the action corresponding to the current state is determined from the preset treatment plan knowledge base, including:

[0036] in, For the The actions in the steps, is the greedy probability, For the The status of the step, For the Execute all actions in the step and state The maximum value obtained when value, To select the maximum The action corresponding to the value is the Actions in a step. Actions refer to treatment plans selected from the preset treatment plan knowledge base.

[0037] For example, For low-intensity exercise, For moderate intensity exercise, For high-intensity exercise, For physical therapy.

[0038] Preset treatment plan knowledge base, including: treatment plan, applicable conditions, expected effects, side effects and risks.

[0039] Treatment options: Treatment options include low-intensity exercise, moderate-intensity exercise, high-intensity exercise, and physical therapy.

[0040] Applicable conditions: Each treatment plan is applicable to a specific first-goal group. For example, low-intensity exercise is suitable for the first goal of initial rehabilitation, while high-intensity exercise is suitable for the first goal of better muscle strength recovery.

[0041] Expected results: Physiological indicators that can be improved by each treatment plan, such as heart rate, muscle strength, range of motion, etc.

[0042] Side effects and risks: These include possible adverse reactions or risks associated with each treatment option, such as muscle strain from excessive exercise.

[0043] Regarding step S13, according to the current state of the first target and the corresponding action, a reward corresponding to the current state is determined.

[0044] Specifically include:

[0045] in, For the The reward in the step, i.e. the state Execute an action Rewards, to All are preset weights. for The heart rate difference between the first step and the last step, for The difference in muscle strength between the first step and the last step, for The difference in pain between the first step and the last step, for The difference between the blood oxygen saturation of the step and the previous step, for The weight difference between the step and the previous step, for The temperature difference between the first step and the last step, for The blood pressure difference between the step and the previous step, for The difference in respiratory rate between the first step and the last step, for The difference between the emotional state of the step and the previous step, for The difference between the ambient temperature of the next step and the previous step, for The difference in range of motion between the next step and the previous step, for The difference in fatigue between the first step and the last step, for The difference in ambient humidity between the first step and the previous step, for The height difference between the next step and the previous step.

[0046] After each action is performed, the reward can be calculated based on the feedback of the first goal. For example, in a certain calculation, =17. If t is 1, then the initial It can be recorded as a preset value, such as 5.

[0047] Regarding step S14, based on the Q-learning algorithm and according to the reward corresponding to the current state, the Q value is updated, and the expected effect of the action corresponding to the current state is determined.

[0048] Based on the Q-learning algorithm and the reward corresponding to the current state, the Q value is updated, including:

[0049] in, for Steps value, for The learning rate of the step, for Automatic discount factor for steps, for Execute all actions in the step The maximum after value.

[0050] Among them, the learning rate includes:

[0051] in, for The learning rate of the step, is the first preset adjustment coefficient, for The error between the step and the previous step.

[0052] Regarding errors, including:

[0053] in, Status Execute an action rewards.

[0054] about ,include:

[0055] in, for The long-term rewards of steps is the second preset adjustment coefficient; in, include:

[0056] in, For the Rewards in steps.

[0057] It should be noted that the Q value represents the long-term expected reward of choosing an action in the current state. If the Q value increases significantly, it means that the action is effective in the current state; if the Q value decreases, it means that the treatment action is not effective.

[0058] The treatment plan can be dynamically adjusted according to the change of Q value. For the treatment plan with increased Q value, this plan can be given priority in subsequent treatment.

[0059] Specifically, after each Q-value update, treatment plans with higher Q-values ​​can be prioritized, gradually guiding the primary target toward the optimal recovery path. As learning progresses, treatment plans can be continuously adjusted based on the primary target's recovery progress and feedback, ensuring that the primary target always receives the most effective treatment.

[0060] In summary, the present invention provides a multimodal rehabilitation decision-making dynamic adjustment method based on reinforcement learning. The method includes: obtaining the current state of the first target based on the physiological information, emotional subjective feedback information and environmental information of the first target; determining the action corresponding to the current state from a preset treatment plan knowledge base based on a greedy strategy, and the preset treatment plan knowledge base includes several rehabilitation treatment plans; determining the reward corresponding to the current state based on the current state of the first target and the corresponding action; updating the Q value based on the Q-learning algorithm and the reward corresponding to the current state, and determining the expected effect of the action corresponding to the current state. By introducing comprehensive physiological data, multi-task reward functions, automated learning rates and discount factor adjustment mechanisms, the present invention can more accurately evaluate the effect of each treatment plan, thereby optimizing the rehabilitation treatment plan of the first target, improving rehabilitation efficiency and satisfaction of the first target. Based on the Q-learning algorithm, and using a greedy strategy, treatment actions are selected from the preset treatment plan knowledge base, and rewards are calculated based on patient feedback to update the Q value, thereby gradually converging to the optimal treatment path. The adjustment and optimization mechanism provided by the present invention not only improves treatment efficiency and effectiveness, but also reduces risks, realizes personalized and adaptive rehabilitation treatment, and ensures that patients always receive the most effective treatment plan.

[0061] Based on the same inventive concept, the present invention provides Figure 2 The multimodal rehabilitation decision dynamic adjustment device based on reinforcement learning shown includes: An acquisition module 21 is configured to obtain a current state of the first target based on the first target's physiological information, emotional subjective feedback information, and environmental information; An action module 22 is configured to determine an action corresponding to the current state from a preset treatment plan knowledge base based on a greedy strategy, wherein the preset treatment plan knowledge base includes a plurality of rehabilitation treatment plans; A reward module 23 is configured to determine a reward corresponding to the current state based on the current state of the first target and the corresponding action; The expectation estimation module 24 is used to update the Q value based on the Q-learning algorithm and the reward corresponding to the current state, and determine the expected effect of the action corresponding to the current state.

[0062] Based on the same inventive concept, the present invention further provides an electronic device, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to execute to implement the multimodal rehabilitation decision dynamic adjustment method based on reinforcement learning as provided above.

[0063] Based on the same inventive concept, the present invention also provides a non-temporary computer-readable storage medium. When the instructions in the storage medium are executed by the processor of an electronic device, the electronic device is able to implement the multimodal rehabilitation decision dynamic adjustment method based on reinforcement learning as provided above.

[0064] Since the electronic device described in this embodiment is an electronic device used to implement the information processing method in the embodiment of the present invention, based on the information processing method described in the embodiment of the present invention, those skilled in the art will be able to understand the specific implementation of the electronic device of this embodiment and its various variations. Therefore, how the electronic device implements the method in the embodiment of the present invention will not be described in detail here. As long as the electronic device used by those skilled in the art to implement the information processing method in the embodiment of the present invention falls within the scope of protection of the present invention.

[0065] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0066] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0067] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0068] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0069] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0070] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A multimodal rehabilitation decision-making dynamic adjustment method based on reinforcement learning, characterized by: The method comprises: Obtaining a current state of the first target based on the first target's physiological information, emotional subjective feedback information, and environmental information; Based on the greedy strategy, the action corresponding to the current state is determined from the preset treatment plan knowledge base, which includes several rehabilitation treatment plans; Determine a reward corresponding to the current state based on the current state of the first target and the corresponding action; Based on the Q-learning algorithm and according to the reward corresponding to the current state, the Q value is updated and the expected effect of the action corresponding to the current state is determined.

2. The multimodal rehabilitation decision dynamic adjustment method based on reinforcement learning according to claim 1, characterized in that: Based on the greedy strategy, the action corresponding to the current state is determined from the preset treatment plan knowledge base, including: in, For the The actions in the steps, is the greedy probability, For the The status of the step, For the Execute all actions in the step and state The maximum value obtained when value, To select the maximum The action corresponding to the value is the Actions in steps.

3. The multimodal rehabilitation decision dynamic adjustment method based on reinforcement learning according to claim 2, characterized in that: According to the current state of the first goal and the corresponding action, the reward corresponding to the current state is determined, including: in, For the The reward in the step, i.e. the state Execute an action Rewards, to All are preset weights. for The heart rate difference between the first step and the last step, for The difference in muscle strength between the first step and the last step, for The difference in pain between the first step and the last step, for The difference between the blood oxygen saturation of the step and the previous step, for The weight difference between the step and the previous step, for The temperature difference between the first step and the last step, for The blood pressure difference between the step and the previous step, for The difference in respiratory rate between the first step and the last step, for The difference between the emotional state of the step and the previous step, for The difference between the ambient temperature of the next step and the previous step, for The difference in range of motion between the next step and the previous step, for The difference in fatigue between the first step and the last step, for The difference in ambient humidity between the first step and the previous step, for The height difference between the next step and the previous step.

4. The multimodal rehabilitation decision dynamic adjustment method based on reinforcement learning according to claim 3, characterized in that: Based on the Q-learning algorithm and the reward corresponding to the current state, the Q value is updated, including: in, for Steps value, for The learning rate of the step, for Automatic discount factor for steps, for Execute all actions in the step The maximum after value.

5. The multimodal rehabilitation decision dynamic adjustment method based on reinforcement learning according to claim 4, characterized in that: Learning rate, including: in, for The learning rate of the step, is the first preset adjustment coefficient, for The error between the step and the previous step.

6. The multimodal rehabilitation decision dynamic adjustment method based on reinforcement learning according to claim 5, characterized in that: Errors, including: in, Status Execute an action rewards.

7. The multimodal rehabilitation decision dynamic adjustment method based on reinforcement learning according to claim 5, characterized in that: ,include: in, for The long-term rewards of steps is the second preset adjustment coefficient; in, include: in, For the Rewards in steps.

8. A multimodal rehabilitation decision-making dynamic adjustment device based on reinforcement learning, characterized in that: The device comprises: an acquisition module, configured to obtain a current state of the first target based on the first target's physiological information, emotional subjective feedback information, and environmental information; The action module is used to determine the action corresponding to the current state from the preset treatment plan knowledge base based on a greedy strategy. The preset treatment plan knowledge base includes several rehabilitation treatment plans; A reward module, configured to determine a reward corresponding to the current state based on the current state of the first target and the corresponding action; The expected estimation module is used to update the Q value based on the Q-learning algorithm and the reward corresponding to the current state, and determine the expected effect of the action corresponding to the current state.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute to implement the multimodal rehabilitation decision dynamic adjustment method based on reinforcement learning as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to implement the multimodal rehabilitation decision dynamic adjustment method based on reinforcement learning as described in any one of claims 1 to 7.