Intelligent motion intervention follow-up visit method combining MLP and EnvelopeQ-Learning

By combining intelligent sports intervention follow-up methods with MLP and Envelope_Q-Learning, the follow-up strategy is dynamically adjusted, which solves the problem that traditional follow-up mode is difficult to achieve personalized management, and significantly improves the patient's exercise compliance and self-efficacy.

CN120183755AActive Publication Date: 2025-06-20ANHUI PROVINCIAL HOSPITAL
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510651096.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-20
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The traditional sports intervention follow-up model is difficult to achieve personalized and precise management, resulting in low patient compliance and limited improvement in self-efficacy, affecting the intervention effect.

Method used

Combining the intelligent exercise intervention follow-up method of MLP and Envelope_Q-Learning, we set reward and punishment rules through reinforcement learning, and dynamically adjust the follow-up strategy based on the real-time assessment of patients' exercise compliance and self-efficacy.

Benefits of technology

In-depth analysis of individual differences between patients, accurately analyze exercise compliance and self-efficacy status, output appropriate follow-up strategies, improve patient exercise compliance and self-efficacy, and reduce the follow-up planning burden of medical staff.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183755A_ABST
    Figure CN120183755A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent motion intervention follow-up visit method combining MLP and EnvelopeQ-Learning. The intelligent motion intervention follow-up visit method comprises a follow-up visit strategy output method for performing follow-up visit strategy output according to target input and a parameter updating method for performing feedback adjustment on the follow-up visit strategy output method according to a target execution result after the follow-up visit strategy is executed. According to the method, reward and punishment rules are set by using reinforcement learning, and the follow-up visit strategy is dynamically adjusted according to the real-time evaluation of the motion compliance and self-efficiency of the patient. A Q value function is updated and optimized by means of an EnvelopeQ-Learning algorithm, an MLP approximation value function is utilized, the state of a patient is input, an action Q value is output, an optimal action is selected from various follow-up visit modes such as telephone, short message and face-to-face, a personalized exercise follow-up visit strategy is formed, and the exercise compliance and self-efficiency of the patient are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sports intervention follow-up methods, and particularly to an intelligent sports intervention follow-up method combining MLP and Envelope_Q-Learning. Background Art

[0002] With the improvement of people's health awareness, the importance of sports intervention in the fields of rehabilitation treatment, chronic disease management, and health promotion has become increasingly prominent. Through a reasonable exercise plan, it can effectively improve the physical function of patients, enhance the quality of life, and assist in disease treatment. However, the traditional sports intervention model faces many challenges in the follow-up link, making it difficult to achieve personalized and precise management, resulting in low sports compliance of patients and limited improvement in sports self-efficacy, ultimately affecting the intervention effect.

[0003] In past practices, conventional sports follow-up methods were often relatively single, mostly using fixed-cycle telephone follow-up or outpatient follow-up, and the follow-up frequency was also relatively fixed. For example, during the baseline period (1 - 4 weeks), there were 2 telephone follow-ups per week, during the intensive period (5 - 12 weeks), there was 1 follow-up per week, and during the consolidation period (13 - 24 weeks), there was 1 telephone follow-up every two weeks. This unified model fails to fully consider the individual differences of patients. Different patients are at different levels in terms of sports compliance and self-efficacy and have different needs. For example, patients with low compliance and poor self-efficacy may require more frequent and targeted supervision and guidance, while patients with high sports enthusiasm and strong self-management ability may be bored with overly frequent follow-up and even develop resistance, which is instead not conducive to long-term adherence to exercise.

[0004] Facing the dilemma of the traditional follow-up model, the patent application number is: 202410724948.5, which discloses a model-based intelligent follow-up system and follow-up method that uses a pre-trained model to generate personalized follow-up questionnaires, improving the pertinence and efficiency of follow-up and reducing manual operations, but lacking sufficient attention to real-time feedback during patient follow-up; the patent application number is: 202410417642.5, which discloses an Internet-based intelligent patient care follow-up system and method that uses deep learning technology to analyze patients' rehabilitation videos and condition information to provide intelligent sports guidance. However, it does not attach importance to the impact of differences in patients' sports compliance and self-efficacy on the follow-up effect. Summary of the Invention

[0005] The present invention aims to at least partly solve one of the technical problems in the related technologies. To this end, an object of the present invention is to propose an intelligent sports intervention follow-up method combining MLP and Envelope_Q-Learning, which uses reinforcement learning to set reward and punishment rules and dynamically adjusts the follow-up strategy based on the real-time evaluation of patients' sports compliance and self-efficacy.

[0006] An intelligent exercise intervention follow-up method combining MLP and Envelope_Q-Learning proposed according to the present invention includes a follow-up strategy output method for outputting a follow-up strategy based on a target state and a parameter update method for feedback adjustment of the follow-up strategy output method according to the target execution result after the follow-up strategy is executed.

[0007] The steps of the follow-up strategy output method are as follows:

[0008] S1: Input the target state into the MLP network, and the MLP network outputs the output value Q of each follow-up strategy according to the weight W and the bias b.

[0009] S2: Use the -greedy strategy to select the follow-up strategy with the largest output value Q as the target follow-up strategy for output.

[0010] S3: After the target follow-up strategy is executed, collect and store the target execution result.

[0011] S4: Apply the target execution result collected and stored in step S3, and calculate the reward and punishment score according to the preset reward and punishment strategy.

[0012] S5: Input the reward and punishment score into the Envelope_Q-Learning model for training to obtain the final Q value.

[0013] The steps of the parameter update method are as follows:

[0014] A1: Apply the final Q value obtained in step S5 to calculate the loss function between the final Q value and the Q value obtained by the MLP ;

[0015] A2: Update the weight W and the bias b in the MLP network by the gradient descent method, so as to update the Q value obtained by the MLP.

[0016] A3: Judge whether the preset condition is reached. If not, execute step A1. If so, terminate the update of the weight W and the bias b, and obtain the final weight and bias.

[0017] Apply the final weight W and bias b obtained in step A3 to replace the weight W and bias b in the MLP network in step S1 to complete the parameter update.

[0018] Preferably, in step S1:

[0019] The exercise compliance and exercise self-efficacy of the patient this week are used to form a two-dimensional grid chart of the patient's exercise status through statistical charts. The exercise compliance and exercise self-efficacy of the patient correspond one-to-one with the position matrix formed by the two-dimensional grid chart of the patient's exercise status. The target input in step S1 is the position matrix reflecting the exercise compliance and exercise self-efficacy of the patient this week;

[0020] In step S1, the MLP network includes an input layer, a hidden layer, and an output layer, where:

[0021] The input layer contains two neurons, and the two neurons correspond to the position matrix coordinates of the exercise compliance and exercise self-efficacy of the patient this week respectively;

[0022] The hidden layer contains five neurons and uses the ReLU activation function;

[0023] The output layer contains eight neurons, which correspond to eight follow-up strategies respectively. Each neuron outputs an output value Q, representing the expected cumulative reward for implementing the corresponding follow-up strategy. The eight follow-up strategies include weekly phone follow-up, weekly text message follow-up, weekly face-to-face follow-up, phone follow-up every three days, text message follow-up every three days, phone follow-up every two weeks, text message follow-up every two weeks, and face-to-face follow-up every two weeks.

[0024] Preferably, in step S2:

[0025] Adopt a linear decreasing strategy to gradually reduce the value:

[0026]

[0027] Among them, is the initial value, is the final value, t is the number of steps of the current training, is the total number of training steps.

[0028] Preferably, the takes a value of 0.9, takes a value of 0.1, takes a value of 1000.

[0029] Preferably, in step S4:

[0030] The exercise compliance is divided into five grades from high to low according to the exercise compliance rate, and the exercise self-efficacy is divided into three grades from high to low by exercise efficacy, namely: excellent, good, and slightly poor;

[0031] Compliance improvement reward: If the patient's exercise compliance level this week is one level higher than last week, the reward value is +2; if it is two levels higher, the reward value is +4; if it is three levels higher, the reward value is +6; if it is four levels higher, the reward value is +8; otherwise, the value is 0.

[0032] Compliance decline penalty: If the patient's exercise compliance level this week is one level lower than last week, the reward value is -2; if it is two levels lower, the reward value is -4; if it is three levels lower, the reward value is -6; if it is four levels lower, the reward value is -8; otherwise, the value is 0.

[0033] Exercise self-efficacy reward: If the patient's exercise self-efficacy score improves from "slightly poor" to "good", the reward value is +3; if it improves from "good" to "excellent", the reward value is +5; if the self-efficacy score reaches "excellent" and remains within this range for at least one week, the reward value is +8. Otherwise, the value is 0.

[0034] Exercise self-efficacy penalty: If the patient's exercise self-efficacy declines from "good" to "slightly poor", the penalty value is -3; if it declines from "excellent" to "good", the penalty value is -5; if the self-efficacy score drops to the "slightly poor" range and remains there for at least one week, the penalty value is -8. Otherwise, the value is 0.

[0035] Preferably, in step S5:

[0036] The calculation formula for updating the target Q value in the Envelope_Q-Learning model is:

[0037]

[0038] Where, is the learning rate, is the discount factor, is the current reward, represents applying the target input in step S1, represents all possible follow-up strategies, represents obtaining the target input by applying the target execution results collected and stored in step S3, represents all possible follow-up strategies under the target input, represents under the target input, all possible follow-up strategies output the maximum Q value.

[0039] Preferably, takes the value of 0.9, takes the value of 0.5.

[0040] Preferably, in step A1:

[0041] A11: Divide the training data into a preset number of mini - batch training data sets, and randomly sample the data of a mini - batch training data set for parameter update;

[0042] A12: Calculate the loss function. In each step of the update process, calculate the loss based on the gap between the current Q - value and the target Q - value, and use the root - mean - square error as the loss function :

[0043]

[0044] where, represents the target Q - value, which is the final Q - value output by the Envelope_Q - Learning model after being trained in step S5; represents the current Q - value output by the MLP network.

[0045] Preferably, in step A2:

[0046] A21: Calculate the gradient of all parameters of the MLP network according to the loss function L: and ;

[0047] A22: Use the gradient descent method to update the network parameters, and the update rule is as follows:

[0048]

[0049]

[0050] where is the learning rate, which controls the step size of each update.

[0051] Preferably, in step A3:

[0052] The preset condition is that the change in the gap between the current Q - value and the target Q - value is less than the preset threshold K or the number of training times in step A2 reaches the preset number N. The value of K is 0.01, and the value of N is 1000.

[0053] The beneficial effects of the present invention are as follows: It provides an intelligent exercise intervention follow-up method combining MLP and EnvelopeQ-Learning. Compared with traditional exercise intervention follow-up methods, it can deeply analyze the individual differences of patients, accurately analyze the exercise compliance and self-efficacy status of patients with MLP, output the Q value of each follow-up action, provide accurate quantitative guidance for strategy selection, and cooperate with the EnvelopeQ-Learning algorithm to dynamically optimize the Q value function according to environmental rewards and punishments, thereby quickly mastering the best follow-up strategies suitable for different patients and greatly improving the exercise compliance and self-efficacy of patients. In addition, medical staff can get rid of manual follow-up planning. Only by combining simple initial settings and subsequent monitoring on the software platform can they independently and continuously customize personalized follow-up plans for patients, enabling them to focus on the core work of professional medical care. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In the drawings:

[0055] Figure 1 is the logic block diagram of the intelligent exercise intervention follow-up method combining MLP and Envelope_Q-Learning proposed by the present invention.

[0056] Figure 2 is the two-dimensional grid diagram of the exercise state proposed by the present invention.

[0057] Figure 3 is the architecture diagram of the intelligent exercise intervention follow-up system combining MLP and Envelope_Q-Learning proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] Referring to Figure 1 , an intelligent exercise intervention follow-up method combining MLP and Envelope_Q-Learning includes a follow-up strategy output method for outputting follow-up strategies according to the target state and a parameter update method for feedback adjustment of the follow-up strategy output method according to the target execution result after the follow-up strategy is executed;

[0059] The steps of the follow-up strategy output method are as follows:

[0060] S1: Input the target state into the MLP network, and the MLP network outputs the output value Q of each follow-up strategy according to the weight W and the bias b;

[0061] In this embodiment:

[0062] The exercise compliance and exercise self-efficacy of the patient this week are formed into a two-dimensional grid diagram of the patient's exercise state through statistical charts, where the exercise compliance and exercise self-efficacy of the patient correspond one-to-one with the position matrix formed by the two-dimensional grid diagram of the patient's exercise state. The target input in step S1 is the position matrix reflecting the exercise compliance and exercise self-efficacy of the patient this week;

[0063] The MLP network includes an input layer, a hidden layer, and an output layer, where:

[0064] The input layer contains two neurons, and the two neurons respectively correspond to the position matrix coordinates of the patient's exercise compliance and exercise self-efficacy this week;

[0065] The hidden layer contains five neurons and uses the ReLU activation function;

[0066] The output layer contains eight neurons, which respectively correspond to eight follow-up strategies. Each neuron outputs an output value Q, representing the expected cumulative reward for executing the corresponding follow-up strategy. The eight follow-up strategies include weekly phone follow-up, weekly text message follow-up, weekly face-to-face follow-up, phone follow-up every 3 days, text message follow-up every 3 days, phone follow-up every two weeks, text message follow-up every two weeks, and face-to-face follow-up every two weeks.

[0067] S2: Use - The greedy strategy selects the follow-up strategy with the largest output value Q as the target follow-up strategy for output;

[0068] In this embodiment:

[0069] Adopt a linear decreasing strategy to gradually reduce value:

[0070]

[0071] Among them, is the initial value, is the final value, t is the number of steps of the current training, is the total number of training steps.

[0072] Specifically, the takes the value of 0.9, takes the value of 0.1, takes the value of 1000.

[0073] After the target follow-up strategy is executed, collect and store the target execution result;

[0074] In this embodiment:

[0075] Exercise compliance is collected through a worn instrument, such as a smart watch, a smart bracelet, etc.;

[0076] Exercise self-efficacy is collected after filling out a chart. The self-efficacy evaluation scale is shown in the following table (scoring criteria: completely inconsistent: 1 point, relatively inconsistent: 2 points, relatively consistent: 3 points, completely consistent: 4 points).

[0077] Table 1 Self-Efficacy Evaluation Scale

[0078] Serial number Problem Result 1 If I try my best, I can always solve the problems encountered during the exercise □ Completely do not conform □ Relatively do not conform □ Relatively conform □ Completely conform 2 I am confident that I can find an exercise method suitable for me □ Completely do not conform □ Relatively do not conform □ Relatively conform □ Completely conform 3 I am confident that I can complete the daily exercise goals □ Completely do not conform □ Relatively do not conform □ Relatively conform □ Completely conform 4 Facing the problems encountered during exercise, I can usually find several solutions □ Completely do not conform □ Relatively do not conform □ Relatively conform □ Completely conform 5 When I feel tired, I am confident that I can keep exercising □ Completely do not conform □ Relatively do not conform □ Relatively conform □ Completely conform 6 When I feel frustrated, I will still keep exercising □ Completely do not conform □ Relatively do not conform □ Relatively conform □ Completely conform 7 Even without the support of family or friends, I will still keep exercising □ Completely do not conform □ Relatively do not conform □ Relatively conform □ Completely conform 8 Even without the help of medical staff, I will still keep exercising □ Completely do not conform □ Relatively do not conform □ Relatively conform □ Completely conform 9 Even after stopping exercising for a period of time, I will still start exercising again □ Completely do not conform □ Relatively do not conform □ Relatively conform □ Completely conform 10 Even without professional exercise venues and equipment, I will still keep exercising □ Completely do not conform □ Relatively do not conform □ Relatively conform □ Completely conform

[0079] S4: Apply the target execution results collected and stored in step S3, and calculate the reward and punishment scores according to the preset reward and punishment strategy;

[0080] In this embodiment:

[0081] The exercise compliance is divided into five levels in descending order of the exercise compliance rate, and the exercise self-efficacy is divided into three levels in descending order of the exercise efficacy, namely: excellent, good, and slightly poor;

[0082] Reward for improved compliance: If the patient's exercise compliance this week is improved by one level compared with last week, the reward value is +2; improved by two levels, the reward value is +4; improved by three levels, the reward value is +6, improved by four levels, the reward value is +8, and other values are assigned 0;

[0083] Punishment for decreased compliance: If the patient's exercise compliance this week is decreased by one level compared with last week, the reward value is -2; decreased by two levels, the reward value is -4; decreased by three levels, the reward value is -6, decreased by four levels, the reward value is -8, and other values are assigned 0;

[0084] Reward for exercise self-efficacy: If the patient's exercise self-efficacy score is improved from "slightly poor" to "good", the reward value is +3, from "good" to "excellent", the reward value is +5, the self-efficacy score reaches "excellent" and remains within this range for at least one week, the reward value is +8. Other values are assigned 0;

[0085] Punishment for exercise self-efficacy: If the patient's exercise self-efficacy is decreased from "good" to "slightly poor", the punishment value is -3; from "excellent" to "good", the punishment value is -5; the self-efficacy score drops to the "slightly poor" area and remains for at least one week, the punishment value is -8. Other values are assigned 0.

[0086] S5: Input the reward and punishment scores into the Envelope_Q-Learning model for training to obtain the final Q value;

[0087] In this embodiment:

[0088] The calculation formula for updating the target Q value in the Envelope_Q-Learning model is:

[0089]

[0090] Among them, is the learning rate, is the discount factor, is the current reward, Represents the target input in application step S1. Represents all possible follow-up strategies. Represents the target input obtained by applying the target execution result collected and stored in application step S3. Represents All possible follow-up strategies under the target input. Represents Under the target input, All possible follow-up strategies output the maximum Q value.

[0091] Specifically, The value is 0.9, The value is 0.5.

[0092] Specifically, the rule of R is the sum of the above-mentioned exercise self-efficacy reward and punishment rule and compliance rule. For example, when the user's current state is [compliance is between 0% and 20%, self-efficacy is slightly poor], and when performing a certain action (such as weekly telephone follow-up), it jumps to the next state (such as [compliance is between 40% and 60%, self-efficacy is good]), then the current reward of R = 4 (the compliance level is increased by two levels) + 3 (the self-efficacy is improved from slightly poor to good) = 7. That is, the R reward obtained by performing the current action in the current state is 7.

[0093] The method for updating the parameters is as follows:

[0094] A1: Apply the final Q value obtained in step S5 to calculate the loss function between the final Q value and the Q value obtained by the MLP. ;

[0095] In this embodiment:

[0096] A11: Divide the training data into a preset number of mini-batch training data sets, and randomly sample the data of a mini-batch training data set for parameter update;

[0097] A12: Calculate the loss function. In each step of the update process, calculate the loss according to the gap between the current Q value and the target Q value, and use the root mean square error as the loss function. :

[0098]

[0099] Among them, Represents the target Q value, which is the final Q value output by the Envelope_Q-Learning model after being trained in step S5. Represents the current Q value output by the MLP network.

[0100] A2: Update the weights W and biases b in the MLP network through the gradient descent method, thereby updating the Q values obtained by the MLP;

[0101] In this embodiment:

[0102] A21: Calculate the gradients for all parameters of the MLP network according to the loss function L: and ;

[0103] A22: Use the gradient descent method to update the network parameters, and the update rules are as follows:

[0104]

[0105]

[0106] where is the learning rate, which controls the step size of each update.

[0107] A3: Determine whether the preset conditions are met. If not, execute step A1. If so, terminate the update of the weights W and biases b and obtain the final weights and biases;

[0108] In this embodiment:

[0109] The preset conditions are that the change in the gap between the current Q value and the target Q value is less than the preset threshold K or the number of training times in step A2 reaches the preset number N. The value of K is 0.01, and the value of N is 1000.

[0110] Apply the final weights W and biases b obtained in step A3 to replace the weights W and biases b of the MLP network in step S1 to complete the update of the parameters.

[0111] Figure 3 , To more clearly illustrate the solutions and effects of this embodiment, in combination with the system built based on the intelligent motion intervention follow-up method of MLP and Envelope_Q-Learning, and in combination with the accompanying drawings for example, the system includes an action execution module, an agent, an MLP network (multi-layer perceptron), and a state feedback module:

[0112] At the beginning of each round of loop, based on the grid position where the current agent is located, that is, the current patient's exercise compliance and self-efficacy status, the agent inputs this state information into the MLP network (step S1). The MLP quickly calculates and outputs the Q values corresponding to each action. At this time, the agent makes action selections according to -greedy strategy. In the initial stage of training, A higher value (initial value 0.9) means that the agent has a greater probability of randomly selecting actions, which helps to widely explore the effects of various follow-up strategies under different patient states; as the number of training steps t increases, the value gradually decreases according to a linear decreasing strategy until it finally reaches 0.1. At this time, the agent is increasingly inclined to select the action with the largest Q value, that is, based on the experience accumulated from the previous exploration, select the follow-up strategy that is most likely to bring high rewards (step S2).

[0113] Once the agent selects an action, the action execution module starts to operate. For example, it triggers the telephone follow-up system to conduct weekly telephone follow-ups, or arranges medical staff for face-to-face follow-ups, etc., to effectively implement the selected follow-up action.

[0114] After the action is executed, the environmental feedback module immediately comes into play. It is used to monitor the subsequent exercise compliance and self-efficacy changes of the patient (step S3), and give corresponding feedback according to the pre-set reinforcement learning reward and punishment rules (step S4). If after this follow-up intervention, the patient's exercise compliance has increased by one level and the self-efficacy has improved from "slightly poor" to "good", then the environment will give the agent a reward value of +5 (2 + 3). At the same time, according to the patient's new state, a new coordinate point is located in the two-dimensional grid, and the new state information represented by this new coordinate point is also fed back to the agent.

[0115] After the agent receives the rewards and new states feedback from the environment, it uses the Envelope_Q-Learning update rule (step S5) to perform the key strategy optimization step. First, calculate the target Q value according to the formula. This formula fully considers the current reward, future potential rewards (reflecting the degree of emphasis on future rewards through the discount factor), and the maximum Q value of the next state, comprehensively measuring the impact of the current action selection on subsequent benefits. Then, adjust the parameters of the MLP through the backpropagation algorithm. Specifically:

[0116] The system quantifies and evaluates the gap between the Q value output by the current MLP and the just calculated target Q value using the root mean square error as the loss function, accurately positioning the deviation degree between the current output of the MLP and the ideal target (step A1).

[0117] With the help of the backpropagation algorithm, according to the chain rule, the loss error is propagated backward from the output layer to the hidden layer and the input layer, carefully calculating the contribution of each parameter (weight W and bias b) in the MLP network to the error, that is, finding the gradient of each parameter.

[0118] Using the gradient descent method, according to the update rule, with the learning rate controlling the step size, the network parameters are updated and adjusted, enabling the MLP network to gradually approximate the optimal value function estimation, so that the output Q value can better guide the follow-up strategy selection (step A2).

[0119] To improve the training efficiency and stability, the training data is also divided into 10 small batches of patches, and a small batch of data is randomly sampled each time to perform the parameter update operation, avoiding problems such as overfitting or waste of computing resources caused by processing a large amount of data at one time.

[0120] The entire training process continues in a loop, as Figure 1 shown, repeating steps A1 and A2 until the training termination conditions are met: firstly, when the change in the Q value is less than the set threshold (0.01), which means that the Q value output by the MLP has tended to be stable, and it is difficult for the system to obtain a better strategy through continuous training; secondly, reaching the specified number of training times (1000 times), ensuring from the perspective of time and resource investment that the training process will not continue indefinitely (step A3). When the termination conditions are met, the agent has successfully learned the optimal follow-up strategy for different patients. Thereafter, in actual applications, personalized and precise follow-up plans can be quickly given based on the real-time status of the patients, helping patients continuously improve their exercise compliance and self-efficacy during the exercise intervention process, and medical staff can also be freed from the cumbersome follow-up planning work and focus their energy on the core medical affairs of the profession.

Claims

1. An intelligent exercise intervention follow-up method combining MLP and Envelope_Q-Learning, characterized by: It includes a follow-up strategy output method for outputting the follow-up strategy according to the target state and a parameter updating method for feedback-adjusting the follow-up strategy output method according to the target execution result after the follow-up strategy is executed; The steps of the follow-up strategy output method are as follows: S1: Input the target state into the MLP network, and the MLP network outputs the output value Q of each follow-up strategy based on the weight W and bias b; S2: Exploitation -The greedy strategy selects the follow-up strategy with the largest output value Q value as the target follow-up strategy for output; S3: After the target follow-up strategy is executed, the target execution results are collected and accessed; S4: Apply the target execution results collected and stored in step S3 to calculate reward and punishment scores according to the preset reward and punishment strategy; S5: Input the reward and punishment scores into the Envelope_Q-Learning model for training to obtain the final Q value; The parameter updating method steps are as follows: A1: Apply the final Q value obtained in step S5 and calculate the loss function between the final Q value and the Q value obtained by MLP ; A2: Update the weights W and bias b in the MLP network by gradient descent, thereby updating the Q value obtained by the MLP; A3: Determine whether the preset condition is met. If not, execute step A1. If yes, terminate the update of weight W and bias b to obtain the final weight and bias. Apply the final weight W and bias b obtained in step A3 to replace the MLP network weight W and bias b in step S1 to complete the parameter update.

2. According to claim 1, the intelligent exercise intervention follow-up method combining MLP and Envelope_Q-Learning is characterized in that: In step S1: The patient's exercise compliance and exercise self-efficacy this week are used to form a two-dimensional grid diagram of the patient's exercise status through statistical charts, wherein the patient's exercise compliance and exercise self-efficacy correspond one-to-one to the position matrix formed by the two-dimensional grid diagram of the patient's exercise status, and the target input in step S1 is a position matrix reflecting the patient's exercise compliance and exercise self-efficacy this week; The MLP network in step S1 includes an input layer, a hidden layer, and an output layer, where: The input layer contains two neurons, which correspond to the position matrix coordinates of the patient's exercise compliance and exercise self-efficacy this week; The hidden layer contains five neurons and uses the ReLU activation function; The output layer contains eight neurons, corresponding to eight follow-up strategies. Each neuron outputs an output value Q, which expresses the expected cumulative reward for executing the corresponding follow-up strategy. The eight follow-up strategies include weekly telephone follow-up, weekly text message follow-up, weekly face-to-face follow-up, telephone follow-up every 3 days, text message follow-up every 3 days, telephone follow-up every two weeks, text message follow-up every two weeks, and face-to-face follow-up every two weeks.

3. According to claim 1, the intelligent exercise intervention follow-up method combining MLP and Envelope_Q-Learning is characterized in that: In step S2: Use a linear reduction strategy to gradually reduce value: ; in, is the initial value, is final value, t is the number of steps in the current training, is the total number of training steps.

4. The intelligent exercise intervention follow-up method combining MLP and Envelope_Q-Learning according to claim 3, characterized in that: Said The value is 0.

9. The value is 0.

1. The value is 1000.

5. According to claim 1, the intelligent exercise intervention follow-up method combining MLP and Envelope_Q-Learning is characterized in that: In step S4: Exercise compliance was divided into five levels from high to low according to the exercise compliance rate, and exercise self-efficacy was divided into three levels from high to low according to the exercise efficacy: excellent, good and slightly poor; Compliance improvement reward: If the patient's exercise compliance this week increases by one level compared to last week, the reward value is +2; if it increases by two levels, the reward value is +4; if it increases by three levels, the reward value is +6; if it increases by four levels, the reward value is +8, and the rest are assigned 0; Compliance decline penalty: If the patient's exercise compliance this week drops by one level compared to last week, the reward value is -2; If the level drops by two, the reward value is -4; if the level drops by three, the reward value is -6; if the level drops by four, the reward value is -8, and the rest are assigned 0; Exercise self-efficacy reward: If the patient's exercise self-efficacy score improves from "slightly poor" to "good", the reward value is +3, and if it improves from "good" to "excellent", the reward value is +5. If the self-efficacy score reaches "excellent" and remains in this range for at least one week, the reward value is +8. Other values ​​are 0; Exercise self-efficacy penalty: If the patient's exercise self-efficacy drops from "good" to "slightly poor", the penalty value is -3; if it drops from "excellent" to "good", the penalty value is -5; if the self-efficacy score drops to the "slightly poor" area and remains there for at least one week, the penalty value is -8, and other values ​​are 0.

6. The intelligent exercise intervention follow-up method combining MLP and Envelope_Q-Learning according to claim 2, characterized in that: In step S5: The calculation formula for the target Q value update in the Envelope_Q-Learning model is: ; in, is the learning rate, is the discount factor, For the current reward, represents the target input in the application step S1, represents all possible follow-up strategies, represents the target input obtained by applying the target execution result collected and stored in step S3, express All possible follow-up strategies under target input, express Under target input, All possible follow-up strategies output the maximum Q value.

7. The intelligent exercise intervention follow-up method combining MLP and Envelope_Q-Learning according to claim 6, characterized in that: The value is 0.

9. The value is 0.

5.

8. The intelligent exercise intervention follow-up method combining MLP and Envelope_Q-Learning according to claim 2, characterized in that: In step A1: A11: Divide the training data into a preset number of small batch training data sets, and randomly sample data from a small batch training data set to update the parameters; A12: Calculate the loss function. In each update process, the loss is calculated based on the gap between the current Q value and the target Q value, using the root mean square error as the loss function. : ; in, represents the target Q value, which is the final Q value output by the Envelope_Q-Learning model after training in step S5. Represents the current Q value output by the MLP network.

9. The intelligent exercise intervention follow-up method combining MLP and Envelope_Q-Learning according to claim 2, characterized in that: In step A2: A21: Calculate the gradient of all parameters of the MLP network according to the loss function L: and ; A22: Use the gradient descent method to update the network parameters. The update rules are as follows: ; ; in is the learning rate, which controls the step size of each update.

10. The intelligent exercise intervention follow-up method combining MLP and Envelope_Q-Learning according to claim 8, characterized in that: In step A3: The preset condition is that the difference between the current Q value and the target Q value is less than a preset threshold K or the number of training times in step A2 reaches a preset number N, where the value of K is 0.01 and the value of N is 1000.

Citation Information

Patent Citations

  • Internet-based intelligent patient care follow-up system and method

    CN118016326B

  • Model-based intelligent follow-up visit system and follow-up visit method

    CN118522389A

  • Mobile robot path planning based on improved depth Q network algorithm

    CN115344046A

  • Off-line reinforcement learning method based on same-strategy regularization strategy evaluation

    CN117875451A

  • Device and method for improved policy learning for robots

    US20240311640A1