Follow-up control method, device and equipment of upper limb rehabilitation robot and storage medium

By using a pre-trained second control strategy model and personalized muscle activation regulation, the problem of upper limb rehabilitation robots not considering rehabilitation levels is solved, and more effective rehabilitation training results are achieved.

CN121625173BActive Publication Date: 2026-06-09CHENGDU UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU UNIV OF TRADITIONAL CHINESE MEDICINE
Filing Date
2026-02-04
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing upper limb rehabilitation robots fail to consider the patient's rehabilitation level when assisting rehabilitation training, resulting in poor training effects.

Method used

A pre-trained second control strategy model is adopted. By acquiring the transfer tuples during the rehabilitation training process, including the starting and ending observation states and action rewards, the model is adjusted to adapt to individual needs. Personalized rehabilitation training is achieved by utilizing muscle group activation level rewards and compensatory muscle group inhibition regulation.

Benefits of technology

It improves the effectiveness of rehabilitation training by guiding patients to train in the correct neuromuscular pattern through personalized control strategies, thereby enhancing rehabilitation outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121625173B_ABST
    Figure CN121625173B_ABST
Patent Text Reader

Abstract

The application discloses a servo control method and device of an upper limb rehabilitation robot, equipment and a storage medium, and relates to the field of robot control. The method comprises the following steps: assisting a target object to perform rehabilitation training by using an upper limb rehabilitation robot which is deployed with a pre-trained second control strategy model; acquiring a second transition tuple of each time step in the rehabilitation training process; and after the target object performs several times of rehabilitation training with the assistance of the upper limb rehabilitation robot, performing transfer learning on the second control strategy model by taking several second transition tuples as training data, so as to adjust the second control strategy model. The application realizes the improvement of the rehabilitation training effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot control technology, and in particular to a follow-up control method, device, equipment and storage medium for an upper limb rehabilitation robot. Background Technology

[0002] Stroke is a leading cause of disability and death worldwide, with approximately 80% of survivors experiencing motor dysfunction, particularly in the upper limbs, where recovery is often difficult. The chronic phase of stroke refers to the period six months after onset, when conventional treatments become less effective, leading to functional stagnation and psychological burnout. Upper limb rehabilitation robots, with their advantages of high-intensity, repetitive, and task-oriented training, offer a new approach to overcoming the challenges of rehabilitation.

[0003] In upper limb rehabilitation robot control technology, a single feature feedback (force feedback or electromyography) is usually used to identify the patient's movement intention, and then the upper limb rehabilitation robot is controlled to assist the patient in rehabilitation training according to the patient's movement intention. However, this method does not take into account the patient's rehabilitation level, resulting in poor rehabilitation training effect. Summary of the Invention

[0004] The main purpose of this application is to provide a follow-up control method, device, equipment and storage medium for an upper limb rehabilitation robot, aiming to solve the technical problem of poor training effect of rehabilitation robot.

[0005] To achieve the above objectives, this application proposes a follow-up control method for an upper limb rehabilitation robot, comprising:

[0006] An upper limb rehabilitation robot equipped with a pre-trained second control strategy model is used to assist the target object in rehabilitation training.

[0007] The second transition tuple is obtained for each time step during the rehabilitation training process. The second transition tuple includes the starting observation state, the target action, the action reward, and the ending observation state. The observation state includes the human-computer interaction torque vector, angular velocity vector, and angle vector of the upper limb rehabilitation robot, as well as the target active muscle group activation vector and the compensatory muscle group activation vector of the target object. The action reward includes at least the muscle group activation level reward. The muscle group activation level reward is used to positively stimulate the activation of the target active muscle group and inhibit the activation of the compensatory muscle group.

[0008] After the target object undergoes several rehabilitation training sessions assisted by the upper limb rehabilitation robot, the second control strategy model is transferred and learned using several second transfer tuples as training data to adjust the second control strategy model.

[0009] In one embodiment, obtaining the second transition tuple for each time step during the rehabilitation training process includes:

[0010] The starting point observation state is collected, and the second control strategy model is used to output the target action based on the starting point observation state, so that the upper limb rehabilitation robot can perform the target action.

[0011] After the upper limb rehabilitation robot performs the target action, the endpoint observation status is collected;

[0012] The action reward is calculated based on the starting point observation status and the ending point observation status;

[0013] The second transition tuple is constructed based on the starting observation state, target action, action reward, and ending observation state.

[0014] In one embodiment, calculating the action reward based on the starting observation state and the ending observation state includes:

[0015] Based on the target impedance parameters and the human-machine interaction torque vector and angular velocity vector of the endpoint observation state, the synchronization and coordination reward is calculated.

[0016] The motion smoothness reward is calculated based on the angular acceleration and human-computer interaction torque of the starting and ending observation states.

[0017] Based on the target active muscle group activation vector and the compensatory muscle group activation vector at the endpoint observation state, the muscle group activation level reward is calculated.

[0018] The motion reward is calculated based on the synchronization and coordination reward, the motion fluency reward, and the motion fluency reward.

[0019] In one embodiment, the observation status at the data acquisition starting point includes:

[0020] Electromyographic signals of the target active muscle groups and compensatory muscle groups were collected;

[0021] The activation vector of the target agonist muscle group is calculated based on the electromyographic signal of the target agonist muscle group and the maximum voluntary contraction value of the target agonist muscle group.

[0022] The activation vector of the compensating muscle group is calculated based on the electromyographic signals of the compensating muscle group and the maximum voluntary contraction value of the compensating muscle group.

[0023] In one embodiment, prior to the step of assisting the target object in rehabilitation training using an upper limb rehabilitation robot deployed with a pre-trained second control strategy model, the method further includes:

[0024] In a simulation environment, a kinematic model of the simulated robot and its upper limbs, connected by flexible constraints, as well as an initial control strategy model, are constructed.

[0025] The upper limb kinematics model is used to perform a motion task, and the first transition tuple of each time step in the motion task is obtained in real time. The observed state of the first transition tuple includes the human-computer interaction torque vector, angular velocity vector and angle vector of the upper limb rehabilitation robot.

[0026] The initial control strategy model is trained using several first transition tuples as training data to obtain the first control strategy model.

[0027] The first control strategy model is structurally modified to obtain the second control strategy model.

[0028] In one embodiment, performing the motor task using an upper limb kinematic model includes:

[0029] The upper limb dynamics model is driven to move actively using a stochastic trajectory planning method.

[0030] In one embodiment, the step of performing transfer learning on the second control policy model using a plurality of second transfer tuples as training data to adjust the second control policy model includes:

[0031] Initialize the second control strategy model;

[0032] In the first phase of training, the weights of the physical flow branch of the second control strategy model after initialization are frozen, and the biological flow branch of the second control strategy model is trained using the second transition tuple.

[0033] In the second phase of training, the weights of the physical flow branches are unfrozen, and end-to-end global fine-tuning is performed.

[0034] Furthermore, to achieve the above objectives, this application also proposes a follow-up control device for an upper limb rehabilitation robot, comprising:

[0035] The follow-up control module is used to assist the target object in rehabilitation training by using an upper limb rehabilitation robot with a pre-trained second control strategy model deployed on it.

[0036] The training data acquisition module is used to acquire the second transition tuple for each time step in the rehabilitation training process. The second transition tuple includes the starting observation state, the target action, the action reward, and the ending observation state. The observation state includes the human-computer interaction torque vector, angular velocity vector, and angle vector of the upper limb rehabilitation robot, as well as the target active muscle group activation vector and the compensatory muscle group activation vector of the target object. The action reward includes at least the muscle group activation level reward. The muscle group activation level reward is used to positively stimulate the activation of the target active muscle group and inhibit the activation of the compensatory muscle group.

[0037] The training module is used to perform transfer learning on the second control strategy model using several second transfer tuples as training data after the target object has undergone several rehabilitation training sessions assisted by the upper limb rehabilitation robot, so as to adjust the second control strategy model.

[0038] In addition, to achieve the above objectives, this application also proposes a follow-up control device for an upper limb rehabilitation robot, comprising: a memory, a processor, and a computer program stored in the memory and running on the processor, the computer program being configured to implement the steps of the follow-up control method for the upper limb rehabilitation robot as described above.

[0039] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the above-described follow-up control method for an upper limb rehabilitation robot.

[0040] One or more technical solutions proposed in this application have at least the following technical effects:

[0041] A follow-up control method for an upper limb rehabilitation robot is proposed. First, the upper limb rehabilitation robot with a pre-trained second control strategy model is used to assist the target object in rehabilitation training. Then, the second transition tuples at each time step during the rehabilitation training process are obtained, and several second transition tuples are used as training data to perform transfer learning on the second control strategy model to adjust the second control strategy model. This allows the second control strategy model to learn control strategies that meet the individualized needs of the target object, guiding the target object to train in the correct neuromuscular pattern, thereby improving the rehabilitation training effect. Attached Figure Description

[0042] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0043] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 A flowchart illustrating the first embodiment of the follow-up control method for the upper limb rehabilitation robot provided in this application;

[0045] Figure 2 A schematic diagram of the training process of the first control strategy model, which is an example of the follow-up control method for the upper limb rehabilitation robot provided in this application;

[0046] Figure 3 A schematic diagram of the rehabilitation training process, illustrating the follow-up control method for the upper limb rehabilitation robot provided in this application;

[0047] Figure 4 A schematic diagram of the training process of the second control strategy model, which is an example of the follow-up control method for the upper limb rehabilitation robot provided in this application.

[0048] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0049] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application. To better understand the technical solutions of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0050] This application provides a follow-up control method for an upper limb rehabilitation robot.

[0051] In the first embodiment of the follow-up control method for the upper limb rehabilitation robot of this application, referring to Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the follow-up control method for the upper limb rehabilitation robot of this application. The follow-up control method for the upper limb rehabilitation robot may include steps S20 to S40:

[0052] Step S20: The upper limb rehabilitation robot, which is equipped with a pre-trained second control strategy model, assists the target object in rehabilitation training.

[0053] It should be noted that the pre-trained second control strategy model has a dual-stream input structure, consisting of a physical flow branch and a biological flow branch. The physical flow branch receives the angle vector, angular velocity vector, and human-machine interaction torque vector of the upper limb rehabilitation robot. Specifically, since the upper limb rehabilitation robot has multiple joints, the angle vector can include the angle vectors of each joint. Similarly, the angular velocity vector and the human-machine interaction torque vector can also include the corresponding physical state information of each joint. The biological flow branch receives the target active muscle group activation vector and the compensatory muscle group activation vector of the target object.

[0054] It should also be noted that, within each time step, after receiving the angle vector, angular velocity vector, and human-machine interaction torque vector of the upper limb rehabilitation robot, and the target active muscle group activation vector and compensatory muscle group activation vector of the target object, the rehabilitation robot outputs the target action based on the starting observation state including the above parameters. The target action is the expected angular velocity of each joint of the upper limb rehabilitation robot. Then, the target action is sent to the controller of the upper limb rehabilitation robot so that the controller drives the upper limb rehabilitation robot to assist the target object in rehabilitation training according to the target action.

[0055] Step S30: Obtain the second transition tuple for each time step during the rehabilitation training process. The second transition tuple includes the starting observation state, the target action, the action reward, and the ending observation state. The observation state includes the human-computer interaction torque vector, angular velocity vector, and angle vector of the upper limb rehabilitation robot, as well as the target active muscle group activation vector and the compensating muscle group activation vector of the target object. The action reward includes at least the muscle group activation level reward. The muscle group activation level reward is used to positively stimulate the activation of the target active muscle group and simultaneously inhibit the activation of the compensating muscle group.

[0056] It should be noted that during rehabilitation training, there may be situations where the target subject's compensatory muscle group activation level is too high and the target active muscle group activation level is insufficient. This situation will affect the target subject's rehabilitation training effect. Therefore, by configuring muscle group activation level reward items, the activation of the target active muscle group can be positively stimulated, while the activation of the compensatory muscle group can be inhibited and regulated. This guides the second control strategy model to optimize in the direction of neural remodeling, thereby improving the rehabilitation training effect.

[0057] In one feasible implementation, step S30, "obtaining the second transition tuple for each time step during rehabilitation training," may include steps A11-A14:

[0058] Step A11: Collect the starting point observation state, and use the second control strategy model to output the target action based on the starting point observation state so that the upper limb rehabilitation robot can perform the target action.

[0059] Step A12: After the upper limb rehabilitation robot performs the target action, collect the endpoint observation status.

[0060] Step A13: Calculate the action reward based on the starting point observation state and the ending point observation state.

[0061] Step A14: Construct the second transition tuple based on the starting observation state, target action, action reward, and ending observation state.

[0062] In one feasible implementation, step A11 may include steps A111 to A113:

[0063] Step A111: Collect electromyographic signals of the target active muscle group and the compensatory muscle group.

[0064] Step A112: Calculate the activation vector of the target agonist muscle group based on the electromyographic signal of the target agonist muscle group and the maximum voluntary contraction value of the target agonist muscle group.

[0065] Step A113: Calculate the activation vector of the compensating muscle group based on the electromyographic signals and the maximum voluntary contraction value of the compensating muscle group.

[0066] It should be noted that the maximum voluntary contraction values ​​of the target active muscle group and the compensatory muscle group are collected in advance before rehabilitation training. The electromyography (EMG) signals of the target active muscle group and the compensatory muscle group can be collected in real time by deploying surface electromyography (sEMG) electrodes on the target active muscle group and the compensatory muscle group. After the raw signals are collected, the raw signals can be first processed by bandpass filtering and full-wave rectification, and then the root mean square of the corresponding EMG signal can be calculated in a sliding time window. The root mean square is divided by the corresponding maximum voluntary contraction value to obtain the standardized activation vector of the target active muscle group and the activation vector of the compensatory muscle group.

[0067] In one feasible implementation, step A13 may include steps A131 to A134:

[0068] Step A131: Calculate the synchronization and coordination reward based on the target impedance parameters and the human-machine interaction torque vector and angular velocity vector of the endpoint observation state.

[0069] It should be noted that the synchronous coordination reward is used to quantify the degree to which the current human-computer interaction conforms to the ideal dynamic relationship of "on-demand assistance".

[0070] Step A132: Calculate the motion smoothness reward based on the angular acceleration and human-computer interaction torque of the starting and ending observation states.

[0071] It should be noted that motion smoothness rewards are used to penalize unsmooth motion and unstable interactions.

[0072] Step A133: Calculate the muscle activation level reward based on the target active muscle group activation vector and the compensatory muscle group activation vector at the endpoint observation state.

[0073] Step A134: Calculate the motion reward based on the synchronization coordination reward, motion fluency reward, and motion fluency reward.

[0074] Specifically, the action reward function of the second strategy control model includes a synchronization coordination reward term, a synchronization coordination reward term, and a relative reward term based on muscle activation level. The expression for the action reward function is as follows:

[0075] ;

[0076] ;

[0077] ;

[0078] ;

[0079] In the formula, As a reward for action, To synchronize and coordinate rewards, Rewards for smoother movement. Rewards for muscle activation levels. H and K The target impedance parameter is set based on rehabilitation medicine theory. The jerkiness is calculated by numerical difference of the angular velocity. The rate of change of the human-computer interaction torque is obtained by numerically differencing the human-computer interaction torque. active For the target active muscle group assembly, compensatory As a compensatory muscle group, For the first i The activation level of the target active muscle group For the first i The positive reward weighting of the activation level of each target active muscle group. For the first j The activation level of each compensatory muscle group For the first j The negative reward or penalty weighting of the activation level of each compensatory muscle group. , and For weight hyperparameters.

[0080] Step S40: After the target object undergoes several rehabilitation training sessions assisted by the upper limb rehabilitation robot, the second control strategy model is transferred and learned using several second transfer tuples as training data to adjust the second control strategy model.

[0081] It should be noted that during each training process, all second transfer tuples are recorded in real time to form a personalized dataset for the patient. This data includes the patient's movement habits and muscle activation characteristics. The second control strategy model is iteratively optimized through the personalized dataset to make the model more suitable for the individual characteristics of the target object.

[0082] In one feasible implementation, step S40 may include steps S41 to S43:

[0083] Step S41: Initialize the second control strategy model.

[0084] Step S42: In the first stage of training, freeze the weights of the physical flow branch of the second control strategy model after initialization, and train the biological flow branch of the second control strategy model using the second transition tuple.

[0085] Step S43: In the second phase of training, unfreeze the weights of the physical flow branches and perform end-to-end global fine-tuning.

[0086] This embodiment provides a follow-up control method for an upper limb rehabilitation robot. First, the upper limb rehabilitation robot, which is equipped with a pre-trained second control strategy model, assists the target object in rehabilitation training. Then, the second transition tuples at each time step during the rehabilitation training are obtained, and several second transition tuples are used as training data to perform transfer learning on the second control strategy model to adjust the second control strategy model. This allows the second control strategy model to learn control strategies that meet the individualized needs of the target object, guiding the target object to train in the correct neuromuscular pattern, thereby improving the rehabilitation training effect.

[0087] In one feasible implementation, the follow-up control method for the upper limb rehabilitation robot may further include step S10 prior to step S20, and step S10 may include steps S101 to S104:

[0088] Step S101: In the simulation environment, construct the kinematic model of the simulated robot and its upper limbs, as well as the initial control strategy model, which are connected by flexible constraints.

[0089] It should be noted that the simulation robot is a precise virtual model of the upper limb rehabilitation robot. The simulation robot has the same kinematic and dynamic parameters as the upper limb rehabilitation robot, including link length, joint type, joint limit, mass and inertia tensor. The upper limb dynamic model is an equivalent kinematic model based on human anatomical structure and is configured with corresponding biomechanical properties.

[0090] Step S102: Execute a motion task using an upper limb kinematic model and acquire the first transition tuple for each time step in the motion task in real time. The observed state of the first transition tuple includes the human-computer interaction torque vector, angular velocity vector, and angle vector of the upper limb rehabilitation robot.

[0091] It should be noted that the first transition tuple also includes the starting observation state, target action, action reward, and ending observation state within a single time step. Unlike the observation states in the second transition tuple, the observation states in the first transition tuple do not include the target active muscle group activation vector and the compensatory muscle group activation vector, and the action reward in the first transition tuple does not include muscle group activation level reward. For example, the action reward includes synchronization coordination reward and movement fluency reward.

[0092] In one feasible implementation, a stochastic trajectory planning method is used to drive the active movement of the upper limb dynamics model.

[0093] It should be noted that, in order to simulate the active intention of the target object, a series of diverse three-dimensional spatial motion trajectories simulating the desktop object retrieval task can be generated based on polynomial interpolation and Gaussian noise, and the upper limb kinematic model can be driven to track these trajectories.

[0094] Step S103: Use several first transition tuples as training data to train the initial control strategy model to obtain the first control strategy model.

[0095] It should be noted that both the initial control strategy model and the first control strategy model only include physical flow branches. After learning and training the initial control strategy model, a general physical cooperative strategy basic model, namely the first control strategy model, is obtained.

[0096] Step S104: Modify the structure of the first control strategy model to obtain the second control strategy model.

[0097] It should be noted that the first control strategy model can be deployed on a rehabilitation robot, and its network structure can be transformed into a dual-input model, namely the second control strategy model. Specifically, the entire structure and weights of the first control strategy model constitute the physical flow branch of the second control strategy model, and the gradient updates of all its layers are temporarily frozen. A new biological flow branch with a similar structure but randomly initialized weights is added to obtain the second control strategy model.

[0098] The follow-up control method for the upper limb rehabilitation robot in this embodiment first trains the initial control strategy model in a simulation environment to obtain a general physical coordination strategy basic model. Then, the basic model is structurally modified and deployed to obtain a pre-trained second control strategy model. In the simulation environment, massive amounts of training data can be generated infinitely and without risk, which avoids the high cost and potential safety risks of repeated trial and error in the real world. Through training in simulation, the initial control strategy model can quickly learn basic control capabilities that are highly related to the physical world, such as how to maintain smooth motion and how to make appropriate responses according to human-machine interaction torque. Only fine-tuning is needed in the real environment, and it can converge to the optimal solution at a faster speed.

[0099] For example, to aid in understanding the follow-up control method of the upper limb rehabilitation robot in the embodiments of this application, such as Figures 2 to 4As shown, an example application of a follow-up control method for an upper limb rehabilitation robot is provided. Specifically, the target subject has impaired right arm mobility, clinically manifested as insufficient shoulder abduction and elbow extension. When attempting to perform a "tabletop retrieval" rehabilitation training task, compensatory contractions of the pectoralis major and biceps brachii muscles are often observed. This example uses a seven-DOF upper limb rehabilitation robot to assist the target subject in rehabilitation training. The follow-up control method of the upper limb rehabilitation robot may include steps S1001~S1010:

[0100] Step S1001: Construct the initial control strategy model M 1. The initial control policy model is a deep reinforcement learning model based on the Soft Actor-Critic (SAC) algorithm. Both the Actor network and the Critic network employ a multilayer perceptron (MLP) with two hidden layers (256 neurons per layer) and ReLU as the activation function. Simultaneously, an experience replay buffer with a capacity of 1,000,000 transition tuples is initialized. .

[0101] Step S1002: Build a virtual simulation environment. In the MuJoCo physics simulation engine, load a simulation robot with the same DH parameters, mass, and inertia tensor as the 7-DOF rehabilitation robot. Construct a simplified dynamic model of the right upper limb. Connect the dynamic model of the right upper limb to the end of the simulation robot through virtual constraints of simulated flexible straps. A virtual six-dimensional torque sensor is set at the connection point.

[0102] Step S1003: Execute the movement task using the upper limb kinematic model and obtain the first transition tuple for each time step in the movement task in real time.

[0103] Based on polynomial interpolation and Gaussian noise, several three-dimensional spatial motion trajectories for simulating a desktop object retrieval task are generated, and the right upper limb dynamic model is driven to track these trajectories.

[0104] At each time step of the simulation loop, the starting observation state is directly read from the MuJoCo engine. S t This includes robot joint angle vectors (angles of 7 joints), robot joint angular velocity vectors (angular velocities of 7 joints), and human-machine interaction torque vectors (obtained by a six-dimensional torque sensor).

[0105] Start-up observation status S t Input initial control strategy model M A 1 Actor network, the Actor network outputs a 7-dimensional action vector. At The action vector A t Interpreted as the expected angular velocity of the robot's seven joints ω ref .

[0106] Desired angular velocity ω ref The commands are sent to the PD controller of the simulated robot to drive the robot to move and obtain the end-point observation state of that time step. S t+1 .

[0107] Based on the endpoint observation status S t+1 Calculate action rewards R sim λ1=1.0, λ2=0.1, H =0 (zero torque control), k= 15 Nms / rad (Provides a damping force proportional to the angular velocity).

[0108] Step S1004: Store the first transition tuple into the experience replay buffer P: Set the starting observation state S t Action vectors A t Action Rewards R sim and endpoint observation status S t+1 Together they form a transfer tuple And store it in the experience replay cache P.

[0109] Step S1005: Train the initial control policy model M 1. When the amount of data in the experience replay buffer P exceeds 2000, at each simulation time step, a mini-batch of 256 samples is randomly sampled from the experience replay buffer P to test the initial control strategy model. M The M1 network (Actor and Critic) is trained and updated once; training lasts for 5,000,000 time steps. Training terminates when the average reward for each round during the evaluation phase converges and no longer shows significant increase after 100 consecutive rounds. The weights of the M1 network at this point are saved.

[0110] Step S1006: Load the trained control policy model M1 and transform its network structure into a dual-stream input model M2. The entire structure and weights of M1 constitute the physical flow branch, and the gradient updates of all its layers are temporarily frozen. Add a biological flow branch with a similar structure but randomly initialized weights. This model M2 is then deployed to the control host of the real rehabilitation robot.

[0111] Step S1007: The upper limb rehabilitation robot, equipped with a control strategy model M2, assists the target patient in rehabilitation training. The patient's right arm is fixed to the 7-DOF rehabilitation robot. Surface electromyography (sEMG) electrodes are then attached to the anterior deltoid and triceps brachii (target muscle groups), as well as the muscle bellies of the pectoralis major and biceps brachii (primary compensatory muscle groups). Before starting, the patient is instructed to contract these four muscles as much as possible, and their maximum voluntary contraction (MVC) values ​​are collected.

[0112] Step S1008: Obtain the second transition tuple for each time step during the rehabilitation training process. In the desktop object retrieval training task, the control system collects the starting point observation status in real time at a frequency of 100Hz. S real The joint angles and angular velocities (7 dimensions each) are derived from the robot motor encoder, and the interaction torque... (6-dimensional) signals are derived from the end effector force sensor. Simultaneously, the sEMG system acquires four channels of electromyography (EMG) signals, which are then bandpass filtered (20-450Hz), rectified, and have their root mean square (RMS) values ​​calculated within a 100ms sliding window. These RMS values ​​are then normalized by dividing by their respective MVC values, ultimately yielding a 2-dimensional target muscle group activation vector. E active and 2D compensatory muscle activation vectors E compensatory These data collectively constitute the 24-dimensional starting observation state. S real The state vector S real The physical component is input into the physical flow branch of M2, and the physiological component is input into the biological flow branch. After forward propagation calculation, the model outputs the expected joint angular velocity in 7 dimensions. ω ref The 7-DOF rehabilitation robot is based on... ω ref Assist patients in performing actions such as retrieving items from the table.

[0113] Step S1009: If the target object completes a set of rehabilitation training consisting of 20 tabletop object retrievals, after completion, all transition tuples generated during these 20 training sessions are saved as a session dataset, denoted as... P 1'.

[0114] Synchronous calculation at each time step , , and , The weight λ3 was set to 0.5. Based on the patient's condition, this was done to encourage the target muscle groups and inhibit compensatory muscle groups. The weights in the formula are set as follows: .

[0115] At each time step of training, the second transition tuple will be generated in real time. Save the exercise. After completing one training movement (one retrieving of an object), prepare for the next movement until the entire set of training is completed.

[0116] Step S1010: Perform transfer learning offline for model fine-tuning and iteration. After a day of rehabilitation treatment, use the collected dataset... P 1' Perform offline fine-tuning of the model. For example... Figure 4 As shown, a progressive training strategy is used to train the model. M 2. Make fine adjustments.

[0117] Phase 1: Freeze all weights of the physical flow branches. Use only the dataset collected that day. P 1' Train the biological flow branch and the fusion layer of the network backend. Use the Adam optimizer with a learning rate set to... Training is conducted for 10 epochs.

[0118] Phase Two: After completing Phase One training, unfreeze the weights of the physics flow branch. Use a smaller learning rate. For the whole M The network was fine-tuned end-to-end and trained for 5 rounds.

[0119] After fine-tuning, a personalized model for the patient's current condition is obtained. M 2_personalized This finely tuned, personalized model will be used in the next rehabilitation session. M 2_personalized This serves as a new foundational model for a new round of treatment and data collection, and the model is further optimized after treatment. Through this cycle, the control model can continuously adapt to the patient's rehabilitation progress, achieving progressive personalization. It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the follow-up control method of the upper limb rehabilitation robot of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0120] This application also provides a follow-up control device for an upper limb rehabilitation robot, which may include:

[0121] The follow-up control module is used to assist the target object in rehabilitation training by using an upper limb rehabilitation robot with a pre-trained second control strategy model deployed on it.

[0122] The training data acquisition module is used to acquire the second transition tuple for each time step in the rehabilitation training process. The second transition tuple includes the starting observation state, the target action, the action reward, and the ending observation state. The observation state includes the human-computer interaction torque vector, angular velocity vector, and angle vector of the upper limb rehabilitation robot, as well as the target active muscle group activation vector and the compensatory muscle group activation vector of the target object. The action reward includes at least the muscle group activation level reward. The muscle group activation level reward is used to positively stimulate the activation of the target active muscle group and inhibit the activation of the compensatory muscle group.

[0123] The training module is used to perform transfer learning on the second control strategy model using several second transfer tuples as training data after the target object has undergone several rehabilitation training sessions assisted by the upper limb rehabilitation robot, so as to adjust the second control strategy model.

[0124] The follow-up control device for the upper limb rehabilitation robot provided in this application adopts the follow-up control method for the upper limb rehabilitation robot in the above embodiments, which can solve the main technical problems. Compared with related technologies, the beneficial effects of the follow-up control device for the upper limb rehabilitation robot provided in this application are the same as the beneficial effects of the follow-up control method for the upper limb rehabilitation robot provided in the above embodiments, and other technical features in the follow-up control device for the upper limb rehabilitation robot are the same as the features disclosed in the follow-up control method for the upper limb rehabilitation robot in the above embodiments, and will not be repeated here.

[0125] This application also provides a follow-up control device for an upper limb rehabilitation robot. The follow-up control device for the upper limb rehabilitation robot may include: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the follow-up control method for the upper limb rehabilitation robot in the above embodiments.

[0126] The follow-up control device for the upper limb rehabilitation robot provided in this application, employing the follow-up control method for the upper limb rehabilitation robot in the above embodiments, can solve the main technical problems. Compared with related technologies, the beneficial effects of the follow-up control device for the upper limb rehabilitation robot provided in this application are the same as the beneficial effects of the follow-up control method for the upper limb rehabilitation robot provided in the above embodiments, and other technical features in the follow-up control device for the upper limb rehabilitation robot are the same as the features disclosed in the follow-up control method for the upper limb rehabilitation robot in the above embodiments, and will not be repeated here.

[0127] This application also provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the follow-up control method of the upper limb rehabilitation robot in the above embodiments.

[0128] The storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described follow-up control method for the upper limb rehabilitation robot, thereby solving the main technical problem. Compared with related technologies, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the follow-up control method for the upper limb rehabilitation robot provided in the above embodiments, and will not be repeated here.

[0129] The above are only some embodiments of this application and do not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A follow-up control method for an upper limb rehabilitation robot, characterized in that, include: An upper limb rehabilitation robot equipped with a pre-trained second control strategy model is used to assist the target object in rehabilitation training. The second transition tuple is obtained for each time step during the rehabilitation training process. The second transition tuple includes the starting observation state, the target action, the action reward, and the ending observation state. The observation state includes the human-computer interaction torque vector, angular velocity vector, and angle vector of the upper limb rehabilitation robot, as well as the target active muscle group activation vector and the compensatory muscle group activation vector of the target object. The action reward includes at least the muscle group activation level reward. The muscle group activation level reward is used to positively stimulate the activation of the target active muscle group and inhibit the activation of the compensatory muscle group. After the target object is assisted by the upper limb rehabilitation robot to perform several rehabilitation training sessions, the second control strategy model is transferred to the training data using several second transfer tuples in order to adjust the second control strategy model. The activation vector of the target active muscle group is calculated based on the electromyographic signal of the target active muscle group and the maximum voluntary contraction value of the target active muscle group. The activation vector of the compensating muscle group is calculated based on the electromyographic signal of the compensating muscle group and the maximum voluntary contraction value of the compensating muscle group. The muscle group activation level reward is calculated based on the target active muscle group activation vector and the compensatory muscle group activation vector. The step of using a plurality of second transition tuples as training data to perform transfer learning on the second control policy model to adjust the second control policy model includes: Initialize the second control strategy model; in the first stage of training, freeze the weights of the physical flow branch of the initialized second control strategy model, and train the biological flow branch of the second control strategy model using the second transition tuple; In the second phase of training, the weights of the physics flow branches are unfrozen, and end-to-end global fine-tuning is performed. The pre-trained second control strategy model is a dual-stream input structure, consisting of a physical flow branch and a biological flow branch. The physical flow branch receives the angle vector, angular velocity vector, and human-machine interaction torque vector of the upper limb rehabilitation robot, while the biological flow branch receives the target active muscle group activation vector and compensatory muscle group activation vector of the target object.

2. The follow-up control method for the upper limb rehabilitation robot as described in claim 1, characterized in that, The second transition tuple for each time step in the rehabilitation training process includes: The starting point observation state is collected, and the second control strategy model is used to output the target action based on the starting point observation state, so that the upper limb rehabilitation robot can perform the target action. After the upper limb rehabilitation robot performs the target action, the endpoint observation status is collected; The action reward is calculated based on the starting point observation status and the ending point observation status; The second transition tuple is constructed based on the starting observation state, target action, action reward, and ending observation state.

3. The follow-up control method for the upper limb rehabilitation robot as described in claim 2, characterized in that, The calculation of the action reward based on the starting point observation state and the ending point observation state includes: Based on the target impedance parameters and the human-machine interaction torque vector and angular velocity vector of the endpoint observation state, the synchronization and coordination reward is calculated. The motion smoothness reward is calculated based on the angular acceleration and human-computer interaction torque of the starting and ending observation states. Based on the target active muscle group activation vector and the compensatory muscle group activation vector at the endpoint observation state, the muscle group activation level reward is calculated. The action reward is calculated based on the synchronization and coordination reward, the motion fluency reward, and the muscle group activation level reward.

4. The follow-up control method for the upper limb rehabilitation robot as described in claim 2, characterized in that, The observation status at the data acquisition starting point includes: Electromyographic signals of the target active muscle groups and compensatory muscle groups were collected; The activation vector of the target agonist muscle group is calculated based on the electromyographic signal of the target agonist muscle group and the maximum voluntary contraction value of the target agonist muscle group. The activation vector of the compensating muscle group is calculated based on the electromyographic signals of the compensating muscle group and the maximum voluntary contraction value of the compensating muscle group.

5. The follow-up control method for the upper limb rehabilitation robot as described in claim 1, characterized in that, Prior to the step of using an upper limb rehabilitation robot equipped with a pre-trained second control strategy model to assist the target object in rehabilitation training, the method further includes: In a simulation environment, a kinematic model of the simulated robot and its upper limbs, connected by flexible constraints, as well as an initial control strategy model, are constructed. The upper limb kinematics model is used to perform a motion task, and the first transition tuple of each time step in the motion task is obtained in real time. The observed state of the first transition tuple includes the human-computer interaction torque vector, angular velocity vector and angle vector of the upper limb rehabilitation robot. The initial control strategy model is trained using several first transition tuples as training data to obtain the first control strategy model. The first control strategy model is structurally modified to obtain the second control strategy model.

6. The follow-up control method for the upper limb rehabilitation robot as described in claim 5, characterized in that, The use of upper limb kinematic models to perform motor tasks includes: The upper limb kinematic model is driven to move actively using a stochastic trajectory planning method.

7. A follow-up control device for an upper limb rehabilitation robot, characterized in that, The device is configured to implement the steps of the follow-up control method for an upper limb rehabilitation robot as described in any one of claims 1 to 6, wherein the control device includes: The follow-up control module is used to assist the target object in rehabilitation training by using an upper limb rehabilitation robot with a pre-trained second control strategy model deployed on it. The training data acquisition module is used to acquire the second transition tuple for each time step in the rehabilitation training process. The second transition tuple includes the starting observation state, the target action, the action reward, and the ending observation state. The observation state includes the human-computer interaction torque vector, angular velocity vector, and angle vector of the upper limb rehabilitation robot, as well as the target active muscle group activation vector and the compensatory muscle group activation vector of the target object. The action reward includes at least the muscle group activation level reward. The muscle group activation level reward is used to positively stimulate the activation of the target active muscle group and inhibit the activation of the compensatory muscle group. The training module is used to perform transfer learning on the second control strategy model using several second transfer tuples as training data after the target object has undergone several rehabilitation training sessions assisted by the upper limb rehabilitation robot, so as to adjust the second control strategy model.

8. A follow-up control device for an upper limb rehabilitation robot, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and running on the processor, the computer program being configured to implement the steps of the follow-up control method for the upper limb rehabilitation robot as described in any one of claims 1 to 6.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the follow-up control method for the upper limb rehabilitation robot as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • CN110652295A

  • CN112757275A