A robotic reinforcement learning mirror training control system for hemiplegic rehabilitation

By combining active and passive robots, and employing a reference impedance model and reinforcement learning control, the problems of existing robots having difficulty determining impedance parameters and limited rewards have been solved, enabling autonomous, safe, and highly adaptable hemiplegic rehabilitation training.

CN116570468BActive Publication Date: 2025-10-28NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310597016.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-25
Publication Date
2025-10-28
Estimated Expiration
2043-05-25

AI Technical Summary

Technical Problem

Existing robots used for hemiplegic rehabilitation training have difficulty accurately determining the patient's impedance parameters, and rehabilitation training robots based on reinforcement learning have long learning times, limited rewards, and insufficient muscle activation in patients.

Method used

The system employs a combination of active and passive robots. The active robot, controlled adaptively using a reference impedance model, is worn on the patient's healthy limb. The passive robot, controlled through reinforcement learning, is worn on the paralyzed limb. By combining the patient's electromyographic signals and emotional feedback, the intensity of rehabilitation training is optimized.

Benefits of technology

It enables autonomous rehabilitation training with robot assistance, reduces human resources and economic costs, provides safe and comfortable rehabilitation training, is suitable for patients with different types and degrees of paralysis, and improves muscle activation and training effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116570468B_ABST
    Figure CN116570468B_ABST
Patent Text Reader

Abstract

This invention provides a robotic reinforcement learning mirror training control system for hemiplegic rehabilitation. It employs wearable bilateral active and passive rehabilitation robots to assist hemiplegic patients in rehabilitation training. The active robot is worn on the patient's healthy side and uses model-referenced adaptive impedance control; the passive robot is worn on the patient's paralyzed side and uses reinforcement learning control. Motion data from the active robot is transmitted to the passive robot, allowing the paralyzed limb to mimic the movements of the healthy side with the robot's assistance, autonomously completing rehabilitation training. The patient can adjust the movements of the healthy and paralyzed robots using their own muscle strength to achieve a more comfortable and appropriate intensity of rehabilitation training. This invention can effectively improve muscle activation in the paralyzed limb while ensuring patient safety, thereby enhancing the rehabilitation efficacy for hemiplegic patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical rehabilitation robots, and in particular relates to a robot reinforcement learning mirror training control system for hemiplegic rehabilitation. Background Technology

[0002] Stroke, also known as cerebrovascular accident, is an acute cerebrovascular disease caused by the sudden rupture or blockage of blood vessels in the brain, preventing blood flow and resulting in brain tissue damage. Stroke often leads to unilateral limb paralysis. If patients do not receive effective and correct rehabilitation training after paralysis, it can cause osteoporosis, muscle atrophy, and a gradual decline in physical fitness. Furthermore, prolonged incorrect rehabilitation training can easily damage the ligaments and tissues around the knee joint, leading to knee hyperextension and ankle inversion. Therefore, patients should exercise under the guidance of a rehabilitation therapist.

[0003] To save the time and money costs associated with physical therapists, rehabilitation robots have been developed. Mirror therapy has proven to be an effective treatment for hemiplegia. However, existing robots used for hemiplegic rehabilitation training struggle to accurately determine the uncertain impedance parameters of various patients; existing reinforcement learning-based rehabilitation training robots have long learning times, limited rewards, and limited muscle activation levels in patients. Summary of the Invention

[0004] Purpose of the invention: The technical problem to be solved by the present invention is to address the shortcomings of the prior art by providing a robot reinforcement learning mirror training control system for hemiplegic rehabilitation, which enables hemiplegic patients to complete rehabilitation training autonomously with the assistance of the robot, making the rehabilitation training process safer and more comfortable.

[0005] The present invention adopts the following technical solution to solve the technical problem: the system includes an active robot and a passive robot. The active robot is worn on the healthy limb of the patient and is adaptively controlled by a reference impedance model; the passive robot is worn on the paralyzed limb of the patient and is controlled by reinforcement learning.

[0006] The dynamic model formula of the system is:

[0007]

[0008]

[0009] Where q i These are the robot joint position coordinates. Let q represent an n-dimensional real matrix, i = m, s, where m represents the active robot, i represents the passive robot, and q represents the active robot. m q represents the joint position coordinates of an active robot. i This represents the joint position coordinates of the driven robot, where n represents the number of robot joints. and These represent the joint velocity and acceleration of the active robot, respectively. and These represent the joint velocity and acceleration of the driven robot, respectively.

[0010] M m (q m ) represents the inertial matrix of the active robot. This represents the centripetal torque and Coriolis torque matrix of an active robot. For the gravitational torque of the active robot, For the frictional torque of the active robot, This refers to the robot torque generated by the actuator of an active robot;

[0011] M s (q s ) represents the inertia matrix of the driven robot. This represents the centripetal torque and Coriolis torque matrix of the driven robot. For the gravitational torque of the driven robot, For the frictional torque of the driven robot, This represents the robot torque generated by the actuator of the driven robot;

[0012] This represents the interaction torque between the active robot and the healthy limb. This represents the interaction torque between the driven robot and the paralyzed limb.

[0013] The joint position and velocity information of the active robot are delayed by time T. m The torque is transmitted to the driven robot, thereby causing the paralyzed limb to follow the movement; the interaction torque between the driven robot and the paralyzed limb is delayed by time T. s The data transmitted to the active robot is represented as follows:

[0014] q sd (t)=K q q m (tT m )

[0015]

[0016] τ FLd (t)=K q τ IL (tT s )

[0017] Where t represents the current time. It is the desired position of the slave robot at the current moment. It is the expected speed of the driven robot at the current moment. It is the interaction torque transmitted from the paralyzed side to the healthy side, K q =diag(K) q1 , ..., K qn ) represents a mirror matrix, the diag function is used to construct a diagonal matrix, K qn Represents the mirror matrix K q The element in the q-th row and n-th column; to accommodate the mirror effect between the healthy and paralyzed sides, K q1 ~K qn The value can be 1 or -1.

[0018] The formula for the reference impedance model is:

[0019]

[0020] Where M md B md and K md Let these represent the desired inertia, damping, and stiffness matrices of the active robot, respectively. It is the output of the reference impedance model, representing the desired location, therefore and These represent the desired velocity and acceleration of the active robot, respectively.

[0021] The reference impedance model receives the interaction torques between the active robot, the driven robot, and the healthy and paralyzed limbs, and generates the ideal motion trajectory of the active robot. At this time, the patient's healthy limb can adjust the motion trajectory of the active robot by applying muscle force.

[0022] The controller of the active robot aims to track the ideal motion trajectory of the active robot generated by the reference impedance model, and establishes the sliding mode variable s. m , is represented as:

[0023]

[0024] in λ1 represents the error between the actual position and the desired position of the active robot, and is a constant. This represents the error between the actual speed and the expected speed of the active robot.

[0025] Estimated acceleration of active robots for:

[0026]

[0027] Where λ2 is a constant.

[0028] The controller formula for the active robot in the joint space is:

[0029]

[0030] The slave robot achieves reinforcement learning control through a reinforcement learning controller, which includes a state s, an action a, and a reward r. The state s includes the robot state and the patient's limb state, as shown in the formula:

[0031] s = [s R s H ] T

[0032] Where s R Indicates the robot's state, s H This represents the patient's limb status, and T represents the matrix transpose.

[0033] The robot state includes joint positions and joint velocities, as expressed in the formula:

[0034]

[0035] Where q mi Let q be the position of the i-th joint of the active robot, i = 1, 2, ..., n, q si Let i be the position of the i-th joint of the driven robot. Let be the joint velocity of the i-th joint of the active robot. Let be the joint velocity of the i-th joint of the driven robot;

[0036] The patient's limb condition is characterized by the amplitude of electromyographic signals on the skin surface, as shown in the formula:

[0037] s H =[E HL1 ...E HLk E IL1 ...E ILk ] T

[0038] Where E HLi E represents the electromyographic signal amplitude of the i-th muscle on the healthy side, where i = 1, 2, ..., k, and k is the number of muscles being tested. ILi The amplitude of the electromyographic signal of the i-th muscle on the paralyzed side;

[0039] The action 'a' is the output torque of the joint actuator of the driven robot, and the formula is:

[0040] a=[τ s1 ...τ sn ] T

[0041] Where, τsn This represents the actuator torque of the nth joint on the driven side.

[0042] The reward function of the reinforcement learning controller aims to maximize muscle activation in the paralyzed limbs, minimize trajectory tracking error between the active and passive robots, and minimize acceleration of the passive robot. Simultaneously, the user's emotions are incorporated into the reinforcement learning controller as an additional influencing factor to control the intensity of rehabilitation training exercises in real time. The reward function r is formulated as follows:

[0043]

[0044] Where Λ is a diagonal positive matrix representing the weights of the tracking error term in the reward function, and q sd This indicates the desired position of the driven robot. γ represents the desired acceleration of the driven robot. Ei E is the weighted value that balances the contribution of the electromyography signal of the i-th tested muscle. FLi E indicates the degree of muscle activation in the patient's healthy limb. ILi Indicates the degree of muscle activation in the affected limb; V FER V is a constant representing the patient's emotion recognition outcome; when the patient exhibits positive facial expressions, V... FER >0; When the patient exhibits a negative facial expression, V FER <0; γ represents the force exerted by the driven robot on the affected limb. u and γ A These are the weighted values ​​for the robot's driving torque and acceleration, respectively.

[0045] Furthermore, the position and speed information of the active robot are transmitted to the passive robot, thereby causing the paralyzed limb to follow the movement; the interaction torque between the passive robot and the paralyzed limb is transmitted to the active robot.

[0046] Furthermore, the active robot employs model reference adaptive impedance control, receiving the interactive torques from the master and slave robots and the patient's bilateral limbs to generate an ideal motion trajectory for the active robot. At this point, the patient's healthy side can adjust the active robot's motion trajectory by applying muscle force.

[0047] Furthermore, the state variables of the reinforcement learning controller include the joint positions and velocities of the master and slave robots, as well as the physiological electrical signals of the patient's healthy and paralyzed limbs; the action variables of the reinforcement learning controller are the output torques of the joint actuators of the slave robot; the reward function of the reinforcement learning controller is constructed with the goal of maximizing muscle activation in the paralyzed limbs, minimizing the trajectory tracking error between the master and slave robots, and minimizing the acceleration of the slave robot. In addition, this invention incorporates the user's emotions into the reinforcement learning controller as an additional influencing factor to control the intensity of rehabilitation training exercises in real time. When the patient exhibits positive emotions, the robot's movements can be appropriately enhanced; when the patient exhibits negative emotions, the training intensity should be reduced or even stopped.

[0048] Furthermore, the reinforcement learning can be implemented using algorithms such as Deep Deterministic Policy Gradient (DDPG) and Normalized Advantage Functions (NAF).

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] 1. Compared with traditional rehabilitation training methods, this invention removes physical therapists from rehabilitation training, saving rehabilitation training costs and reducing the human resources and economic pressure on patients during the rehabilitation training process.

[0051] 2. The active robot described in this invention adopts reference impedance model control, receives the interaction torque between the master and slave robots and the healthy and paralyzed limbs, and generates the ideal motion trajectory of the active robot. This allows the patient to adjust the robot's motion trajectory through the muscle strength of the healthy limb, so as to obtain a more comfortable and appropriate rehabilitation training intensity.

[0052] 3. The slave robot described in this invention adopts reinforcement learning control, which eliminates the disturbances and uncertainties existing in the robot dynamics model or impedance model, enabling the system to adapt to the rehabilitation training of patients with different types and degrees of paralysis.

[0053] 4. This invention integrates multimodal perception of the robot and the patient, incorporating the robot's motion trajectory, the patient's physiological electrical signals, and the patient's emotions into the reinforcement learning control controller of the slave robot. By setting an appropriate reward function, it can efficiently obtain the intensity of rehabilitation training exercises that can be adjusted in real time based on the patient's facial expressions. Attached Figure Description

[0054] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.

[0055] Figure 1 This is a diagram showing the overall architecture of the system of the present invention.

[0056] Figure 2 This is a schematic diagram of the information transmission method of the system of the present invention. Detailed Implementation

[0057] like Figure 1 As shown, this invention provides a robot reinforcement learning mirror training control system for hemiplegic rehabilitation. The robot for hemiplegic rehabilitation is the Aidong Rehabilitation Training Robot developed by Beijing Daai Robotics Co., Ltd. The system includes an active robot and a passive robot. The active robot is worn on the patient's healthy limb and uses model reference adaptive impedance control; the passive robot is worn on the patient's paralyzed limb and interacts with the patient using reinforcement learning control.

[0058] Preferably, the dynamic model formula of the bilateral master-slave robot system for hemiplegic rehabilitation in joint space is as follows:

[0059]

[0060]

[0061] in i = m, s, (where m represents the active robot, s represents the passive robot, and n represents the number of robot joints) are the robot joint position coordinates. and These represent joint velocity and acceleration, respectively. The inertia matrix, Here are the centripetal torque and Coriolis torque matrices. For gravitational torque, For frictional torque, This represents the robot torque generated by the actuator. This represents the interaction torque between the active robot and the healthy limb. This represents the interaction torque between the driven robot and the paralyzed limb.

[0062] like Figure 2 As shown, the joint position and velocity information of the active robot are delayed by time T. m The torque is transmitted to the slave robot, thereby causing the paralyzed limb to follow the movement; the interaction torque between the slave robot and the paralyzed limb is delayed by time T. s The data transmitted to the active robot is represented as follows:

[0063] q sd (t)=K q q m (tT m )

[0064]

[0065] τ FLd (t)=K q τ IL (tT s )

[0066] Where t represents the current time. It is the desired position of the slave robot. It is the interaction torque transmitted from the paralyzed side to the healthy side, K q =diag(K) q1 , ..., K qn ) represents the mirror matrix, adapting to the mirror effect between the healthy and paralyzed sides, K q1 ~K qn The value can be either 1 or -1. Taking lower limb rehabilitation as an example, when the degrees of freedom of movement are hip flexion / extension, knee flexion / extension, and ankle dorsiflexion / plantarflexion, K... qi =1; when the degree of freedom of motion is hip abduction / adduction, K qi =-1.

[0067] Preferably, a reference impedance model for the active robot is established, with the following formula:

[0068]

[0069] in and Let these represent the desired inertia, damping, and stiffness matrices of the active robot, respectively. This is the output of the impedance model, representing the desired location, therefore and These represent the desired velocity and acceleration, respectively.

[0070] The reference impedance model receives the interaction torques between the master and slave robots and the healthy and paralyzed limbs, generating the ideal motion trajectory of the active robot. Simultaneously, the patient's healthy side adjusts the active robot's trajectory by applying appropriate forces. When the paralyzed side experiences pain or discomfort during treatment, the healthy side can reduce muscle strength to constrain movement; conversely, when the paralyzed side experiences resistance and cannot actively complete a given task, the healthy side can increase muscle strength to expand the range of motion. Therefore, the patient can adjust the muscle strength on the healthy side to regulate the movement of both the healthy and paralyzed sides to achieve appropriate movement intensity.

[0071] The controller of the active robot aims to track the ideal motion trajectory of the active robot generated by the reference impedance model. A sliding mode variable s is established. m as follows:

[0072]

[0073] in λ1 is a constant.

[0074] Estimated acceleration of active robots for:

[0075]

[0076] Where λ2 is a constant.

[0077] The controller of the active robot aims to track the ideal motion trajectory of the active robot generated by the reference impedance model. The controller formula for the active robot in joint space is:

[0078]

[0079] Preferably, the reinforcement learning control controller of the slave robot consists of a state s, an action a, and a reward r. The state s comprises the robot state and the patient's limb state. The formula is:

[0080] s = [s R s H ] T

[0081] Where s R Indicates the robot's state, s H It indicates the patient's limb condition.

[0082] The robot's state includes joint positions and joint velocities, as expressed in the formula:

[0083]

[0084] Where q mi (i = 1, 2, ..., n) represents the position of the i-th joint of the active robot, q si (i = 1, 2, ..., n) represents the position of the i-th joint of the driven robot. For the joint speed of the active robot, The speed of the driven robot joints.

[0085] The patient's limb condition is characterized by the amplitude of electromyographic signals on the skin surface, as shown in the formula:

[0086] s H =[E HL1 ...E HLk E IL1 ...E ILk ] T

[0087] Where E HLi (i = 1, 2, ..., k, where k is the number of muscles being tested) represents the electromyographic signal amplitude of the i-th muscle on the healthy side, E ILi The amplitude of the electromyographic signal of the i-th muscle on the paralyzed side is denoted as .

[0088] The action 'a' of the reinforcement learning controller is the output torque of the joint actuator of the slave robot, and the formula is:

[0089] a=[τ s1 ...τ sn ] T

[0090] The reward function of the reinforcement learning controller aims to maximize muscle activation in the paralyzed limbs, minimize trajectory tracking error between the active and passive robots, and minimize acceleration in the passive robot. Furthermore, this invention incorporates the user's emotions into the reinforcement learning controller as an additional influencing factor to control the intensity of rehabilitation training exercises in real time. When the patient exhibits positive facial expressions, the robot's movements can be appropriately enhanced; when the patient exhibits negative facial expressions, the training intensity should be reduced or even stopped. In the reinforcement learning controller, all three factors are included in the reward function, as shown in the formula:

[0091]

[0092] Where Λ is a diagonal positive matrix representing the weights of the tracking error term in the reward function, and the parameter γ... Ei It is the weight value that balances the contribution of the electromyography signal of the i-th tested muscle. V FER V is a constant representing the patient's emotion recognition outcome; when the patient exhibits positive facial expressions, V... FER >0, take V FER =0.5; V = 0.5 when the patient exhibits a negative facial expression. FER <0, take V FER = -3.5. γ u and γ A These are the weight values ​​for the robot's driving torque and acceleration, respectively, with values ​​of 0.3 and 0.5.

[0093] Preferably, the weight values ​​of the above terms can be determined using a relative entropy inverse reinforcement learning algorithm. (Referring to Boularias, A., Kober, J. and Peters, J. (2011), “Relative entropy inverse reinforcement learning,” Proceedings of Artificial Intelligences and Statistics, pp. 20-27.)

[0094] Preferably, the reinforcement learning can be implemented using algorithms such as Deep Deterministic Policy Gradient (DDPG) and Normalized Advantage Functions (NAF). (Adapted from Mnih, V., Badia, AP, Mirza, M., Graves, A., Harley, T., Lillicrap, TP., Silver, D., and Kavukcuoglu, K. (2016), “Asynchronous methods for deep reinforcement learning,” Proceedings of International Conference Machine Learning, pp. 1928-1937.)

[0095] The aforementioned control system was validated on a motion rehabilitation training robot developed by Beijing Daai Robotics Technology Co., Ltd. The root mean square error of the robot's active side trajectory tracking was 0.02 rad, and the root mean square error of the passive side trajectory tracking was 0.05 rad. This indicates that the controller has very small trajectory tracking error and good tracking effect.

[0096] Before rehabilitation training, the patient's Functional Motor Assessment (FMA) score was 21.8±2.2, the Berg Rating Scale (BBS) score was 25.8±4.2, and the Modified Ashworth Scale (MAS) score was 2.5±0.5. After rehabilitation training, the scores on the above three scales were 29.5±2.5, 44.6±3.4, and 1.4±0.6, respectively. It should be noted that the higher FMA and BBS scores and the lower MAS score indicate that rehabilitation has improved. Therefore, this rehabilitation training system has a significant rehabilitation effect on hemiplegic patients.

[0097] This invention provides a robot reinforcement learning mirror training control system for hemiplegic rehabilitation. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.

Claims

1. A robotic reinforcement learning mirror training control system for hemiplegic rehabilitation, characterized in that, It includes an active robot and a passive robot. The active robot is worn on the healthy limb of the patient and is adaptively controlled using a reference impedance model. The passive robot is worn on the paralyzed limb of the patient and is controlled using reinforcement learning. The dynamic model formula of the system is: Where q i These are the robot joint position coordinates. Let q represent an n-dimensional real matrix, i = m, s, where m represents the active robot, i represents the passive robot, and q represents the active robot. m q represents the joint position coordinates of an active robot. i This represents the joint position coordinates of the driven robot, where n represents the number of robot joints. and These represent the joint velocity and acceleration of the active robot, respectively. and These represent the joint velocity and acceleration of the driven robot, respectively. M m (q m ) represents the inertial matrix of the active robot. This represents the centripetal torque and Coriolis torque matrix of an active robot. For the gravitational torque of the active robot, For the frictional torque of the active robot, This refers to the robot torque generated by the actuator of an active robot; M s (q s ) represents the inertia matrix of the driven robot. This represents the centripetal torque and Coriolis torque matrix of the driven robot. For the gravitational torque of the driven robot, For the frictional torque of the driven robot, This represents the robot torque generated by the actuator of the driven robot; This represents the interaction torque between the active robot and the healthy limb. This represents the interaction torque between the driven robot and the paralyzed limb; The joint position and velocity information of the active robot are delayed by time T. m The torque is transmitted to the driven robot, thereby causing the paralyzed limb to follow the movement; the interaction torque between the driven robot and the paralyzed limb is delayed by time T. s The data transmitted to the active robot is represented as follows: q sd (t)=K q q m (t-T m ) t FLd (t)=K q t IL (tT s ) Where t represents the current time. It is the desired position of the slave robot at the current moment. It is the expected speed of the driven robot at the current moment. It is the interaction torque transmitted from the paralyzed side to the healthy side, K q =diag(K) q1 ,…,K qn ) represents a mirror matrix, the diag function is used to construct a diagonal matrix, K qn Represents the mirror matrix K q The element in the q-th row and n-th column; to accommodate the mirror effect between the healthy and paralyzed sides, K q1 ~K qn The value can be 1 or -1; The formula for the reference impedance model is: Where M md B md and K md Let these represent the desired inertia, damping, and stiffness matrices of the active robot, respectively. It is the output of the reference impedance model, representing the desired location, therefore and These represent the desired velocity and acceleration of the active robot, respectively. The reference impedance model receives the interaction torques between the active robot, the driven robot, and the healthy and paralyzed limbs, and generates the ideal motion trajectory of the active robot. At this time, the patient's healthy limb can adjust the motion trajectory of the active robot by applying muscle force. The controller of the active robot aims to track the ideal motion trajectory of the active robot generated by the reference impedance model, and establishes the sliding mode variable s. m ; The sliding mode variable s m Expressed as: in λ1 represents the error between the actual position and the desired position of the active robot, and is a constant. This represents the error between the actual speed and the expected speed of the active robot. Estimated acceleration of active robots for: Where λ² is a constant; The controller formula for the active robot in the joint space is: The slave robot achieves reinforcement learning control through a reinforcement learning controller, which includes a state s, an action a, and a reward r. The state s includes the robot state and the patient's limb state, as shown in the formula: s=[s R ,s H ] T Where s R Indicates the robot's state, s H This represents the patient's limb status, and T represents the matrix transpose. The robot state includes joint positions and joint velocities, as expressed in the formula: Where q mi Let q be the position of the i-th joint of the active robot, i = 1, 2, ..., n. si Let i be the position of the i-th joint of the driven robot. Let be the joint velocity of the i-th joint of the active robot. Let be the joint velocity of the i-th joint of the driven robot; The patient's limb condition is characterized by the amplitude of electromyographic signals on the skin surface, as shown in the formula: s H =[And HL1 …AND HLk ,AND IL1 …AND ILk ] T Where E HLi E represents the electromyographic signal amplitude of the i-th muscle on the healthy side, where i = 1, 2, ..., k, and k is the number of muscles being tested. ILi The amplitude of the electromyographic signal of the i-th muscle on the paralyzed side; The action 'a' is the output torque of the joint actuator of the driven robot, and the formula is: a=[τ s1 …t sn ] T Where, τ sn This represents the actuator torque of the nth joint on the driven side; The reward function of the reinforcement learning controller aims to maximize muscle activation in the paralyzed limbs, minimize trajectory tracking error between the active and passive robots, and minimize acceleration of the passive robot. Simultaneously, the user's emotions are incorporated into the reinforcement learning controller as an additional influencing factor to control the intensity of rehabilitation training exercises in real time. The reward function r is formulated as follows: Where Λ is a diagonal positive matrix representing the weights of the tracking error term in the reward function, and q sd This indicates the desired position of the driven robot. γ represents the desired acceleration of the driven robot. Ei E is the weighted value that balances the contribution of the electromyography signal of the i-th tested muscle. FLi E indicates the degree of muscle activation in the patient's healthy limb. ILi Indicates the degree of muscle activation in the affected limb; V FER V is a constant representing the patient's emotion recognition outcome; when the patient exhibits positive facial expressions, V... FER >0; When the patient exhibits a negative facial expression, V FER <0; γ represents the force exerted by the driven robot on the affected limb. u and γ A These are the weighted values ​​for the robot's driving torque and acceleration, respectively.