Control method for finger rehabilitation exoskeleton combining metamorphic principle and rope driving module
By establishing a human finger dynamics model and a friction compensation model, and combining the TD3 algorithm to optimize the PID controller, the problems of friction interference and inaccurate joint motion recognition in rope-driven exoskeletons were solved. This enabled high-precision cable tension tracking and orderly joint motion, improving the control stability and human-computer interaction compliance of the exoskeleton.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEFEI UNIV OF TECH
- Filing Date
- 2026-03-27
- Publication Date
- 2026-05-12
AI Technical Summary
Existing finger rehabilitation exoskeletons suffer from problems such as inaccurate tension calculations due to friction interference in rope-driven structures, unadapted friction changes, poor control effects, and inaccurate recognition of joint motion configurations, resulting in transmission jamming and poor human-computer interaction smoothness.
A human finger dynamics model and a friction compensation model were established. The PID controller was optimized by combining the TD3 algorithm. Joint angle data was collected in real time by multi-angle sensors, the motion configuration was identified, and the PID parameters were dynamically adjusted to achieve orderly coordination between cable tension tracking and joint movement.
It improves the accuracy of cable tension tracking, enhances the smoothness of human-computer interaction, avoids transmission jamming, and improves the safety and effectiveness of rehabilitation training.
Smart Images

Figure CN122005265A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of finger rehabilitation exoskeleton robot technology, specifically a control method for a finger rehabilitation exoskeleton that combines the principle of variable cell structure with a rope drive module. Background Technology
[0002] Finger rehabilitation exoskeletons are wearable human-machine collaborative systems that generate controllable assistive torques at the wearer's finger joints. Compared to traditional manual rehabilitation therapy, finger rehabilitation exoskeletons are characterized by highly targeted rehabilitation and precise controllable training intensity. They feature a kinematic structure similar to the human hand and are typically designed with actuators, sensors, controllers, and power supplies to provide controllable assistive forces / torques to the patient's finger joints, helping them regain finger muscle strength and motor control function. After wearing the exoskeleton, patients with finger dysfunction can achieve finger flexion, extension, grasping, and release movements just like normal people, greatly improving their daily living self-care abilities and confidence in life.
[0003] For finger rehabilitation exoskeletons, significant individual variability and the susceptibility of rope-driven structures to friction interference lead to numerous shortcomings in existing technologies. Generally, current devices rely on a single sensor to collect data or simplified models to calculate the desired cable tension. However, due to nonlinear interferences such as rope-sleeve and rope-pulley friction in rope-driven systems, tension data collected by a single tension sensor is susceptible to noise, resulting in significant errors and inaccurate desired tension calculations. Furthermore, simply establishing a basic dynamic model to output the desired tension fails to accommodate individual differences in the movement angles of various finger joints and does not consider friction variations under different motion configurations. This not only results in high computational complexity and slow system response but also insufficient tension tracking accuracy. In addition, existing devices often employ conventional PID control for closed-loop tension tracking, requiring manual parameter adjustments based on experience, which is time-consuming and labor-intensive. Moreover, it cannot dynamically adapt to parameter changes in the sequential movement of multiple finger joints, leading to poor control performance. Meanwhile, existing exoskeletons do not incorporate the principle of cellular transformation into their configuration recognition mechanism, making it difficult to achieve orderly and coordinated switching between metacarpophalangeal joints and proximal and distal interphalangeal joints. This results in inaccurate joint movement configuration recognition and discontinuous tension changes during configuration transitions, which can easily lead to transmission jamming, poor human-computer interaction smoothness, and may even cause secondary damage to the patient's fingers. Summary of the Invention
[0004] To quickly and accurately obtain the desired cable tension, this invention establishes a collaborative computing mechanism combining a human finger dynamics model, a finger exoskeleton kinematics model, and a friction compensation model (including rope-rope loop and rope-pulley friction models). By combining joint angle data to determine the motion configuration and selectively matching the friction compensation model, the precise desired cable tension under different configurations is obtained. Furthermore, the TD3 algorithm is used to optimize the PID controller, dynamically outputting PID parameter adjustments to improve the accuracy of the drive motor control signal and the tension tracking precision. To address the issues of inaccurate multi-joint timing coordination control and configuration recognition, multi-angle sensors are used to collect joint angle data in real time, identifying finger motion configurations. Automatic switching is triggered when the metacarpophalangeal joint reaches its flexion or extension limit, achieving orderly movement of the metacarpophalangeal joint and the proximal and distal interphalangeal joints. Simultaneously, the configuration of the cable tension and motor angle is established for closed-loop control of the drive motor, improving the smoothness of human-machine interaction, avoiding transmission jamming, and ensuring the safety and effectiveness of rehabilitation training.
[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: A control method for a finger rehabilitation exoskeleton combining the principle of cell variation and a rope-driven module includes the following steps: S10. Collect kinematic data of the exoskeleton during flexion and extension movements through a sensor system, including cable tension data collected by a tension sensor and rotation angle data of each joint collected by an angle sensor. S11. Perform filtering preprocessing on the collected angle data and cable tension data respectively; S12. Based on the motion angle trajectory of each joint, the expected torque of each joint is calculated through the dynamic model of the human finger. Then, the force analysis and calculation of each cable is performed through the kinematic model of the finger exoskeleton. Combined with the obtained expected torque of the joint, the tension of each cable required for finger bending / extension is calculated. Finally, the friction compensation model under different configurations is obtained by judging the angle of the metacarpophalangeal joint, and the final expected tension is obtained. S13. The expected values of cable tension and joint angles under each configuration, as well as the tension and joint angles collected by the actual sensors, are input into the PID controller optimized by the TD3 algorithm. The PID parameters are obtained by training the reward function, policy network and value network of the TD3 algorithm and continuously updating the network accordingly. This enables cable tension tracking when the motor is working in speed mode. S14. Use root mean square error to evaluate the performance of the PID controller optimized by the TD3 algorithm and measure the degree of error between the expected tension and the actual value. S15. The data processing software built into the host computer compares the calculated expected tension curve with the actual tension and angle values collected by the sensor, automatically adjusts the PID controller parameters optimized by the TD3 algorithm, and controls the drive mechanism through the lower computer to adjust the tension of each cable, thereby controlling the tension of each cable.
[0006] Furthermore, in step S11, the specific method for angle data preprocessing is as follows: Linear interpolation is performed on the joint angle data collected by the angle sensor to eliminate the loss of data points during the data acquisition process. Then, a 50Hz notch filter is used to remove power frequency interference, and a Butterworth bandpass filter with a bandwidth of 20Hz-500Hz is used to filter high frequency noise, resulting in a smooth joint angle curve. The specific method for preprocessing cable tension data is as follows: The tension data collected by the tension sensor is first filtered by a hardware filtering circuit to remove low-frequency drift. Then, it is filtered by software using the same 50Hz notch filter and 20Hz-500Hz Butterworth bandpass filter as the angle data. The preprocessed tension data is then used as a window with a preset number of sampling points to extract the root mean square (RMS) of the time-domain features for feature analysis. The formula for calculating RMS is as follows: ; in, This is a cable tension signal; This refers to the window size.
[0007] Furthermore, in step S12, the formula for calculating the desired torque at any joint of the finger is: ; in, Let be the mass matrix of the finger at that joint. , , These are the angular acceleration vector, velocity vector, and angle vector of the joint, respectively. For Coriolis matrix, Joint torque caused by gravity, This is the viscous damping coefficient matrix. This is the joint stiffness coefficient matrix.
[0008] Furthermore, the friction compensation model includes a rope-rope friction model and a rope-pulley friction model; The rope-rope friction model adopts the static transmission model of Coulomb friction, and the calculation formula is as follows: ; in, To input rope tension, The coefficient for the direction of rope motion. The coefficient of friction of the rope loop, The angle of the rope loop bending; The rope-pulley friction model is based on Coulomb friction, and the calculation formula is as follows: ; in, The number of fixed pulleys, , , These are the constants identified experimentally. The wrap angle of the rope on the pulley is the sum of the wrap angles of the movable pulley and the corresponding fixed pulley during the bending process, and the sum of the wrap angles of the extended movable pulley and the corresponding fixed pulley during the extension process.
[0009] Furthermore, the formula for calculating the cable foundation tension, combined with the expected torque of the joint, is as follows: ; in, This represents the expected torque of the cable during flexion and extension of the corresponding joint of the finger. This represents the conversion relationship between the tension of the cable and the tension required for flexion and extension of the corresponding finger joint. The equivalent force arm of the cable tension generating flexion-extension torques at the metacarpophalangeal, proximal interphalangeal, and distal interphalangeal joints. These are the angle parameters of the metacarpophalangeal joints during flexion. These are the angle parameters of the metacarpophalangeal joints during extension; Final expected tensile force The formula for calculating the sum of the base tensile force and the friction compensation value is as follows: ; ; The first and third configurations represent the flexion and extension states of the metacarpophalangeal joints, respectively, while the second and fourth configurations represent the flexion and extension states of the remaining joints.
[0010] Furthermore, in step S13, the specific process of obtaining optimized PID parameters in the TD3 algorithm-optimized PID controller through training the reward function, policy network, and value network of the TD3 algorithm and continuously updating the network accordingly is as follows: The tension error, error derivative, error integral, joint angle, and actual tension are used as the state space parameters of the TD3 algorithm. These state space parameters are then used as the input parameters of the reward function of the TD3 algorithm. After processing by the reward function, the output is sent to the reinforcement learning module of the TD3 algorithm. The policy network and value network of the reinforcement learning module continuously update the parameters to obtain the optimized PID parameter gain. This gain is then input into the action space of the TD3 algorithm to adjust the PID parameter gain, and finally outputs the control signal of the PID controller.
[0011] Furthermore, the state space of the TD3 algorithm Defined as: ; in, The error between the expected tension and the actual tension. For the error derivative, For the error integral, This is the actual measured angle of the thumb metacarpophalangeal joint. The desired tension in the cable; Action space of TD3 algorithm Defined as: ; in, , , These are the incremental adjustments for the PID proportional, integral, and derivative parameters, respectively. Reward function of TD3 algorithm Designed as follows: ; in, , , These are weighting coefficients, used to balance tension error, thumb metacarpophalangeal joint angle deviation, and control energy consumption, respectively. This refers to the deviation between the actual angle and the desired angle of the thumb metacarpophalangeal joint. To control the magnitude of the signal; The formula for updating the value network is: ; In the formula, This represents the loss function value of the value network. This refers to the batch size, which is the number of samples taken during each update. For the current value network in state The following measures During the action Value estimation; For the first Instant reward for each sample; This is a discount factor used to balance the weights of current and future rewards, and its value ranges from [value range missing]. ; The target value network is used for stable training. This is the target policy network, used to output the target action for the next state; • The formula for updating the delayed update policy network is: • ; In the formula, The gradient of the objective function of the policy network; These are the parameters of the policy network; For value network to action The gradient; Output the gradient of the policy network with respect to its own parameters; For the policy network in state The action to be output; The target network for soft updates is: ; In the formula, These are the parameters of the target network, which includes the target value network and the target policy network. These are the parameters for the currently online network; This is a soft update coefficient that controls the update speed of the target network; PID controller control signals The calculation formula is: ; in, , , These are the initial proportional, integral, and derivative gains of the PID controller. For tensile force error, This is the differential of the error.
[0012] Furthermore, the policy network of the TD3 algorithm is a 4-layer deep neural network, including a state input layer, 2 hidden layers and an action output layer, with the number of neurons being 5, 64, 48 and 3 respectively. The activation function is LeakyReLU and the output layer uses the Tanh function. The value network is a 5-layer deep neural network, including an input layer, 3 hidden layers and a Q-value output layer, with the number of neurons being 8, 48, 48, 24 and 1 respectively. The activation function is LeakyReLU, and the hyperparameter λ of LeakyReLU is 0.001.
[0013] Furthermore, in step S14, the root mean square error is used. The formula is: ; in, It is the actual tension of the cable. It is the expected tension of the cable obtained through kinematics, dynamics, and friction models. It is the data length of the test sample sequence.
[0014] A finger rehabilitation exoskeleton combining the principle of cell variation with a rope-driven module is also provided, which uses the control method described above.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Overcoming core technical challenges and improving control precision and stability: Addressing the common frictional nonlinearity (rope-rope loop, rope-pulley friction) and multi-joint timing control difficulties in rope-driven exoskeletons, a closed-loop control strategy is developed by constructing a dedicated friction compensation model and variable cell configuration recognition mechanism, combined with the TD3 algorithm to optimize PID control parameters. This achieves high-precision tracking of cable tension, significantly reduces the root mean square error between the expected tension and the actual tension, effectively solves problems such as large force tracking deviation and joint movement stuttering under traditional control methods, and improves the stability and reliability of exoskeleton control. 2. Integrating Reinforcement Learning with Traditional Control to Enhance Human-Computer Interaction Compliance: The innovative TD3 algorithm optimizes the PID controller, using multi-dimensional data such as tension error, error integral and derivative, and joint angle as state inputs to dynamically adjust PID parameters, achieving a synergistic integration of reinforcement learning and traditional PID control. Compared to conventional PID control, this strategy better adapts to the dynamic characteristics of human finger movements, reduces interference between the exoskeleton and the human finger, improves wearing comfort and human-computer interaction compliance, and avoids secondary damage to the patient's fingers during training. 3. Multi-sensor fusion design ensures training safety and data accuracy: Integrating cable tension sensors and multi-joint angle sensors, a comprehensive data acquisition and preprocessing system is constructed. Signal interference is removed through notch filtering and Butterworth bandpass filtering to improve data acquisition accuracy. Simultaneously, the host computer monitors joint angles, cable tension, and control errors in real time. When abnormalities occur, it automatically alarms, adjusts parameters, or pauses training. Combined with a precise alignment design during wear, this comprehensively ensures the safety of rehabilitation training and provides reliable data support for subsequent training program optimization. 4. Based on the principle of variable cell structure, it achieves orderly and coordinated movement of multiple joints: Through angle sensors, it identifies the movement configuration of finger joints in real time and automatically completes the sequential switching of "metaphalangeal joint flexion - proximal and distal interphalangeal joint flexion - metaphalangeal joint extension - proximal and distal interphalangeal joint extension", accurately matching the natural flexion and extension movement pattern of human fingers. This configuration adaptive switching mechanism can not only simulate the movement trajectory of normal fingers, but also specifically exercise the mobility of each joint, improving the scientific nature and effectiveness of rehabilitation training.
[0016] 5. Wide applicability, significant rehabilitation effects, and convenient operation: This exoskeleton is effectively suitable for patients with finger flexion and extension disorders caused by stroke, trauma, and other reasons. It can specifically assist patients in completing finger rehabilitation training and promote the recovery of finger motor function. The device is fixed with Velcro straps, making it easy to wear. The training process can achieve automatic configuration switching and tension adjustment without manual intervention. At the same time, the host computer intuitively displays training data, allowing medical staff to monitor the training progress in real time. Balancing practicality and ease of use, it is suitable for promotion and application in clinical rehabilitation and home rehabilitation scenarios. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the structure of the finger rehabilitation exoskeleton of the present invention; Figure 2 This is a schematic diagram of the index finger rope drive device of the finger rehabilitation exoskeleton of the present invention. Figure 3 This is a schematic diagram of the thumb rope drive device of the finger rehabilitation exoskeleton of the present invention. Figure 4 This is a schematic diagram of the structure of a hand joint actuator; Figure 5 This is a flowchart of the control method for the finger rehabilitation exoskeleton of the present invention; Figure 6 The desired joint torque for the flexion and extension of the index finger and thumb; Figure 7 Comparison of force trajectory tracking between PID optimized for TD3 algorithm and ordinary PID; Figure 8 A comparison chart of the force tracking errors of the PID optimized for the TD3 algorithm and the ordinary PID. Detailed Implementation
[0018] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby providing a clearer and more explicit definition of the scope of protection of the present invention.
[0019] It should be noted that when a component is said to be "installed on" another component, it can be directly on the other component or it may be in a component that is centered on it. When a component is said to be "set on" another component, it can be directly set on the other component or it may also be in a component that is centered on it. When a component is said to be "fixed to" another component, it can be directly fixed to the other component or it may also be in a component that is centered on it.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "or / and" as used herein includes any and all combinations of one or more of the associated listed items.
[0021] This invention is an algorithmic optimization of the control method based on the applicant's prior application for a rope-driven hand rehabilitation exoskeleton with alignment capability and its control method (publication number CN118319684A). Therefore, the specific structure and working principle of the finger rehabilitation exoskeleton combining the variable cell principle and the rope-driven module described in this application can be found in the aforementioned literature. However, for the sake of clarity in explaining the technical solution of this application, the following is a brief description of the structural composition of the finger rehabilitation exoskeleton, using the structure of the thumb and index finger as examples.
[0022] A finger rehabilitation exoskeleton combining the principle of variable cell structure with a rope-driven module includes a finger rehabilitation exoskeleton body, a power module, a control system, a sensor system acquisition module, a data processing module, and a control strategy module based on TD3 algorithm-optimized PID. The finger rehabilitation exoskeleton includes a rope drive device and a hand joint actuator. The rope drive device and the hand joint actuator are flexibly connected by a cable to transmit power. The rope drive device is fixed to the human arm with Velcro straps, and the hand joint actuator is fixed to the back of the human hand and the phalanges of each finger with Velcro straps. The whole device is easy to wear and is firmly fixed without any obvious jamming or displacement.
[0023] like Figures 1 to 4 As shown, the rope drive device includes an upper arm support 31 fixed to the upper arm and a lower arm support 28 fixed to the forearm. A motor support 33 is fixedly installed at the center of the top surface of the upper arm support 31, and bearing mounting plates 5 are fixedly installed on the top of both side walls. An index finger drive motor 1 is fixed on the motor support 33. The index finger drive motor 1 is a DC drive motor, and its output shaft is fixedly connected to a driving bevel gear 2. An index finger take-up reel axle 32 is rotatably mounted on the bearing mounting plate 5 via rolling bearings. An index finger take-up reel 3 and a driven bevel gear 4 that meshes with the driving bevel gear 2 are fixed on the take-up reel axle 32. An index finger bending drive rope 11 and an index finger extension drive rope 10 for driving index finger bending and extension are respectively connected to the two sides of the groove of the index finger take-up reel 3. A thumb drive motor 30 and a thumb take-up reel 29, which are fixedly connected to the output shaft of the thumb drive motor 30, are fixedly mounted on the forearm support plate 28. Thumb bending drive rope 9 and thumb extension drive rope 27, which are used to drive thumb bending and extension, are respectively connected to both sides of the wheel groove of the thumb take-up reel 29. Tension sensors 6 are connected in series on all the cables to collect cable tension data in real time.
[0024] The hand joint actuator includes a back support base 8 fixed to the back of the hand, an index finger actuator, and a thumb actuator. The index finger actuator consists of an index finger pulley mechanism 13, an index finger parallelogram mechanism 15, and an index finger crank-slider mechanism 16. The pulley mechanism 13 is connected between the back support base 8 and the proximal phalanx sleeve 21, achieving the alignment of the metacarpophalangeal (MCP) joint axis through an asymmetric rope pulley structure. The index finger parallelogram mechanism 15 is connected between the proximal phalanx sleeve 21 and the middle phalanx sleeve 20, achieving automatic alignment of the proximal interphalangeal (PIP) joint. The crank-slider mechanism 16 is connected between the middle phalanx sleeve 20 and the distal phalanx sleeve 19 and is drivenly connected to the parallelogram mechanism 15, achieving automatic alignment of the distal interphalangeal (DIP) joint. The thumb action execution mechanism includes a thumb pulley mechanism and a thumb parallelogram mechanism 23, which are respectively connected between the back of the hand support seat 8 and the thumb proximal phalanx sleeve 24, and the thumb proximal phalanx sleeve 24 and the thumb distal phalanx sleeve 22, to achieve precise alignment of the thumb metacarpophalangeal (MCP) joint and the interphalangeal (IP) joint.
[0025] The sensor system acquisition module includes a tension sensor 6 mounted on the cable, a first angle sensor 14 mounted at the hinge of the index finger parallelogram mechanism 15, a second angle sensor 17 mounted at the hinge of the crank-slider mechanism 16, a third angle sensor 18 mounted on the distal phalanx finger sleeve 19, and a fourth angle sensor 25 mounted at the hinge of the thumb parallelogram mechanism 23. The tension sensor 6 collects cable tension data at a sampling rate, and each angle sensor collects corresponding joint rotation angle data. Together, they obtain the joint angle positions during the movement of the index finger and thumb, as well as the real-time tension information of each cable, providing data support for the control strategy.
[0026] The control system includes a host computer (PC) and a slave computer (microcontroller). The host computer has built-in MATLAB data processing software and communicates with the slave computer via serial port with a preset baud rate of 115200. It is responsible for data reception, storage, and model calculation. The slave computer is connected to the data processing module, index finger drive motor 1, and thumb drive motor 30. The data processing module is connected to the sensor acquisition module and is responsible for filtering, converting, and transmitting sensor data. The microcontroller outputs PWM signals according to the control commands from the host computer to drive the motors and achieve closed-loop speed control.
[0027] The control strategy refers to the error between the actual tension collected by sensors and the desired tension of each cable obtained through the dynamic model of the human finger, the kinematic model of the finger exoskeleton, and the friction compensation model. A PID controller optimized with the TD3 algorithm is used to perform closed-loop tension tracking control on the drive motor, and the control tension of the cable is adjusted in real time. Then, the finger bending / extension configuration is determined by the angle sensor. Based on the cell variation principle, the automatic sequential switching from metacarpophalangeal joint flexion to proximal and distal interphalangeal joint flexion, then to metacarpophalangeal joint extension, and finally to proximal and distal interphalangeal joint extension is realized to achieve compliant control of the exoskeleton.
[0028] See appendix Figure 5 A control method for a finger rehabilitation exoskeleton combining the principle of cell variation and a rope-driven module includes the following steps: S10. Collect kinematic data of the exoskeleton during flexion and extension movements through a sensor system, including cable tension data collected by a tension sensor and rotation angle data of each joint collected by an angle sensor.
[0029] S11. Perform filtering preprocessing on the collected angle data and cable tension data respectively.
[0030] (1) The specific method for angle data preprocessing is as follows: Linear interpolation was performed on the joint angle data collected by various angle sensors to eliminate the loss of data points during the data acquisition process. Then, a 50Hz notch filter was used to remove power frequency interference, and a Butterworth bandpass filter with a bandwidth of 20Hz-500Hz was used to filter high frequency noise, resulting in a smooth joint angle curve.
[0031] (2) The specific method for preprocessing cable tension data is as follows: For the tension data collected by tension sensor 6, low-frequency drift is first removed by hardware filtering circuit, and then software filtering is performed using the same 50Hz notch filter and 20Hz-500Hz Butterworth bandpass filter as the angle data. The root mean square (RMS) of the time-domain features of the preprocessed tension data is extracted using a preset number of sampling points as a window for feature analysis. The formula for calculating RMS is as follows: ; in, This is a cable tension signal; This represents the window size; in this example, the value is 30.
[0032] S12. Based on the motion angle trajectories of the MCP, PIP, and DIP joints of the index finger and the MCP and IP joints of the thumb, the desired torque of each joint is calculated using a dynamic model of the human finger. Figure 6As shown. Then, the force analysis and calculation of each cable are performed by the finger exoskeleton kinematic model. Combined with the obtained joint expected torque, the tension of each cable required for finger bending / extension is calculated. Finally, the friction compensation model under different configurations is obtained by judging the metacarpophalangeal joint angle, and the final expected tension is obtained.
[0033] The formula for calculating the desired torque at any joint of the finger is: ; in, Let be the mass matrix of the finger at that joint. , , These are the angular acceleration vector, velocity vector, and angle vector of the joint, respectively. For Coriolis matrix, Joint torque caused by gravity, This is the viscous damping coefficient matrix. This is the joint stiffness coefficient matrix.
[0034] Friction compensation models include rope-rope friction models and rope-pulley friction models.
[0035] The rope-rope friction model adopts the static transmission model of Coulomb friction, and the calculation formula is as follows: ; in, To input rope tension, The coefficient for the direction of rope motion. The coefficient of friction of the rope loop, The angle of the rope loop bending; The rope-pulley friction model is based on Coulomb friction, and the calculation formula is as follows: ; in, The number of fixed pulleys, , , These are the constants identified experimentally. The wrap angle of the rope on the pulley is the sum of the wrap angles of the movable pulley and the corresponding fixed pulley during the bending process, and the sum of the wrap angles of the extended movable pulley and the corresponding fixed pulley during the extension process.
[0036] The formula for calculating the cable foundation tension based on the expected moment of the joint is as follows: ; in, This represents the expected torque of the cable during flexion and extension of the corresponding joint of the finger. This represents the conversion relationship between the tension of the cable and the tension required for flexion and extension of the corresponding finger joint. The equivalent force arm of the cable tension generating flexion-extension torques at the metacarpophalangeal, proximal interphalangeal, and distal interphalangeal joints. These are the angle parameters of the MCP joint during flexion. These are the angle parameters of the MCP joint during extension; The final expected tensile force is the sum of the base tensile force and the friction compensation value, and the calculation formula is as follows: ; ; in, For the desired cable tension, the first and third configurations represent the bending and extension states of the MCP joint, respectively, and the second and fourth configurations represent the bending and extension states of the remaining joints (DIP joint, PIP joint, or IP joint).
[0037] S13. The calculated expected values of cable tension and joint angles under each configuration, as well as the actual tension and joint angles collected by the sensors, are input into the PID controller optimized by the TD3 algorithm. The PID parameters are obtained by training the reward function, policy network and value network of the TD3 algorithm and continuously updating the network accordingly. This enables cable tension tracking when the motor is working in speed mode.
[0038] In the PID controller optimized by the TD3 algorithm, the specific process of training and continuously updating the network using the reward function, policy network, and value network of the TD3 algorithm to obtain the optimized PID parameters is as follows: The tension error, error derivative, error integral, joint angle, and actual tension are used as the state space parameters of the TD3 algorithm. These state space parameters are then used as the input parameters of the reward function of the TD3 algorithm. After processing by the reward function, the output is sent to the reinforcement learning module of the TD3 algorithm. The policy network and value network of the reinforcement learning module continuously update the parameters to obtain the optimized PID parameter gain. This gain is then input into the action space of the TD3 algorithm to adjust the PID parameter gain, and finally outputs the control signal of the PID controller.
[0039] In this embodiment, the state space of the TD3 algorithm Defined as: ; in, The error between the expected tension and the actual tension. For the error derivative, For the error integral, This is the actual measured angle of the thumb metacarpophalangeal joint. The desired tension in the cable; Action space of TD3 algorithm Defined as: ; in, , , These are the incremental adjustments for the PID proportional, integral, and derivative parameters, respectively. Reward function of TD3 algorithm Designed as follows: ; in, , , These are weighting coefficients, used to balance tension error, thumb metacarpophalangeal joint angle deviation, and control energy consumption, respectively. This refers to the deviation between the actual angle and the desired angle of the thumb metacarpophalangeal joint. To control the magnitude of the signal; The formula for updating the value network (Critic) is: ; In the formula, This represents the loss function value of the value network. This refers to the batch size, which is the number of samples taken during each update. For the current value network in state The following measures During the action Value estimation; For the first Instant reward for each sample; This is a discount factor used to balance the weights of current and future rewards, and its value ranges from [value range missing]. ; The target value network is used for stable training. This is the target policy network, used to output the target action for the next state.
[0040] The formula for updating the delayed update policy (Actor) network is: • ; In the formula, The gradient of the objective function of the policy network; These are the parameters of the policy network; For value network to action The gradient; Output the gradient of the policy network with respect to its own parameters; For the policy network in state The action to output.
[0041] Soft update target network (Polyak average, )for: ; In the formula, These are the parameters of the target network, which includes the target value network and the target policy network. These are the parameters for the currently online network; This is the soft update coefficient, which controls the update speed of the target network.
[0042] The policy network (Actor) of the TD3 algorithm is a 4-layer deep neural network, including a state input layer, 2 hidden layers, and an action output layer, with the number of neurons being 5, 64, 48, and 3 respectively. The activation function is LeakyReLU, and the output layer uses the Tanh function. The value network (Critic) is a 5-layer deep neural network, including an input layer, 3 hidden layers, and a Q-value output layer, with the number of neurons being 8, 48, 48, 24, and 1 respectively. The activation function for all layers is LeakyReLU, and the hyperparameter λ of LeakyReLU is set to 0.001.
[0043] PID controller control signals The calculation formula is: ; in, , , These are the initial proportional, integral, and derivative gains of the PID controller. For tensile force error, This is the differential of the error.
[0044] S14. Use the root mean square error to evaluate the performance of the PID controller optimized by the TD3 algorithm and measure the degree of error between the expected tension and the actual value.
[0045] In this embodiment, the root mean square error is used. The formula is: ; in, It is the actual tension of the cable. It is the expected tension of the cable obtained through kinematics, dynamics, and friction models. It is the data length of the test sample sequence.
[0046] S15. The data processing software built into the host computer compares the calculated expected tension curve with the actual tension and angle values collected by the sensor, automatically adjusts the PID controller parameters optimized by the TD3 algorithm, and controls the drive mechanism through the lower computer to adjust the tension of each cable, thereby controlling the tension of each cable.
[0047] This invention also provides a finger rehabilitation exoskeleton combining the principle of cell variation with a rope-driven module, which employs the control method described above. The method for performing finger rehabilitation training using this exoskeleton includes the following steps: S20: Before rehabilitation training, the subject performs 5-10 minutes of relaxation activities, focusing on the index finger, thumb, and hand joints. The skin surface of the back of the hand and fingers is cleaned and hair is removed to ensure that the skin is dry and clean. The connection stability and integrity of the tension sensor and the first to fourth angle sensors are checked. Then, the finger rehabilitation exoskeleton is put on, and the upper arm support plate, forearm support plate, back of hand support plate, and finger sleeves of each finger bone are fixed to the corresponding parts by Velcro straps. It is ensured that the pulley mechanism, parallelogram mechanism, crank slider mechanism of the exoskeleton are precisely aligned with the MCP, PIP, DIP and IP joints of the human finger without any jamming or displacement.
[0048] S21: Activate the exoskeleton emergency stop switch and the lower-level machine switch. The sensor system acquisition module starts working, and the upper-level machine runs the program to initialize the system. During initialization, the upper-level machine automatically detects the communication links between each sensor and the lower-level machine, the motor drive status, and the stability of the serial port connection. If any abnormality is found, an audible and visual alarm is issued and the fault type is displayed. The system restarts after the fault is cleared. Simultaneously, the tension sensor is zero-point calibrated, and the angle sensor's range is calibrated to ensure that the data acquisition accuracy meets the training requirements.
[0049] S22: The sensor system acquisition module begins to acquire data in real time. The host computer's built-in data processing software performs filtering and preprocessing on the acquired angle data and cable tension data. The tension sensor acquires the tension data of each cable, and the first to fourth angle sensors synchronously acquire the corresponding joint rotation angle data. All acquired data is transmitted to the lower computer through the data processing module, and then the lower computer uploads it to the host computer for storage and preliminary processing at a preset baud rate. Finally, the data is processed by the built-in MATLAB for subsequent model training and calculation.
[0050] S23: The expected values of cable tension and joint angle under various configurations, obtained through pre-modeling and calculation using MATLAB, are input together with the actual tension and joint angle data collected in real time by sensors into the PID controller optimized by the TD3 algorithm. Simultaneously, the host computer monitors the MCP joint angle and cable tension in real time. When the MCP joint angle reaches the 48° buckling limit or 0° extension limit, a configuration transition is triggered. Finally, the index finger drive motor and thumb drive motor are controlled to perform corresponding actions, achieving high-precision tracking of cable tension. The results are then compared with those of traditional PID control. Figure 7 and Figure 8 As shown in the figure, it can be seen that, compared with the PID controller of the transmission, the PID controller optimized by the TD3 algorithm of this invention has a significantly more stable and accurate force tracking error.
[0051] S24: During rehabilitation training, the host computer displays the actual and expected values of the joint angles and cable tension in real time, calculates and displays the root mean square error, and monitors the training progress and control accuracy in real time. If the error exceeds the preset range, the PID controller parameters optimized by the TD3 algorithm are automatically adjusted or the training is paused and a prompt is issued. Training resumes after the parameters are adjusted or the abnormality is eliminated.
[0052] S25: Repeat steps S23 and S24 until the subject completes the rehabilitation training.
[0053] S26: After training, turn off the host computer and slave computer switches, the program stops running, the sensor system acquisition module stops working, the subject removes the finger rehabilitation exoskeleton, and cleans the residual impurities on the surface of each sensor.
[0054] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0055] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A control method for a finger rehabilitation exoskeleton combining the principle of cell variation and a rope-driven module, characterized in that, Includes the following steps: S10. Collect kinematic data of the exoskeleton during flexion and extension movements through a sensor system, including cable tension data collected by a tension sensor and rotation angle data of each joint collected by an angle sensor. S11. Perform filtering preprocessing on the collected angle data and cable tension data respectively; S12. Based on the motion angle trajectory of each joint, the expected torque of each joint is calculated through the dynamic model of the human finger. Then, the force analysis and calculation of each cable is performed through the kinematic model of the finger exoskeleton. Combined with the obtained expected torque of the joint, the tension of each cable required for finger bending / extension is calculated. Finally, the friction compensation model under different configurations is obtained by judging the angle of the metacarpophalangeal joint, and the final expected tension is obtained. S13. The expected values of cable tension and joint angles under each configuration, as well as the tension and joint angles collected by the actual sensors, are input into the PID controller optimized by the TD3 algorithm. The PID parameters are obtained by training the reward function, policy network and value network of the TD3 algorithm and continuously updating the network accordingly. This enables cable tension tracking when the motor is working in speed mode. S14. Use root mean square error to evaluate the performance of the PID controller optimized by the TD3 algorithm and measure the degree of error between the expected tension and the actual value. S15. The data processing software built into the host computer compares the calculated expected tension curve with the actual tension and angle values collected by the sensor, automatically adjusts the PID controller parameters optimized by the TD3 algorithm, and controls the drive mechanism through the lower computer to adjust the tension of each cable, thereby controlling the tension of each cable.
2. The control method for the finger rehabilitation exoskeleton combining the principle of cell variation and the rope-driven module according to claim 1, characterized in that: In step S11, the specific method for angle data preprocessing is as follows: Linear interpolation is performed on the joint angle data collected by the angle sensor to eliminate the loss of data points during the data acquisition process. Then, a 50Hz notch filter is used to remove power frequency interference, and a Butterworth bandpass filter with a bandwidth of 20Hz-500Hz is used to filter high frequency noise, resulting in a smooth joint angle curve. The specific method for preprocessing cable tension data is as follows: The tension data collected by the tension sensor is first filtered by a hardware filtering circuit to remove low-frequency drift. Then, it is filtered by software using the same 50Hz notch filter and 20Hz-500Hz Butterworth bandpass filter as the angle data. The preprocessed tension data is then used as a window with a preset number of sampling points to extract the root mean square (RMS) of the time-domain features for feature analysis. The formula for calculating RMS is as follows: ; in, This is a cable tension signal; This refers to the window size.
3. The control method for the finger rehabilitation exoskeleton combining the principle of cell variation and the rope-driven module according to claim 1, characterized in that: In step S12, the formula for calculating the desired torque at any joint of the finger is: ; in, Let be the mass matrix of the finger at that joint. , , These are the angular acceleration vector, velocity vector, and angle vector of the joint, respectively. For Coriolis matrix, Joint torque caused by gravity, This is the viscous damping coefficient matrix. This is the joint stiffness coefficient matrix.
4. The control method for the finger rehabilitation exoskeleton combining the cell-shifting principle and the rope-driven module according to claim 3, characterized in that: The friction compensation model includes a rope-rope friction model and a rope-pulley friction model; The rope-rope friction model adopts the static transmission model of Coulomb friction, and the calculation formula is as follows: ; in, To input rope tension, The coefficient for the direction of rope motion. The coefficient of friction of the rope loop, The angle of the rope loop bending; The rope-pulley friction model is based on Coulomb friction, and the calculation formula is as follows: ; in, The number of fixed pulleys, , , These are the constants identified experimentally. The wrap angle of the rope on the pulley is the sum of the wrap angles of the movable pulley and the corresponding fixed pulley during the bending process, and the sum of the wrap angles of the extended movable pulley and the corresponding fixed pulley during the extension process.
5. The control method for the finger rehabilitation exoskeleton combining the principle of cell variation and the rope-driven module according to claim 4, characterized in that: The formula for calculating the cable foundation tension based on the expected moment of the joint is as follows: ; in, This represents the expected torque of the cable during flexion and extension of the corresponding joint of the finger. This represents the conversion relationship between the tension of the cable and the tension required for flexion and extension of the corresponding finger joint. The equivalent force arm of the cable tension generating flexion-extension torques at the metacarpophalangeal, proximal interphalangeal, and distal interphalangeal joints. These are the angle parameters of the metacarpophalangeal joints during flexion. These are the angle parameters of the metacarpophalangeal joints during extension; Final expected tensile force The formula for calculating the sum of the base tensile force and the friction compensation value is as follows: ; ; The first and third configurations represent the flexion and extension states of the metacarpophalangeal joints, respectively, while the second and fourth configurations represent the flexion and extension states of the remaining joints.
6. The control method for the finger rehabilitation exoskeleton combining the principle of cell variation and the rope-driven module according to any one of claims 1 to 5, characterized in that: In step S13, the specific process of obtaining optimized PID parameters in the TD3 algorithm-optimized PID controller through training using the TD3 algorithm's reward function, policy network, and value network, and continuously updating the network accordingly, is as follows: The tension error, error derivative, error integral, joint angle, and actual tension are used as the state space parameters of the TD3 algorithm. These state space parameters are then used as the input parameters of the reward function of the TD3 algorithm. After processing by the reward function, the output is sent to the reinforcement learning module of the TD3 algorithm. The policy network and value network of the reinforcement learning module continuously update the parameters to obtain the optimized PID parameter gain. This gain is then input into the action space of the TD3 algorithm to adjust the PID parameter gain, and finally outputs the control signal of the PID controller.
7. The control method for the finger rehabilitation exoskeleton combining the cell-variant principle and the rope-driven module according to claim 6, characterized in that: The state space of the TD3 algorithm Defined as: ; in, The error between the expected tension and the actual tension. For the error derivative, For the error integral, This is the actual measured angle of the thumb metacarpophalangeal joint. The desired tension in the cable; Action space of TD3 algorithm Defined as: ; in, , , These are the incremental adjustments for the PID proportional, integral, and derivative parameters, respectively. Reward function of TD3 algorithm Designed as follows: ; in, , , These are weighting coefficients, used to balance tension error, thumb metacarpophalangeal joint angle deviation, and control energy consumption, respectively. This refers to the deviation between the actual angle and the desired angle of the thumb metacarpophalangeal joint. To control the magnitude of the signal; The formula for updating the value network is: ; In the formula, This represents the loss function value of the value network. This refers to the batch size, which is the number of samples taken during each update. For the current value network in state The following measures During the action Value estimation; For the first Instant reward for each sample; This is a discount factor used to balance the weights of current and future rewards, and its value ranges from [value range missing]. ; The target value network is used for stable training. This is the target policy network, used to output the target action for the next state; • The formula for updating the delayed update policy network is: ; In the formula, The gradient of the objective function of the policy network; These are the parameters of the policy network; For value network to action The gradient; Output the gradient of the policy network with respect to its own parameters; For the policy network in state The action to be output; The target network for soft updates is: ; In the formula, These are the parameters of the target network, which includes the target value network and the target policy network. These are the parameters for the currently online network; This is a soft update coefficient that controls the update speed of the target network; PID controller control signals The calculation formula is: ; in, , , These are the initial proportional, integral, and derivative gains of the PID controller. For tensile force error, This is the differential of the error.
8. The control method for the finger rehabilitation exoskeleton combining the principle of cell variation and the rope-driven module according to claim 1, characterized in that: The policy network of the TD3 algorithm is a 4-layer deep neural network, including a state input layer, 2 hidden layers and an action output layer, with the number of neurons being 5, 64, 48 and 3 respectively. The activation function is LeakyReLU and the output layer uses the Tanh function. The value network is a 5-layer deep neural network, including an input layer, 3 hidden layers and a Q-value output layer, with the number of neurons being 8, 48, 48, 24 and 1 respectively. The activation function is LeakyReLU, and the hyperparameter λ of LeakyReLU is 0.
001.
9. The control method for a finger rehabilitation exoskeleton combining the principle of cell variation and a rope-driven module according to any one of claims 1 to 5, or 7 or 8, is characterized in that: In step S14, the root mean square error is used. The formula is: ; in, It is the actual tension of the cable. It is the expected tension of the cable obtained through kinematics, dynamics, and friction models. It is the data length of the test sample sequence.
10. A finger rehabilitation exoskeleton combining the principle of cell variation with a rope-driven module, characterized in that: The finger rehabilitation exoskeleton uses the control method described in any one of claims 1 to 9.