Autonomous walking lower limb exoskeleton robot based on reinforcement learning

Through reinforcement learning method, the design of a twelve degrees of freedom autonomous walking lower limb exoskeleton robot has solved the problem of insufficient balance and anti-interference ability in the existing technology, achieved a more natural gait and higher robustness, and improved user independence and quality of life.

CN120131391APending Publication Date: 2025-06-13DALIAN UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510218517.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing lower limb exoskeleton robots have poor robustness in maintaining balance and anti-interference, and relying on external devices leads to users lacking self-confidence and natural gait to walk independently.

Method used

A reinforcement learning method is used to design a twelve degrees of freedom autonomously walking lower limb exoskeleton robot, which can be constructed to carry out reinforcement learning training, optimize control strategies, and deploy deep neural network models onto real robots to realize intelligent assistance and adaptive adjustment.

Benefits of technology

It realizes the autonomous walking and balance of the lower limb exoskeleton robot in different environments, enhances the user's confidence and comfort, improves independence, flexibility and quality of life, and reduces dependence on external devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120131391A_ABST
    Figure CN120131391A_ABST
Patent Text Reader

Abstract

The invention discloses an autonomous walking lower limb exoskeleton robot based on reinforcement learning, and belongs to the field of medical instruments. Firstly, a mechanical structure of the lower limb exoskeleton robot with twelve degrees of freedom is designed to adapt to wearers with different body types. The joints are driven by motors and are provided with limits, and model machine physical building is completed by using electronic devices such as a microcomputer, a router and a battery, so that the motors of the lower limb exoskeleton robot can move according to instructions. Training is carried out through a reinforcement learning method, the unknown interaction force problem existing between wearers of different weights and heights and the exoskeleton robot is simulated through randomization processing of the model, and the problem that an accurate dynamic model needs to be established in a traditional control method is avoided; the deep neural network control strategy obtained through training can effectively solve the problem of adaptability to unknown wearers, and meanwhile, different terrain environments are constructed in the simulation environment to train the exoskeleton robot, so that the terrain adaptability of the robot and the robustness of the robot to external interference are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical devices. Specifically, the present invention designs a lower limb exoskeleton robot that can walk autonomously based on reinforcement learning. Background Art

[0002] At present, with the increasing number of elderly people and stroke patients, exoskeleton robots have received widespread attention as effective rehabilitation tools. Existing lower limb exoskeleton robots generally weigh between 20 kg and 40 kg, which is less than the weight of the user. Therefore, the user will greatly interfere with the movement of the lower limb exoskeleton robot and easily fall. Therefore, the balance and anti-interference issues of the lower limb exoskeleton robot are particularly important. People have tried different solutions and conducted useful explorations:

[0003] At present, the method of maintaining balance of lower limb exoskeleton robots is mostly to maintain stability and balance by adding external devices, such as crutches, traction ropes and support frames, etc., but this limits the user's freedom of movement, making the user dependent on external devices and lacking the confidence to walk alone. At the same time, external devices will also cause unnatural movements and gaits. Lower limb exoskeleton robots that can walk autonomously without external devices mostly use dynamic methods, which have poor robustness, weak anti-interference ability, and high requirements for the training environment.

[0004] Different from the above method, the present invention uses reinforcement learning method to enable users to walk autonomously and maintain balance while wearing lower limb exoskeleton robots. The length of the thigh and calf is adjustable, and the mass can be adjusted through the simulation environment to adapt to users of different heights and weights. Compared with lower limb exoskeleton robots that rely on crutches and traction ropes, it can simulate the movement of normal walking more naturally. Walking freely in different environments such as uphill and downhill and stairs can also enhance the user's confidence and comfort, improve the user's independence, flexibility and quality of life, reduce dependence on external equipment, and have good robustness. Summary of the invention

[0005] The purpose of the present invention is to overcome the defects of the prior art and design a twelve-degree-of-freedom autonomous walking lower limb exoskeleton robot, which can enable the user to complete autonomous walking only by relying on the lower limb exoskeleton robot and the gait trajectory planning is more natural and closer to humans. The control strategy is optimized using a reinforcement learning algorithm, and the reinforcement learning model is deployed to a real prototype to achieve intelligent assistance and adaptive adjustment of the lower limb exoskeleton robot to the user, thereby providing a new method for controlling a lower limb exoskeleton robot.

[0006] To achieve the above object, the present invention adopts the following technical solution:

[0007] An autonomous walking lower limb exoskeleton robot based on reinforcement learning, where the autonomous walking lower limb exoskeleton robot is an autonomous walking lower limb exoskeleton robot with adjustable leg length, twelve degrees of freedom, and self-balancing ability;

[0008] The twelve degrees of freedom include hip abduction and adduction, hip external rotation and internal rotation, hip flexion and extension, knee flexion and extension, ankle flexion and extension, and ankle abduction and adduction;

[0009] The adjustable leg length includes the adjustment of the thigh length and the calf length;

[0010] The implementation process of the autonomous walking: By constructing a simulation environment for reinforcement learning, importing the three-dimensional model of the lower limb exoskeleton robot into the simulation environment, and performing reinforcement learning training on the motion control of the lower limb exoskeleton robot; During the reinforcement learning training process, it is necessary to add the mass and moment of inertia information of different wearers to the lower limb exoskeleton robot model, so that the reinforcement learning training control strategy can adapt to unknown wearers; And by changing the terrain of the simulation environment, the lower limb exoskeleton robot can learn to carry the wearer and achieve autonomous walking on uneven ground; After the reward function of the reinforcement learning converges, deploy the trained deep neural network model to the real exoskeleton robot.

[0011] Furthermore, the lower limb exoskeleton robot includes: a back plate 1, two hip abduction and adduction drive motors 2, two hip external rotation and internal rotation drive motors 3, two hip flexion and extension drive motors 4, adjustable thigh linkages 5, two knee flexion and extension drive motors 6, adjustable calf linkages 7, two ankle flexion and extension drive motors 8, two ankle abduction and adduction drive motors 9, a sole plate 10, and an electrical box 11. The hip abduction and adduction drive motors 2, hip external rotation and internal rotation drive motors 3, hip flexion and extension drive motors 4, knee flexion and extension drive motors 6, ankle flexion and extension drive motors 8, and ankle abduction and adduction drive motors 9 are all connected to the electrical box 11. The electrical box 11 includes a controller and electronics and a battery, which are used to control the operation of each motor in the device and respectively achieve hip abduction and adduction, hip external rotation and internal rotation, hip flexion and extension, knee flexion and extension, ankle flexion and extension, and ankle abduction and adduction actions.

[0012] Furthermore, the hip abduction and adduction drive motor 2 is disposed on the back plate 1 and is connected to the hip external rotation and internal rotation drive motor 3 through a connecting rod on the outside of the back plate 1. The hip external rotation and internal rotation drive motor 3 is connected to the hip flexion and extension drive motor 4 through a connecting rod on the front side of the hip external rotation and internal rotation drive motor 3. The knee flexion and extension drive motor 6 is connected to the hip flexion and extension drive motor 4 through an adjustable thigh connecting rod 5 below. The ankle flexion and extension drive motor 8 is connected to the knee flexion and extension drive motor 6 through an adjustable calf connecting rod 7 below. The ankle abduction and adduction drive motor 9 is connected to the ankle flexion and extension drive motor 8 through a connecting rod on the rear side. The sole plate 10 is connected to the ankle abduction and adduction drive motor 9 through a connecting rod below. The electrical box is disposed behind the back plate 1.

[0013] Furthermore, a trained deep neural network model is loaded on the controller.

[0014] Further, the specific steps of the autonomous walking implementation process are as follows:

[0015] Step 1: Export the three-dimensional model of the designed lower limb exoskeleton robot to the simulation platform for obtaining the state information of the lower limb exoskeleton robot and subsequent reinforcement learning training.

[0016] Step 2: Define a gait phase, which includes two double support phases DS and two single support phases SS in each gait cycle C T and use a sine wave to generate a reference motion to reflect the repeatability of the motion cycles involving pitch, knee, and ankle movements.

[0017] Step 3. In the motion control of the lower limb exoskeleton robot, train the reinforcement learning model: M = <S, a, T, O, R, γ>, where the state space S defines the current complete state information of the lower limb exoskeleton, which is the input of the model and includes observable information, terrain information, and privileged information. The observable information includes proprioceptive sensor data (basic pose of the robot), clock cycle signal, and speed command. The terrain information includes ground height information, ground friction coefficient, and terrain type. The privileged information includes the precise model of the terrain, the mass of the lower limb exoskeleton robot, the linear velocity of the base, the thrust torque, the trajectory error, and foot collision detection; in addition, it is necessary to randomize the weight of the robot and input it into the training of the deep neural network model as part of the privileged information to solve the problem of user weight deviation after deployment on a real lower limb exoskeleton robot. The action a defines the control actions of the exoskeleton, including the control instructions for the twelve joints of the robot. The transition dynamics T(S'|S, a) defines the probability distribution of transitioning to the new state S' after selecting the action a from the current state S. The reward function R(S, a) defines the reward returned according to the current state and action. The discount factor γ ∈ [0, 1] controls the influence of future rewards on the current decision. O represents the observation space.

[0018] Step 4. Utilize the Proximal Policy Optimization (PPO) algorithm, supplemented by the Asymmetric Actor-Critic method and the integration of privileged information during training, to complete the training of the motion control model of the lower limb exoskeleton. The policy loss is defined as:

[0019] where, is the ratio of the current policy to the old policy, and A πb (o ≤t , a t ) is the advantage function, and c 1 , c 2 are two hyperparameters used to limit the amplitude of each update.

[0020] The advantage function A πb (o ≤t , a t ) adopts the Generalized Advantage Estimation (GAE): A πb (o ≤t , a t ) = δ t + (γλ)A πb (o ≤t+1 , a t+1 )

[0021] where, δ t is the Temporal Difference (TD) error, defined as: δ t = r t + γV θ (s t+1 ) - V θ(s t ), where λ is a hyperparameter that balances bias and variance, and V θ (s t ) is the value function, representing the expected return at state s t ;

[0022] During the training process, in addition to updating the policy network, the value function also needs to be continuously updated: where is the target value;

[0023] Step 5: Since the exoskeleton system in the real world cannot fully perceive all state information, the model is deployed to a real robot by combining sensors and operating with the partially observable Markov decision process (POMDP) framework. Through the mapping of the partially observable state o t ∈ O to the distribution of the action a t ∈ A, the twelve joint angles in a t are input into each motor, enabling the exoskeleton to walk normally. The policy π(a|o ≤t ) is used to maximize the expected return:

[0024] where r t is the immediate reward obtained at time step t.

[0025] Furthermore, the reward function of reinforcement learning consists of four key parts: speed tracking reward, gait reward, contact reward, and regularization term. When defining the reward function, a tracking error metric is used, expressed as φ(e, w) = exp(-w · ||e|| 2 )

[0026] where e represents the tracking error and w is the relevant weight.

[0027] The beneficial effects of the present invention are as follows: The present invention provides a 12-degree-of-freedom lower limb exoskeleton robot with adjustable leg length, including a back plate, hip abduction and adduction drive motors, hip external rotation and internal rotation drive motors, hip flexion and extension drive motors, adjustable thigh linkages, knee flexion and extension drive motors, adjustable calf linkages, ankle flexion and extension drive motors, ankle abduction and adduction drive motors, and a sole plate. The lower limb exoskeleton robot is driven by motors to perform movements, maintaining a good gait and balance during training. The twelve degrees of freedom improve the flexibility of the lower limb exoskeleton robot, allowing for training with a more natural gait. The hip external rotation and internal rotation drive motors are arranged on the outside, solving the problem of difficult donning and doffing of the device for users. Using the reinforcement learning method for training solves the problem of poor robustness of the lower limb exoskeleton robot for autonomous walking. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly describe the accompanying drawings required for the description of the embodiments.

[0029] Figure 1 is the construction and training flowchart of an exoskeleton robot based on reinforcement learning;

[0030] Figure 2 is the schematic diagram of the mechanical structure of a lower limb exoskeleton robot;

[0031] Figure 3 is the schematic diagram of an adjustable thigh connecting rod;

[0032] Figure 4 is the schematic diagram of an adjustable calf connecting rod;

[0033] Figure 5 is the schematic diagram of the electrical box operation module;

[0034] Figure 6 is the flowchart of reinforcement learning training;

[0035] In the figure: 1 backboard, 2 hip abduction and adduction drive motor, 3 hip external rotation and internal rotation drive motor, 4 hip flexion and extension drive motor, 5 adjustable thigh connecting rod, 51 upper thigh connecting rod, 52 lower thigh connecting rod, 6 knee flexion and extension drive motor, 7 adjustable calf connecting rod, 71 upper calf connecting rod, 72 lower calf connecting rod, 8 ankle flexion and extension drive motor, 9 ankle abduction and adduction drive motor, 10 sole plate, 11 electrical box. Specific Embodiments

[0036] The following will clearly and completely describe the concept, specific structure and technical effects generated by the present invention in combination with the embodiments and the accompanying drawings, so as to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, other embodiments obtained by those skilled in the art without creative efforts shall fall within the scope of protection of the present invention.

[0037] Embodiment 1, refer to Figure 1, First, the mechanical structure of the lower limb exoskeleton robot is designed to adapt to wearers of different body types and make it more convenient to put on and take off. Twelve degrees of freedom improve the flexibility of the lower limb exoskeleton robot, enabling it to generate a more natural gait. The joints are driven by motors and limited to prevent excessive joint angles from causing harm to people. The physical prototype is built using electronic devices such as a microcomputer, router, and battery, enabling the motors of the lower limb exoskeleton robot to move according to instructions. The problem of poor robustness of the lower limb exoskeleton robot for autonomous walking is solved through training with reinforcement learning methods. The reinforcement learning model is deployed to the robot prototype to complete the autonomous walking of the lower limb exoskeleton robot.

[0038] In this embodiment, referring to Figure 2 , a lower limb exoskeleton robot with 12 degrees of freedom is provided, including a back plate 1, two hip abduction and adduction drive motors 2, two hip external rotation and internal rotation drive motors 3, two hip flexion and extension drive motors 4, an adjustable thigh link 5, two knee flexion and extension drive motors 6, an adjustable calf link 7, two ankle flexion and extension drive motors 8, two ankle abduction and adduction drive motors 9, a sole plate 10, and an electrical box 11. The hip abduction and adduction drive motors 2, hip external rotation and internal rotation drive motors 3, hip flexion and extension drive motors 4, knee flexion and extension drive motors 6, ankle flexion and extension drive motors 8, and ankle abduction and adduction drive motors 9 are all connected to the electrical box 11. The electrical box 11 includes a controller and electronics and a battery, which are used to control the operation of each motor in the device, respectively realizing hip abduction and adduction, hip external rotation and internal rotation, hip flexion and extension, knee flexion and extension, ankle flexion and extension, and ankle abduction and adduction movements. The hip abduction and adduction drive motors 2 are arranged on the back plate 1 and are connected to the hip external rotation and internal rotation drive motors 3 through a link on the outside of the back plate 1. The hip external rotation and internal rotation drive motors 3 are connected to the hip flexion and extension drive motors 4 through a link in front of the hip external rotation and internal rotation drive motors 3. The knee flexion and extension drive motors 6 are connected to the hip flexion and extension drive motors 4 through the adjustable thigh link 5 below. The ankle flexion and extension drive motors 8 are connected to the knee flexion and extension drive motors 6 through the adjustable calf link 7 below. The ankle abduction and adduction drive motors 9 are connected to the ankle flexion and extension drive motors 8 through a link at the rear. The sole plate 10 is connected to the ankle abduction and adduction drive motors 9 through a link below. The electrical box is arranged behind the back plate 1.

[0039] In this embodiment, referring to Figure 3 , the upper thigh link 51 and the lower thigh link 52 are connected by a pin, and the length of the thigh link is adjusted by passing the pin through different hole diameters. Referring to Figure 4, the upper calf link 71 and the lower calf link 72 are connected by a pin, and the length of the thigh link is adjusted by passing the pin through different hole diameters.

[0040] In this embodiment, referring to Figure 5 , the electrical box 11 is a plastic box with dimensions of 288mm×250mm×100mm, which contains a controller, a router, a battery, and other electrical components. The built-in controller is an InterNUC mini computer as the main controller. The host computer communicates with the controller computer based on WiFi. The node controllers are the hip joint driver 1, hip joint driver 2, hip joint driver 3, knee joint driver, ankle joint driver 1, and ankle joint driver 2, which execute the commands of the main controller and collect sensor data. To ensure the real-time performance of the control algorithm, the node controllers communicate with the main controller through a controller area network (CAN). The battery module in the electrical box 11 powers the entire system. The lower limb exoskeleton system monitors its current state through an encoder set and a host computer. The encoders are integrated in the joint actuators and are used to measure the current state of each joint.

[0041] In this embodiment, for the 12 motors: the hip abduction and adduction drive motor 2, hip external rotation and internal rotation drive motor 3, hip flexion and extension drive motor 4, knee flexion and extension drive motor 6, ankle flexion and extension drive motor 8, and ankle abduction and adduction drive motor 9, the joint angle limits are set according to the functional activity ranges of the human joints to prevent the motors of the lower limb exoskeleton robot from moving beyond the human joint activity ranges and causing harm to the user.

[0042] In this embodiment, referring to Figure 6 , the combination of a reinforcement learning (RL) model and a partially observable Markov decision process (POMDP) framework is used, the proximal policy optimization (PPO) algorithm is employed, and the asymmetric actor-critic method and generalized advantage estimation (GAE) are combined to optimize the motion control strategy of the lower limb exoskeleton.

[0043] First, the designed 12-degree-of-freedom lower limb exoskeleton robot is exported as a URDF to the Isaac Gym simulation platform to obtain the robot state information and for subsequent reinforcement learning training.

[0044] Secondly, a gait phase is defined, which includes two double-support phases (DS) and two single-support phases (SS) in each gait cycle. The cycle time, denoted as C T , is the duration of a complete gait cycle. A sine wave is used to generate the reference motion, which reflects the repeatability of the motion cycles involving pitch, knee, and ankle movements.

[0045] Then, in the motion control of the lower limb exoskeleton, train the reinforcement learning model (such as Figure 6 ): M = <S, a, T, O, R, γ>

[0046] Among them, the state space S defines the current complete state information of the lower limb exoskeleton, which is the input of the model, including observable information, terrain information, and privileged information. The observable information includes proprioceptive sensor data (basic pose of the robot), clock cycle signal, and speed command. The terrain information includes ground height information, ground friction coefficient, terrain type, etc. The privileged information includes the precise model of the terrain, the mass of the robot, the linear velocity of the base, the thrust torque, the trajectory error, and foot collision detection. Here, the weight of the robot is specifically randomized and input into the network training as part of the privileged information to solve the problem of user weight deviation after deployment on a real robot. The action a defines the control actions of the exoskeleton, including control instructions for the twelve joints of the robot. The transition dynamics T(S'|S, a) defines the probability distribution of transitioning to the new state S′ after selecting the action a from the current state S. The reward function R(S, a) defines the reward returned according to the current state and action. The discount factor γ ∈ [0, 1] controls the influence of future rewards on the current decision. O represents the observation space.

[0047] Use the Proximal Policy Optimization (PPO) algorithm, supplemented by the asymmetric actor-critic method and the integration of privileged information during training, to complete the training of the motion control model for the lower limb exoskeleton. The policy loss is defined as:

[0048] Among them, is the ratio of the current policy to the old policy, and A πb (o ≤t , a t ) is the advantage function, and c 1 , c 2 are two hyperparameters used to limit the amplitude of each update.

[0049] The advantage function A πb (o ≤t , a t ) adopts the Generalized Advantage Estimation GAE: A πb (o ≤t , a t ) = δ t + (γλ)A πb (o ≤t+1 , a t+1 )

[0050] Among them, δ t is the temporal difference TD error, defined as: δ t = r t + γV θ (st+1 ) - V θ (s t ), where λ is a hyperparameter that balances bias and variance, and V θ (s t ) is the value function, representing the expected return at state s t .

[0051] During the training process, in addition to updating the policy network, the value function also needs to be continuously updated: where is the target value;

[0052] Finally, since the exoskeleton system in the real world cannot fully perceive all state information, the model is deployed to a real robot by combining sensors and operating with the partially observable Markov decision process (POMDP) framework. Through the partially observable state o t ∈ O mapped to the distribution of action a t ∈ A, the policy π(a|o ≤t ) is used to maximize the expected return:

[0053] where r t is the immediate reward obtained at time step t.

[0054] In this embodiment, the reward function of reinforcement learning consists of four key parts: speed tracking reward, gait reward, contact reward, and regularization term.

[0055] When defining the reward function, a tracking error metric is used, expressed as φ(e, w) = exp(-w · ||e|| 2 )

[0056] where e represents the tracking error and w is the relevant weight. The target reference height is set to 0.93m.

[0057] Inspired by the above ideal embodiment according to the present invention, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. An autonomous walking lower limb exoskeleton robot based on reinforcement learning, characterized in that: The autonomous walking lower limb exoskeleton robot is an autonomous walking lower limb exoskeleton robot with adjustable leg length, twelve degrees of freedom and self-balancing ability; The twelve degrees of freedom include hip abduction and adduction, hip external rotation and internal rotation, hip flexion and extension, knee flexion and extension, ankle flexion and extension, ankle abduction and adduction; The adjustable leg length includes adjustment of thigh length and adjustment of calf length; The autonomous walking realization process is as follows: by constructing a simulation environment for reinforcement learning, a three-dimensional model of the lower limb exoskeleton robot is imported into the simulation environment, and reinforcement learning training is performed on the motion control of the lower limb exoskeleton robot; During the reinforcement learning training process, it is necessary to add the mass and moment of inertia information of different wearers to the lower limb exoskeleton robot model so that the reinforcement learning training control strategy can adapt to unknown wearers. By changing the terrain of the simulated environment, the lower limb exoskeleton robot can learn to carry the wearer and walk autonomously on uneven ground; after the reward function of reinforcement learning converges, the trained deep neural network model is deployed to the real exoskeleton robot.

2. The autonomous walking lower limb exoskeleton robot based on reinforcement learning according to claim 1, characterized in that: The lower limb exoskeleton robot comprises: a back plate (1), two hip joint abduction and adduction drive motors (2), two hip joint external rotation and internal rotation drive motors (3), two hip joint flexion and extension drive motors (4), an adjustable thigh link (5), two knee joint flexion and extension drive motors (6), an adjustable calf link (7), two ankle joint flexion and extension drive motors (8), two ankle joint abduction and adduction drive motors (9), a plantar plate (10) and an electrical box (11); hip joint abduction and adduction drive motors (2), hip joint The external rotation and internal rotation drive motor (3), the hip joint flexion and extension drive motor (4), the knee joint flexion and extension drive motor (6), the ankle joint flexion and extension drive motor (8), and the ankle joint abduction and adduction drive motor (9) are all connected to an electrical box (11). The electrical box (11) includes a controller, electronic devices and batteries, which are used to control the operation of each motor in the device, and respectively realize hip joint abduction and adduction, hip joint external rotation and internal rotation, hip joint flexion and extension, knee joint flexion and extension, ankle joint flexion and extension, ankle joint abduction and adduction.

3. The autonomous walking lower limb exoskeleton robot based on reinforcement learning according to claim 2, characterized in that: The hip joint abduction and adduction drive motor (2) is arranged on the back plate (1), and is connected to the hip joint external rotation and internal rotation drive motor (3) on the outside of the back plate (1) through a connecting rod; the hip joint external rotation and internal rotation drive motor (3) is connected to the hip joint flexion and extension drive motor (4) through a connecting rod at the front side of the hip joint external rotation and internal rotation drive motor (3); the knee joint flexion and extension drive motor (6) is connected to the bottom of the hip joint flexion and extension drive motor (4) through an adjustable thigh connecting rod (5); the ankle joint flexion and extension drive motor (8) is connected to the bottom of the knee joint flexion and extension drive motor (6) through an adjustable shank connecting rod (7); the ankle joint abduction and adduction drive motor (9) is connected to the rear side of the ankle joint flexion and extension drive motor (8) through a connecting rod; the plantar plate (10) is connected to the bottom of the ankle joint abduction and adduction drive motor (9) through a connecting rod; and the electrical box is arranged at the rear of the back plate (1).

4. The autonomous walking lower limb exoskeleton robot based on reinforcement learning according to claim 2, characterized in that: The controller is loaded with a trained deep neural network model.

5. The autonomous walking lower limb exoskeleton robot based on reinforcement learning according to claim 1, characterized in that: The specific steps of the autonomous walking implementation process are as follows: Step 1: Export the designed three-dimensional model of the lower limb exoskeleton robot to the simulation platform to obtain the state information of the lower limb exoskeleton robot and subsequent reinforcement learning training; Step 2: Define a gait phase. In each gait cycle C T The reference motion was generated using a sine wave, which included two double-support phases DS and two single-support phases SS, reflecting the repetitive nature of the motion cycle involving pitch, knee, and ankle; Step 3: Train the reinforcement learning model in the motion control of the lower limb exoskeleton robot: M =<S,a,T,O,R,γ> , where the state space S defines the complete current state information of the lower limb exoskeleton and is the input of the model, including observable information, terrain information and privileged information; observable information includes proprioceptive sensor data (basic posture of the robot), clock cycle signal, and speed command; terrain information includes ground height information, ground friction coefficient, and terrain type; privileged information includes an accurate model of the terrain, the mass of the lower limb exoskeleton robot, the linear velocity of the base, thrust torque, trajectory error, and foot collision detection; in addition, the robot's weight needs to be randomized as part of the privileged information input into the deep neural network model training to solve the problem of user weight deviation after deployment on the real lower limb exoskeleton robot; action a defines the control action of the exoskeleton, including the control instructions of the robot's twelve joints; transition dynamics T(S'|S,a) defines the probability distribution of transitioning to the new state S' after selecting action a from the current state S; reward function R(S,a) defines the reward returned based on the current state and action; discount factor γ∈[0,1] controls the impact of future rewards on current decisions; O represents the observation space; Step 4: Use the proximal strategy to optimize the PPO algorithm, supplemented by the asymmetric actor-critic method and the integration of privileged information during training to complete the training of the lower limb exoskeleton motion control model; the policy loss is defined as: in, is the ratio of the current strategy to the old strategy, A πb (o ≤t ,a t ) is the advantage function, c1, c2 are two hyperparameters used to limit the magnitude of each update; Advantage function A πb (o ≤t ,a t ) Using the generalized advantage function GAE:A πb (o ≤t ,a t )=δ t +(γλ)A πb (o ≤t+1 ,a t+1 ) Among them, δ t is the timing difference TD error, defined as: δ t =r t +γV θ (s t+1 )-V θ (s t ), λ is a hyperparameter that balances bias and variance, V θ (s t ) is the value function, which means that in state s t Expected return under During the training process, in addition to updating the policy network, the value function also needs to be continuously updated: in, is the target value; Step 5: Since the exoskeleton system in the real world cannot fully perceive all state information, the model is deployed on the real robot by combining sensors and partially observable Markov decision process POMDP framework operations; t ∈O maps to action a t ∈A distribution, a t The twelve joint angles in the input to each motor, so that the exoskeleton can walk normally, using the strategy π(a|o ≤t ) to maximize the expected return: Among them, r t is the immediate reward obtained at time step t.

6. The autonomous walking lower limb exoskeleton robot based on reinforcement learning according to claim 5, characterized in that: The reward function of reinforcement learning consists of four key parts: speed tracking reward, gait reward, contact reward and regularization term; when defining the reward function, a tracking error metric is used, expressed as φ(e,w)=exp(-w·||e|| 2 ), where e represents the tracking error and w is the associated weight.

Citation Information

Cited By

  • Mechanical arm self-adaptive impedance control method and system based on inner ring performance feedback and energy tank

    CN121374658A