A method and device for controlling a hexapod robot leg-arm multiplexing based on reinforcement learning

By using a reinforcement learning-based control method to acquire observation information and determine joint torques using PD control, the leg and arm reuse of a hexapod robot is realized. This solves the problems of reduced load and energy loss in hexapod robots during tasks and achieves efficient integration of movement and manipulation.

CN120993711BActive Publication Date: 2026-02-03HARBIN INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511508291.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-02-03
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Existing hexapod robots cannot reuse their legs and arms during tasks, resulting in reduced payload and energy loss, making it impossible to efficiently complete movement and manipulation tasks.

Method used

A reinforcement learning-based control method is adopted. By acquiring observation information and inputting it into the control strategy model, the PD control method is used to determine the joint torque, thereby realizing the leg and arm reuse control of the hexapod robot. The strategy is updated by combining a multi-task training environment and a value network, and the movement and operation modes are integrated.

Benefits of technology

It improves the integration accuracy of the leg-arm reuse control of the hexapod robot, solves the problems of reduced payload and energy loss, realizes the deep integration of movement and operation, and achieves smooth switching between different control modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120993711B_ABST
    Figure CN120993711B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on reinforcement learning's hexapod robot leg arm multiplex control method and device, it is related to motion control field.The method includes: obtaining the observation information of target hexapod robot;Observation information is input to control strategy model, and the expected joint angle of target hexapod robot is obtained;Control strategy model is obtained according to the policy update of training architecture based on value network after using the method of reinforcement learning according to the multi-task training environment constructed;Training architecture is determined based on teacher-student privilege learning;Value network includes successively connected shared feature layer and multi-task head;Training architecture includes: state estimation encoder, terrain information encoder, privilege information encoder, historical information encoder and ontology network;Using PD control method based on expected joint angle determines joint torque, to control target hexapod robot leg arm multiplex.This application can realize the leg arm multiplex control of hexapod robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of motion control, and in particular to a method and apparatus for reusing the legs and arms of a hexapod robot based on reinforcement learning. Background Technology

[0002] Legged robots are expected to operate in complex and rugged environments. Typically, tasks take two forms: locomotion and manipulation. Humanoid robots decouple these two tasks through their unique configuration (arms and legs), while quadrupedal and hexapod robots, which lack arms, can perform manipulation tasks by adding additional robotic arms. However, these methods dedicate the arms to manipulation tasks and the legs to locomotion tasks. The robotic arms, when not in operation, reduce the robot system's payload and generate unnecessary energy loss. Therefore, achieving leg-arm reuse control in hexapod robots is crucial. Summary of the Invention

[0003] The purpose of this application is to provide a reinforcement learning-based method and apparatus for the multiplexing control of the legs and arms of a hexapod robot, which can realize the multiplexing control of the legs and arms of a hexapod robot.

[0004] To achieve the above objectives, this application provides the following solution:

[0005] In a first aspect, this application provides a reinforcement learning-based method for the multiplexing control of the legs and arms of a hexapod robot, including:

[0006] Obtain observational information about the target hexapod robot;

[0007] The observed information is input into the control strategy model to obtain the desired joint angles of the target hexapod robot. The control strategy model is obtained by updating the training architecture based on the value network using reinforcement learning, based on the constructed multi-task training environment. The training architecture is determined based on teacher-student privileged learning. The value network includes a shared feature layer and a multi-task head connected in sequence. The training architecture includes: a state estimation encoder, a terrain information encoder, a privileged information encoder, a history information encoder, and an ontology network.

[0008] The PD control method is used to determine the joint torque based on the desired joint angle in order to control the reuse of the legs and arms of the target hexapod robot.

[0009] Secondly, this application provides a reinforcement learning-based hexapod robot leg-arm multiplexing control device, comprising:

[0010] The information acquisition module is used to acquire observation information of the target hexapod robot;

[0011] The desired joint angle determination module is used to input the observation information into the control strategy model to obtain the desired joint angles of the target hexapod robot. The control strategy model is obtained by updating the training architecture based on the value network using reinforcement learning, based on the constructed multi-task training environment. The training architecture is determined based on teacher-student privileged learning. The value network includes a shared feature layer and a multi-task head connected in sequence. The training architecture includes a state estimation encoder, a terrain information encoder, a privileged information encoder, a history information encoder, and an ontology network.

[0012] The leg-arm reuse control module is used to determine the joint torque based on the desired joint angle using the PD control method, so as to control the leg-arm reuse of the target hexapod robot.

[0013] According to the specific embodiments provided in this application, the following technical effects are disclosed:

[0014] This application provides a reinforcement learning-based method and apparatus for leg-arm reuse control of a hexapod robot. By acquiring observation information of the target hexapod robot, this information is input into a control strategy model to obtain the desired joint angles. A PD control method is then used to determine joint torques based on these desired joint angles, enabling the hexapod robot to perform leg-arm reuse control tasks. This addresses the current problem of task decoupling and the inability of hexapod robots to reuse their legs and arms. Furthermore, the control strategy model in this application is obtained by updating the training architecture based on a value network using a reinforcement learning method within a constructed multi-task training environment. The training architecture is determined based on teacher-student privileged learning. The value network includes a shared feature layer and a multi-task head connected sequentially. The training architecture includes a state estimation encoder, a terrain information encoder, a privileged information encoder, a history information encoder, and an ontology network. Based on this control strategy model, the integration accuracy of leg-arm reuse control for hexapod robots can be improved, solving the problems of reduced effective load and energy loss under different operating conditions, thus enabling effective control of leg-arm reuse in the target hexapod robot. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 The flowchart shows a reinforcement learning-based control method for the reuse of legs and arms in a hexapod robot.

[0017] Figure 2A flowchart of the control strategy technology in the reinforcement learning-based hexapod robot leg and arm reuse control method.

[0018] Figure 3 This is a schematic diagram of the actual deployment of the control strategy in the reinforcement learning-based hexapod robot leg and arm reuse control method.

[0019] Figure 4 This is a structural diagram of a reinforcement learning-based hexapod robot leg and arm reuse control device. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] Leg-arm reuse is a more efficient biomimetic approach that gives legs and feet the duality of movement and manipulation, enabling them to handle the tasks of legged robots more efficiently.

[0022] This application proposes a reinforcement learning-based leg-arm reuse control method and device for a hexapod robot. While enabling base movement, it allows for the independent control of any one or two legs to reach a specific target or trajectory, achieving a deep integration of movement and manipulation control. Furthermore, it integrates three motion modes—robust movement, maneuvering, and stationary maneuvering (i.e., robust base movement mode, stationary whole-body manipulation mode, and dynamic moving maneuvering mode)—into a single control strategy. This unified control strategy eliminates the need for controller switching, enabling smooth transitions between different control modes.

[0023] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0024] In one exemplary embodiment, such as Figure 1 As shown, a reinforcement learning-based method for multiplexing the legs and arms of a hexapod robot is provided, including:

[0025] Step 100: Obtain observation information of the target hexapod robot.

[0026] Step 200: Input the observation information into the control policy model to obtain the desired joint angles of the target hexapod robot. The control policy model is obtained by updating the training architecture based on the value network using reinforcement learning, based on the constructed multi-task training environment. The training architecture is determined based on teacher-student privileged learning. The value network includes a shared feature layer and a multi-task head connected in sequence. The training architecture includes: a state estimation encoder, a terrain information encoder, a privileged information encoder, a history information encoder, and an ontology network.

[0027] The multi-task training environment is defined by adopting a curriculum-based approach and adding a robust training mechanism; the multi-task training environment includes basic motor control task mode, static whole-body operation task mode, and dynamic motion-coupled operation task mode.

[0028] Among them, the basic motion control task mode achieves dynamic tracking of target linear velocity and angular velocity by generating rhythmic gait patterns; the static whole-body operation task mode maintains a stable support posture while specifying the end effector of the limb to track the preset target position, determines all feasible areas in the leg workspace through forward kinematics and collision detection algorithms, and randomly selects a set number of end effectors to track the trajectory of randomly sampled target points; the dynamic motion coupling operation task mode synchronously executes the base velocity command and the joint trajectory tracking of the specified foot end effector.

[0029] The robust training mechanism simulates external disturbances by applying an impulse in any direction to the root of the target hexapod robot during base movement training; during operation training, a combination of omnidirectional force and torque is used to randomly load the foot actuators to simulate load interaction scenarios.

[0030] Step 300: The PD control method is used to determine the joint torque based on the desired joint angle in order to control the reuse of the target hexapod robot's legs and arms.

[0031] The mathematical expression corresponding to the PD control method is:

[0032] .

[0033] in, Joint torque; This is the proportionality coefficient; For the desired joint position; For location feedback; These are the differential coefficients; For speed feedback.

[0034] In one embodiment, the observation information specifically includes: command observation information, ontological observation information, terrain observation information, privileged observation information, and historical observation information.

[0035] The command observation information includes externally input control commands; these commands are used to control the target hexapod robot to switch motion modes; the motion modes include: robust base movement mode, stationary full-body operation mode, and dynamic walking operation mode. Body observation information includes: body quaternions, gravity vector, joint positions, joint velocities, and the action at the previous time step. Terrain observation information includes: terrain elevation sampling information within a set range. Privileged observation information includes: body mass, ground friction, foot-end force, foot-end torque, and the desired joint angle of the manipulating legs. Historical observation information is obtained by stitching together the body observation information corresponding to a set number of time steps.

[0036] In one embodiment, the method for determining the control strategy model specifically includes:

[0037] Based on the constructed multi-task training environment, the motion mode is switched according to command observation information; in any motion mode, the following is performed within the training architecture:

[0038] The ontology observation information is input into the state estimation encoder, and the output is the state estimation information; the terrain observation information is input into the terrain information encoder, and the output is the terrain feature vector; the privileged observation information is input into the privileged information encoder, and the output is the privileged feature vector; the terrain feature vector and the privileged feature vector are concatenated into the teacher feature vector; the historical observation information is input into the historical information encoder, and the output is the student feature vector; the command observation information, ontology observation information, state estimation information, and teacher feature vector are input into the ontology network for action selection, and the output is the corresponding expected joint angle.

[0039] The current output action is determined based on the corresponding expected joint angle, and the current reward value is calculated using a reward function. The action reward is then calculated based on the current reward value, and normalization parameters for the action reward are determined. These normalization parameters include the mean and standard deviation. The normalized action reward is determined based on the value network. The normalization parameters are dynamically updated using an incremental method. The weights and biases of each task head in the value network are adjusted based on the updated normalization parameters. Finally, the shared feature layer and the current task head of the value network are updated based on the normalized action reward.

[0040] Reinforcement learning is used to update the terrain information encoder, privileged information encoder, and ontology network.

[0041] Using the actual fuselage speed in the preset simulation as a label, and aiming to minimize the difference between the state estimation information and the actual fuselage speed, the state estimation encoder is updated using supervised learning. Also, using the teacher feature vector as a label, and aiming to minimize the difference between the student feature vector and the teacher feature vector, the historical information encoder is updated using supervised learning until the preset training convergence state is reached. The training architecture corresponding to the preset training convergence state is determined as the control strategy model.

[0042] The normalization parameters are updated dynamically using an incremental method, and the corresponding mathematical expression is:

[0043] .

[0044] in, For the task exist The mean parameter at any given time; The decay rate of the exponential moving average; For update operator; For the task exist The mean parameter at any given time; For the task exist The reward value in the current sample subset at the given time; To calculate the mean; For the task exist The second moment at time t; For the task exist The second moment at time t; To calculate the variance; For the task exist The variance parameter at time. For serial numbers.

[0045] Determining normalized action rewards based on value networks specifically includes:

[0046] For each task head in the multi-task network, corresponding to a task, it outputs a normalized value by receiving state features extracted from the shared feature layer:

[0047] .

[0048] Determining normalized action rewards using generalized advantage estimation:

[0049] .

[0050] in, For task headers targeting tasks The normalized value of the output; As weight; The state features are output by the shared feature layer; For bias; For the task exist TD residuals at time step; For the task exist The reward value at any given moment; Discount factor; For the task exist Value estimation at any given moment; For the task exist Value estimation at any given moment; For the task exist The advantage of timing estimation; For the smoothing parameter of the generalized dominance estimate; For the task exist TD residuals at time step; This represents the offset for future steps. In return. For sharing feature layers.

[0051] Adjusting the weights and biases of each task head in the value network based on the updated normalized parameters, specifically including:

[0052] .

[0053] in, As weight; For bias; The mean, Standard deviation; The adjusted weights; This is the adjusted bias; For the updated standard deviation; This is the updated mean; This is the update operator.

[0054] The reinforcement learning method employs the PPO algorithm, using CLIP loss for updates, specifically including:

[0055] .

[0056] .

[0057] in, The loss value updated for the controller; for Calculation of expected value at time; Importance sampling ratio; For network parameters; For normalized advantage estimation; This is the clipping function; The cropping area; For network parameters The control network; The strategy before the update; for Momentary action; for Current state; For the task exist Normalized advantage estimation at time; For the task exist Timing advantage estimation; The mean, The standard deviation is denoted as .

[0058] This application proposes a multi-modal reinforcement learning control method for hexapod robots. Based on multi-task reinforcement learning, it achieves integrated movement-manipulation control of the hexapod robot. Through a limb reuse architecture and a multi-head commentator network, it realizes seamless integration of movement, static manipulation, and dynamic in-journey manipulation. In practical applications, the technical solution adopted in this application includes the following steps:

[0059] Training process:

[0060] A multi-task training environment is constructed, the control command range is updated using a course-based approach, and a robust training mechanism is added. The robot's observation information is acquired, including: command observation information, body observation information, terrain observation information, privileged observation information, and historical observation information. Based on different command observation information, the hexapod robot will autonomously switch motion modes, including: robust base movement mode, stationary full-body operation mode, and dynamic walking operation mode. The control process is consistent for all motion modes.

[0061] The system inputs ontology observation information into the state estimation encoder, outputting state estimation information; it inputs terrain observation information into the terrain information encoder, outputting a terrain feature vector; it inputs privileged observation information into the privileged information encoder, outputting a privileged feature vector; it concatenates the terrain feature vector and the privileged feature vector into a teacher feature vector; it inputs historical observation information into the historical information encoder, outputting a student feature vector; it inputs command observation information, ontology observation information, state estimation information, and teacher feature vector into the ontology network, which selects actions based on the input and outputs the desired joint position; it inputs the desired joint position into the joint controller to obtain the robot's current output action; it uses a reward function to calculate the current reward value based on the current output action; and it calculates the reward value based on the current reward value. Calculate the action reward and its normalized parameters: mean and standard deviation; calculate the normalized action reward; dynamically update the mean and standard deviation using an incremental method; adjust the weights and biases of each task head in the value assessment network; update the shared feature layer and current task head of the value assessment network based on the normalized action reward; update the terrain information encoder, privileged information encoder, and ontology network using the PPO algorithm; using the actual aircraft speed in the simulation as a label, update the state estimation encoder using supervised learning with the goal of minimizing the difference between the state estimation information and the actual aircraft speed; using the teacher feature vector as a label, update the history information encoder using supervised learning with the goal of minimizing the difference between the student feature vector and the teacher feature vector; repeat this process until training converges.

[0062] Usage process:

[0063] The operator specifies control commands as command observation information; depending on the different command observation information, the hexapod robot will autonomously switch motion modes. When the command observation information only contains the base velocity command (x, y linear velocity, z angular velocity), robust base movement is implemented; when the command observation information only contains the foot-end manipulator leg specification command (a six-dimensional Boolean vector specifying any 0-2 legs as manipulator legs) and the manipulator leg position command (specifying the three-dimensional coordinates of the manipulator leg in the robot coordinate system), the robot implements stationary whole-body operation control; when the command observation information simultaneously contains the base velocity command, the foot-end manipulator leg specification command, and the manipulator leg position command, the robot implements dynamic movement operation. The robot acquires observation information, including command observation information, body observation information, and historical observation information. The body observation information is input to a state estimation encoder, which outputs state estimation information. Historical observation information is input to a historical information encoder, which outputs a student feature vector. The command observation information, body observation information, state estimation information, and student feature vector are input to a body network, which selects actions based on the input and outputs the desired joint position. The joint torque is calculated by a PD controller and sent to a joint controller, which outputs torque to control the movement of the hexapod robot.

[0064] The multi-mode hexapod robot reinforcement learning control method proposed in this application can move using a full-leg gait when not performing operational tasks; when performing operational tasks, it can achieve trajectory tracking of the end of any specified leg while the remaining moving legs naturally optimize different stable movement gaits; when the base is stationary, it can expand a larger workspace through an effect similar to whole-body control.

[0065] Step 1: Build a multi-task training environment.

[0066] Construct a multi-task modal training environment and randomly generate robot control commands for each task.

[0067] The first task mode is base locomotion control, which achieves robot body pose control by tracking root velocity commands.

[0068] The second task mode is stationary whole-body manipulation, which achieves fixed-point operation by tracking the spatial coordinates of each foot actuator; in this mode, it is necessary to specify the operating leg and the desired position of the operating leg.

[0069] The third task mode is dynamic locomotion-coupled manipulation, which synchronously executes root velocity command tracking and foot position tracking to achieve real-time coordination of motion and manipulation functions.

[0070] The basic motion control task is built on a mature methodology. The robot achieves dynamic tracking of the target linear velocity and angular velocity by generating rhythmic gait patterns. The velocity command range is adaptively scaled through a reward-driven mechanism, in which the allowable velocity limit automatically expands as the system accumulates velocity tracking rewards during training.

[0071] .

[0072] in, This is the command for maximum speed after the update. This is the command for maximum speed before the update. For speed bonus value, This represents the maximum obtainable speed bonus value.

[0073] Static operation task: While maintaining a stable support posture, the robot specifies the end effectors of the limbs to track the preset target position; all feasible areas within the leg workspace are determined through forward kinematics and collision detection algorithms, and 1-2 limb end effectors are randomly selected to track the trajectory of randomly sampled target points; in order to achieve a work domain expansion function similar to a whole-body controller and ensure motion accessibility, the desired base position / posture in the world coordinate system is progressively expanded during the training process, thereby expanding the reachable operation domain of the foot end effectors.

[0074] Dynamic in-journey operation task: Synchronously execute the base speed command and the joint trajectory tracking of the designated foot actuator; when in the speed tracking task, the base's degrees of freedom are constrained to some extent, and the operating foot is confined to its intrinsic working domain, which is defined in the mechanical motion reachable space in the robot's body coordinate system. A working domain extension method similar to a whole-body controller (WBC-like) is not used here.

[0075] Robust training mechanism: In the base movement training, an impulse in any direction is applied to the root of the robot to simulate the external disturbance environment; in the operation training phase, an omnidirectional force-torque combination is used to randomly load the foot actuators to simulate load interaction scenarios.

[0076] The training process was carried out simultaneously in three task environments (pure motion task, static operation task, and mobile operation task) with a scenario distribution ratio of 25%:25%:50%.

[0077] Step 2: Obtain observation information.

[0078] The observation space is divided into a four-dimensional subspace structure: command observation information, ontological observation information, topographic observation information, privileged observation information, and historical observation information.

[0079] During the control process of synchronously executing operations and movement tasks, command observation includes three types of information:

[0080] 1. Base movement command: desired fuselage linear velocity (x, y directions) and angular velocity (z direction).

[0081] 2. Designated operating leg: A six-dimensional Boolean vector (Foot id cmd) is used to binary encode the task mode of each limb. The state value "0" corresponds to the movement task, "1" corresponds to the operation task, and the all-zero vector represents the pure movement condition. The autonomous switching control of single-limb / multi-limb task mode is realized by dynamically updating this vector.

[0082] 3. Manipulating leg position command: Perform a Hadamard product operation on the Boolean encoded vector (Foot id cmd) and the three-dimensional expected position matrix (Foot pos cmd) of the manipulating limb end, so that the expected position parameter of the moving limb end is forced to zero, while retaining the spatial coordinate command of the manipulating limb end, thereby establishing a mapping relationship between task allocation and motion control.

[0083] The body observation information includes: fuselage quaternions, gravity vector, joint position, joint velocity, and the previous time step action, all of which are information that can be obtained from actual sensors.

[0084] Topographic observation information includes: topographic elevation sampling within a specific area. Used only during simulation training.

[0085] Privileged observation information includes: fuselage mass, ground friction, foot-end force, foot-end torque, and desired joint angle of the operating leg. This information is used only during simulation training.

[0086] The historical observation information is pieced together from nearly 20 time steps of ontological observation.

[0087] All spatial location information is converted into the robot coordinate system as observation input.

[0088] Step 3: Control Strategy.

[0089] like Figure 2 As shown, the observation information is input into the control strategy, the desired joint angle of the robot is output, and the joint torque is calculated by the PD controller, thereby controlling the movement of the hexapod robot. Figure 2 In This represents the teacher's feature vector.

[0090] The control strategy is based on a training framework centered on privileged learning for both teachers and students. Specifically:

[0091] The control strategy architecture comprises four encoding modules: a state estimation encoder, a terrain information encoder, a privileged information encoder, a history information encoder, and an ontology network.

[0092] The state estimation encoder is implemented as a multilayer perceptron, which takes the body observation information as input and outputs the robot's state estimation information, namely the estimated values ​​of linear velocity and angular velocity.

[0093] The terrain information encoder is implemented as a multilayer perceptron, which encodes terrain observation information into terrain feature vectors.

[0094] The privileged encoder module is implemented as a multilayer perceptron, which encodes privileged observation information into privileged feature vectors.

[0095] The terrain feature vector and the privileged feature vector are concatenated to form the corresponding teacher feature vector. .

[0096] The historical information encoder is implemented as a TCN network, which uses historical observation information as input and outputs the corresponding student feature vector. This vector is used to replace real-world deployments where full observations are lacking. Input ontology network.

[0097] The ontology network structure consists of a multilayer perceptron with three hidden layers, each containing [512, 256, 128 neurons], using ELU as the activation function. Command observation information, ontology observation information, state estimation information, and teacher feature vectors are input to the ontology network, and the output action space represents the desired joint positions of the 18 driven joints. The joint torques are calculated by the PD controller to control the hexapod robot's motion. The PD control calculation formula is:

[0098] .

[0099] in, Joint torque, unit: ; This is the proportionality coefficient; The desired joint position is expressed in rad. This is for position feedback, measured in rad. These are the differential coefficients; For velocity feedback, the unit is rad / s.

[0100] Step 4: Calculate the reward value.

[0101] Different robot motion scenarios can be evaluated by calculating reward values ​​using a reward function.

[0102] Different tasks have different reward function configurations.

[0103] Most safety-related penalties are universal, such as joint rotation speed and joint torque. The main difference between different tasks lies in the task rewards.

[0104] The walking task encourages the base to track specified linear and angular velocity commands. These are defined using exponential functions.

[0105] During the stationary operation, the specified operating foot is required to track the desired position in space, while penalizing movement of the standing foot and reducing the penalty weight for base stability. A simple tracking reward function is:

[0106] .

[0107] in, For foot tracking rewards, For foot position command, The current position of the foot. This is the position tracking deviation coefficient.

[0108] However, this simple reward design can lead to unsafe leg swinging when the positional error is large. Therefore, the positional tracking reward is divided into two parts:

[0109] .

[0110] in and Separate at the current time (i.e.) (time) and the previous time (i.e.) The foot position error at any given moment is calculated. The desired foot speed is set to 0.1 m / s. The time taken to travel between the actual foot position and the desired position is calculated. If the time is less than the travel time, the foot is guaranteed to move towards the desired position. If the time is greater than the travel time, accurate tracking is encouraged.

[0111] .

[0112] For estimated arrival time, This is the initial foot position. To predict foot speed.

[0113] In mobile operation tasks, the agent must track two tasks simultaneously.

[0114] The specific reward function and weights are shown in Table 1.

[0115] Table 1 Reward Function and Weight Table

[0116]

[0117] in, For linear velocity command, The linear velocity at the current moment. This is the linear velocity tracking deviation coefficient. For angular velocity command, The angular velocity at the current moment, This is the angular velocity tracking deviation coefficient. for Error in foot position at any given moment. for Error in foot position at any given moment. For foot position command, The current position of the foot. This is the position tracking deviation coefficient. Command for the position of the foot landing. For the position of the foot landing, Gravity vector Axial projection, Gravity vector Axial projection, For fuselage height, For the high expectations of the fuselage, For the first The force of contact between the legs and the ground. For the first The maximum permissible ground contact force of a leg. Let the velocity of the fuselage be along the z-axis. For the first Joint torque, For the first Joint rotation speed, For the previous moment Joint rotation speed, For time step. For serial numbers. The number of feet that touch the ground.

[0118] Step 5: Strengthen learning and training.

[0119] In the leg-arm reuse problem of hexapod robots, there is a conflict between the stable movement of the base and the precise tracking of the feet. The movement task relies on periodic gait planning to maintain base stability, while the manipulation task requires temporarily converting the moving leg into a manipulator arm, leading to a contraction of the support area and increasing the risk of base instability. The reward configurations also differ between tasks: the movement task focuses on base velocity tracking, while the manipulation task emphasizes foot position tracking. The policy gradient may get stuck in a local optimum due to the conflict between task objectives, and the scale difference of rewards for different tasks will further exacerbate the instability of training.

[0120] Given this contradiction, naively using reinforcement learning to handle multi-task scenarios often results in poor performance. Agents are typically distracted by a single task, focusing instead on the more prominent one. Normalizing the value estimation makes the value function learning more stable, while effective denormalization is used for policy updates. Combined with a multi-task head critic method, this approach enables leg-arm reuse control of a hexapod robot, effectively mitigating gradient conflicts and value scale differences between the robot's movement and manipulation tasks.

[0121] In the method mentioned in this application, the policy network is task-agnostic, that is, it is general to all tasks. The policy network does not explicitly accept task indices, but infers them implicitly from the commands. This ensures that the generated agent can operate in a completely task-agnostic manner.

[0122] The value network is designed with a shared feature layer and multiple task heads. The backbone network is used to extract state features. Each task head corresponds to a task, accepts state features, and outputs a normalized value.

[0123] .

[0124] Collect experience from multiple tasks in parallel using the current strategy, and add an index to each task. Stored as:

[0125] .

[0126] in, As a reward; For action; State; This is the updated state.

[0127] For the task For each sample, the return is calculated using generalized advantage estimation (GAE):

[0128] .

[0129] The raw value for each task is parameterized into a linear transformation of the normalized value. :

[0130] .

[0131] For each task, updates are performed using mini-batch data, employing an incremental approach to track the new mean and standard deviation of the values:

[0132] .

[0133] In this process, the standardization objective of the values ​​is non-stationary, which can corrupt the learned network, thus affecting the updating of the mean. Standard value At this time, it is necessary to adjust the weight of the last layer of the critic. and This ensures the continuity of the output, thereby avoiding drastic fluctuations in the critic output due to changes in the normalization parameter.

[0134] .

[0135] The value network updates the shared backbone parameters θ and the corresponding weights and biases of the task headers:

[0136] .

[0137] in, For value network sharing feature layer and the first The loss value of each task head; For the task The return value. For serial numbers.

[0138] In the control policy network, the terrain information encoder, privileged information encoder, and ontology network are updated using the reinforcement learning proximal policy optimization algorithm (PPO):

[0139] Advantages after normalization Update the strategy:

[0140] .

[0141] PPO uses CLIP loss for policy updates:

[0142] .

[0143] Each time a rollout occurs, the corresponding value network... The task head is updated and reverted to the normalized target value, while the policy network updates using the denormalized result, independent of the task.

[0144] The state estimation encoder and the history information encoder are updated using supervised learning, and are trained concurrently with reinforcement learning.

[0145] Using the actual fuselage speed in the simulation as a label, and with the goal of minimizing the difference between the state estimation information and the actual fuselage speed, supervised learning is used to update the state estimation encoder.

[0146] Using teacher feature vectors as labels, and aiming to minimize the difference between student and teacher feature vectors, supervised learning is used to update the historical information encoder.

[0147] The optimized state estimation encoder, terrain information encoder, privileged information encoder, history information encoder, and ontology network are obtained through training.

[0148] Step Six: Deploy the control strategy on the actual machine.

[0149] like Figure 3 As shown, the operator specifies control commands as command observation information; depending on the different command observation information, the hexapod robot will autonomously switch motion modes, acquire the robot's observation information, input the body observation information to the state estimation encoder, and output state estimation information.

[0150] Historical observation information is input into the historical information encoder, which outputs the student feature vector. Command observation information, ontology observation information, state estimation information, and student feature vector are input into the ontology network. The ontology network selects actions based on the input and outputs the expected joint position.

[0151] The PD controller calculates the joint torque and sends it to the joint controller, which then outputs torque to control the movement of the hexapod robot. Figure 3 In To command observation information, For ontological observation information, For state estimation information, This represents the student's feature vector.

[0152] In one exemplary embodiment, such as Figure 4 As shown, a reinforcement learning-based hexapod robot leg-arm multiplexing control device is provided, comprising:

[0153] The information acquisition module is used to acquire observation information of the target hexapod robot.

[0154] The desired joint angle determination module is used to input observation information into the control strategy model to obtain the desired joint angles of the target hexapod robot. The control strategy model is obtained by updating the training architecture based on the value network using reinforcement learning, based on the constructed multi-task training environment. The training architecture is determined based on teacher-student privileged learning. The value network includes a shared feature layer and a multi-task head connected in sequence. The training architecture includes: a state estimation encoder, a terrain information encoder, a privileged information encoder, a history information encoder, and an ontology network.

[0155] The leg-arm reuse control module is used to determine the joint torque based on the desired joint angle using the PD control method, so as to control the leg-arm reuse of the target hexapod robot.

[0156] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0157] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A reinforcement learning-based control method for the reuse of legs and arms in a hexapod robot, characterized in that, include: Obtain observational information about the target hexapod robot; The observation information is input into the control strategy model to obtain the desired joint angles of the target hexapod robot; The control strategy model is obtained by updating the training architecture based on the value network using reinforcement learning, based on the constructed multi-task training environment. The training architecture is determined based on teacher-student privileged learning; the value network includes a shared feature layer and a multi-task head connected in sequence; the training architecture includes: a state estimation encoder, a terrain information encoder, a privileged information encoder, a history information encoder, and an ontology network; The PD control method is used to determine the joint torque based on the desired joint angle in order to control the reuse of the legs and arms of the target hexapod robot. The multi-task training environment is defined using a curriculum-based approach with the addition of robust training mechanisms; the multi-task training environment includes basic motion control task modes, static whole-body operation task modes, and dynamic motion-coupled operation task modes. Among them, the basic motion control task mode achieves dynamic tracking of target linear velocity and angular velocity by generating rhythmic gait patterns; The static whole-body operation task mode is to maintain a stable support posture while specifying the end effector of the limb to track the preset target position. The algorithm uses forward kinematics and collision detection to determine all feasible areas in the leg workspace and randomly selects a set number of end effectors to track the trajectory of randomly sampled target points. The dynamic motion coupling operation task mode is the joint trajectory tracking of synchronously executing the base speed command and the specified foot actuator; The robust training mechanism simulates external disturbances by applying an impulse in any direction to the root of the target hexapod robot during base movement training; during operation training, a combination of omnidirectional force and torque is used to randomly load the foot actuators to simulate load interaction scenarios.

2. The reinforcement learning-based hexapod robot leg and arm reuse control method according to claim 1, characterized in that, The observation information specifically includes: command observation information, body observation information, terrain observation information, privileged observation information, and historical observation information; The command observation information includes externally input control commands; the control commands are used to control the target hexapod robot to switch motion modes; the motion modes include: robust base movement mode, stationary whole-body operation mode, and dynamic moving operation mode; The body observation information includes: fuselage quaternions, gravity vector, joint position, joint velocity, and previous time step action; The terrain observation information includes: terrain elevation sampling information within a defined range; The privileged observation information includes: fuselage mass, ground friction, foot-end force, foot-end torque, and desired joint angle of the operating leg; The historical observation information is obtained by splicing together the ontological observation information corresponding to a set number of time steps.

3. The reinforcement learning-based hexapod robot leg and arm reuse control method according to claim 2, characterized in that, The method for determining the control strategy model specifically includes: Based on the constructed multi-task training environment, the motion mode is switched according to command observation information; In any motion mode, the following is performed within the training architecture: The ontology observation information is input into the state estimation encoder, and the output is the state estimation information; The terrain observation information is input into the terrain information encoder, and the output is a terrain feature vector; The privileged observation information is input into the privileged information encoder, and the privileged feature vector is output. The terrain feature vector and the privileged feature vector are concatenated to form the teacher feature vector; The historical observation information is input into the historical information encoder, and the student feature vector is output. The command observation information, ontology observation information, state estimation information and teacher feature vector are input into the ontology network to select actions and output the corresponding expected joint angles. The current output action is determined based on the corresponding expected joint angle, and the current reward value is calculated based on the current output action using a reward function; The action reward is calculated based on the current reward value, and the normalization parameters of the action reward are determined; the normalization parameters include: mean and standard deviation; Determine normalized action rewards based on value networks; The normalized parameters are updated dynamically using an incremental method. Adjust the weights and biases of each task head in the value network according to the updated normalization parameters; Update the shared feature layer and current task head of the value network based on the normalized action reward; Reinforcement learning is used to update the terrain information encoder, privileged information encoder, and ontology network; Using the actual fuselage speed in the preset simulation as a label, and aiming to minimize the difference between the state estimation information and the actual fuselage speed, the state estimation encoder is updated using supervised learning. Also, using the teacher feature vector as a label, and aiming to minimize the difference between the student feature vector and the teacher feature vector, the historical information encoder is updated using supervised learning until the preset training convergence state is reached. The training architecture corresponding to the preset training convergence state is determined as the control strategy model.

4. The reinforcement learning-based hexapod robot leg and arm reuse control method according to claim 3, characterized in that, The normalization parameters are updated dynamically using an incremental method, and the corresponding mathematical expression is: ; in, For the task exist The mean parameter at any given time; The decay rate of the exponential moving average; For update operator; For the task exist The mean parameter at any given time; For the task exist The reward value in the current sample subset at the given time; To calculate the mean; For the task exist The second moment at time t; For the task exist The second moment at time t; To calculate the variance; For the task exist The variance parameter at time; For serial numbers.

5. The reinforcement learning-based hexapod robot leg and arm reuse control method according to claim 3, characterized in that, Determining normalized action rewards based on value networks specifically includes: For each task head in the multi-task network, corresponding to a task, it outputs a normalized value by receiving state features extracted from the shared feature layer: ; Determining normalized action rewards using generalized advantage estimation: ; in, For task headers targeting tasks The normalized value of the output; As weight; The state features are output by the shared feature layer; For bias; For the task exist TD residuals at time step; For the task exist The reward value at any given moment; Discount factor; For the task exist Value estimation at any given moment; For the task exist Value estimation at any given moment; For the task exist The advantage of timing estimation; For the smoothing parameter of the generalized dominance estimate; For the task exist TD residuals at time step; This represents the offset for future steps. In return; For serial numbers.

6. The reinforcement learning-based hexapod robot leg and arm reuse control method according to claim 3, characterized in that, Adjusting the weights and biases of each task head in the value network based on the updated normalized parameters, specifically including: ; in, As weight; For bias; The mean, Standard deviation; The adjusted weights; This is the adjusted bias; For the updated standard deviation; This is the updated mean; This is the update operator.

7. The reinforcement learning-based hexapod robot leg and arm reuse control method according to claim 3, characterized in that, The reinforcement learning method employs the PPO algorithm, using CLIP loss for updates, specifically including: ; ; in, The loss value updated for the controller; for Calculation of expected value at time; Importance sampling ratio; For network parameters; For normalized advantage estimation; This is the clipping function; The cropping area; For network parameters The control network; The strategy before the update; for Momentary action; for Current state; For the task exist Normalized advantage estimation at time; For the task exist Timing advantage estimation; The mean, Standard deviation; For serial numbers.

8. The reinforcement learning-based hexapod robot leg and arm reuse control method according to claim 1, characterized in that, The mathematical expression corresponding to the PD control method is: ; in, Joint torque; This is the proportionality coefficient; For the desired joint position; For location feedback; These are the differential coefficients; For speed feedback.

9. A reinforcement learning-based hexapod robot leg-arm reuse control device, characterized in that, include: The information acquisition module is used to acquire observation information of the target hexapod robot; The expected joint angle determination module is used to input the observation information into the control strategy model to obtain the expected joint angles of the target hexapod robot; the control strategy model is obtained by updating the training architecture based on the value network using reinforcement learning, according to the constructed multi-task training environment. The training architecture is determined based on teacher-student privileged learning; the value network includes a shared feature layer and a multi-task head connected in sequence; the training architecture includes: a state estimation encoder, a terrain information encoder, a privileged information encoder, a history information encoder, and an ontology network; The leg-arm reuse control module is used to determine the joint torque based on the desired joint angle using the PD control method, so as to control the leg-arm reuse of the target hexapod robot. The multi-task training environment is defined using a curriculum-based approach with the addition of robust training mechanisms; the multi-task training environment includes basic motion control task modes, static whole-body operation task modes, and dynamic motion-coupled operation task modes. Among them, the basic motion control task mode achieves dynamic tracking of target linear velocity and angular velocity by generating rhythmic gait patterns; The static whole-body operation task mode is to maintain a stable support posture while specifying the end effector of the limb to track the preset target position. The algorithm uses forward kinematics and collision detection to determine all feasible areas in the leg workspace and randomly selects a set number of end effectors to track the trajectory of randomly sampled target points. The dynamic motion coupling operation task mode is the joint trajectory tracking of synchronously executing the base speed command and the specified foot actuator; The robust training mechanism simulates external disturbances by applying an impulse in any direction to the root of the target hexapod robot during base movement training; during operation training, a combination of omnidirectional force and torque is used to randomly load the foot actuators to simulate load interaction scenarios.

Citation Information

Patent Citations

  • Igniter robot production line control method and system based on deep reinforcement learning

    CN119579120A