Gait learning method, system, device and storage medium based on reinforcement learning

By collecting and clustering the historical execution status of quadruped robots and enriching the initial state library, the problem of insufficient adaptability of quadruped robots' gait control strategy is solved, and more efficient and stable gait learning and control are achieved.

CN119474884BActive Publication Date: 2025-05-23SHANDONG ENERGY GRP CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510052267.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-23
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

The influence of four-legged robots on the same action in different initial states is different, resulting in poor adaptability of strategies derived from deep reinforcement learning.

Method used

By collecting various states generated during the execution of historical expert strategies as initialization states, a pre-constructed deep reinforcement learning model is used to perform gait learning to obtain a gait control strategy. Specific steps include collecting and compiling state data, saving it to the initial state library, and improving the diversity of the initial state through clustering and update processing.

Benefits of technology

The adaptability of the gait control strategy obtained by the deep reinforcement learning model is improved, the learning efficiency and effect are improved, and the stable movement of the four-legged robot under different environments and conditions is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119474884B_ABST
    Figure CN119474884B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of quadruped robots, and specifically provides a gait learning method, system, device and storage medium based on reinforcement learning, including: collecting multiple states generated during the execution of historical expert strategies as initialization states; using a pre-built deep reinforcement learning model, performing gait learning based on the initialization state to obtain a gait control strategy. The present invention enriches the initial state of the deep reinforcement learning model and improves the adaptability of the learned gait control strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of quadruped robots, and in particular relates to a gait learning method, system, device and storage medium based on reinforcement learning. Background Art

[0002] Flexible and efficient motion control is the basis and prerequisite for the realization of specific functions of various mobile robots. To this end, scholars in the field of robotics continue to explore and optimize robot motion control algorithms, and are committed to achieving reliable, accurate and efficient control of complex robots. Compared with wheeled or tracked robots, legged robots represented by quadruped bionic robots have inherent characteristics such as complex mechanical structures, and their motion stability and environmental adaptability need to be improved.

[0003] In recent years, the gait control algorithm of quadrupedal bionic robots based on deep reinforcement learning has gradually emerged, that is, the quadrupedal robot learns the appropriate gait control strategy through continuous trial and error. At present, most reinforcement learning tasks use a fixed initial state.

[0004] However, for quadruped robots, different initial states have different effects on the same action. Therefore, using a fixed initial state makes the strategy derived from deep reinforcement learning less adaptable. Summary of the invention

[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides a gait learning method, system, device and storage medium based on reinforcement learning to solve the above-mentioned technical problems.

[0006] In a first aspect, the present invention provides a gait learning method based on reinforcement learning, comprising:

[0007] Collect various states generated during the execution of historical expert strategies as initialization states;

[0008] Using a pre-built deep reinforcement learning model, gait learning is performed based on the initialization state to obtain a gait control strategy.

[0009] In an optional implementation, multiple states generated during the execution of historical expert strategies are collected as initialization states, including:

[0010] Collecting states generated during the execution of historical expert strategies, wherein the states include positions, postures, angles, and angular velocities of each link of the quadruped robot;

[0011] The state is compiled into an initial state tuple, and the initial state tuple is saved to an initial state repository.

[0012] In an optional embodiment, the method further comprises:

[0013] Defining a sample set , neighborhood parameters ;

[0014] Initialize the core object collection , initialize the number of clusters , initialize the unvisited sample set ,Cluster Partition ;

[0015] Find samples by distance measurement of -Neighborhood subsample set ;

[0016] Confirm that the number of samples in the subsample set meets , the sample Add core object sample collection ;

[0017] In the core object collection In the example above, randomly select a core object , initialize the current cluster core object queue , Initialize the category number , initialize the current cluster sample set , Update the unvisited sample set ;

[0018] In the current cluster core object queue Extract a core object , through the neighborhood distance threshold Find all -Neighborhood subsample set ,make , Update the current cluster sample set , Update the unvisited sample set ,renew ;in, is the neighborhood subsample set The intersection with the set of unvisited samples;

[0019] Confirm the current cluster core object queue , then the current cluster Generation completed, update cluster division , Update the core object collection , and confirm the core object set , the processing is determined to be complete, the cluster division results are output, and samples are evenly selected from multiple clusters as the initial state of each learning cycle of the deep reinforcement learning model.

[0020] In an optional embodiment, the sample is found by distance measurement. of -Neighborhood subsample set ,include:

[0021] For sample The state introduces momentum factor ( , , , ), and the momentum state is obtained ( , , , ),in, For location, For posture, For speed, is the angular velocity, For quality, is the moment of inertia;

[0022] Calculation Sample With sample The Euclidean distance of the momentum state, if the Euclidean distance does not exceed , the sample is classified into the -Neighborhood subsample set .

[0023] In an optional embodiment, the method further comprises:

[0024] Calculate the potential difference between any two samples in the subsample set and calculate the average value of the potential difference;

[0025] For isolated samples that are not included in the neighborhood subsample set, the potential difference between the isolated sample and the core sample of each subsample set is calculated, and the target subsample set whose potential difference does not exceed the corresponding potential difference average value is selected;

[0026] If there is a target sub-sample set, the isolated sample is classified into the target sub-sample set; if there is no target sub-sample set, the isolated sample is determined to be noise.

[0027] In an optional embodiment, the method further comprises:

[0028] The state data generated by the gait control strategy during the test are collected, and the state data are marked with the parameters of the quadruped robot and then imported into the initial state library.

[0029] In a second aspect, the present invention provides a gait learning system based on reinforcement learning, comprising:

[0030] The data collection module is used to collect various states generated during the execution of historical expert strategies as initialization states;

[0031] The gait learning module is used to utilize a pre-built deep reinforcement learning model to perform gait learning based on the initialization state to obtain a gait control strategy.

[0032] In an optional embodiment, the data collection module includes:

[0033] A collecting unit, used for collecting states generated during the execution of historical expert strategies, wherein the states include positions, postures, angles and angular velocities of each link of the quadruped robot;

[0034] The storage unit is used to compile the state into an initial state tuple and save the initial state tuple to an initial state library.

[0035] In a third aspect, a device is provided, including:

[0036] A memory, used for storing a gait learning program based on reinforcement learning;

[0037] A processor is used to implement the steps of the reinforcement learning-based gait learning method provided in the first aspect when executing the reinforcement learning-based gait learning program.

[0038] In a fourth aspect, a computer-readable storage medium is provided, on which a reinforcement learning-based gait learning program is stored. When the reinforcement learning-based gait learning program is executed by a processor, the steps of the reinforcement learning-based gait learning method provided in the first aspect are implemented.

[0039] The beneficial effect of the present invention lies in that the reinforcement learning-based gait learning method, system, device and storage medium provided by the present invention enrich the initial state of the deep reinforcement learning model by collecting multiple states generated by historical expert strategies as initialization states, thereby improving the adaptability of the learned gait control strategy.

[0040] In addition, by clustering and updating the initialization states, it is ensured that the deep reinforcement learning model can learn enough different initial states, thereby improving learning efficiency and effectiveness.

[0041] In addition, the invention has a reliable design principle, a simple structure and a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention.

[0044] Figure 2 is a schematic block diagram of a system according to an embodiment of the present invention.

[0045] Figure 3 A schematic diagram of the structure of a device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0046] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0048] The key terms appearing in the present invention are explained below.

[0049] The Deep Reinforcement Learning (DRL) model is a machine learning model that combines deep learning and reinforcement learning. It uses the powerful feature extraction and expression capabilities of deep learning to solve the problem of processing high-dimensional state and action space in reinforcement learning.

[0050] Deep reinforcement learning model based on policy gradient

[0051] Asynchronous Advantage Actor-Critic (A3C) Algorithm

[0052] Principle: It combines the policy gradient and value function approximation methods, and accelerates convergence by having multiple parallel agents learn simultaneously in different copies of the environment. The A3C algorithm uses a global shared network, including a policy network (actor) and a value network (critic). Each parallel agent obtains parameters from the global network and feeds its own gradient information back to the global network for update.

[0053] Algorithm flow: Initialize the global policy network and value network; multiple agents run in parallel in different copies of the environment, each agent selects actions based on the current strategy, interacts with the environment to obtain rewards and new states; the agent calculates the advantage function, that is, the difference between the actual reward and the value network prediction value; according to the advantage function and the gradient of the policy network, the global policy network is updated; at the same time, according to the advantage function and the gradient of the value network, the global value network is updated.

[0054] Application scenarios: It is widely used in multi-agent systems, robot control in complex environments and other fields.

[0055] Proximal Policy Optimization (PPO) algorithm

[0056] Principle: The PPO algorithm is an improvement on the policy gradient algorithm, which optimizes policy updates by introducing importance sampling. It proposes a truncated policy optimization objective function to ensure that the policy does not change too much at each update, thereby improving the stability and efficiency of training.

[0057] Algorithm flow: Collect multi-step trajectory data; calculate the advantage function; use the truncated policy optimization objective function to update the policy network; optionally update the value network at the same time to better estimate the state value.

[0058] Application scenarios: It has achieved good results in the fields of robot control, autonomous driving, games, etc. For example, in complex tasks such as simulating robot walking and jumping, PPO can quickly learn effective strategies.

[0059] The deep reinforcement learning method uses the process of interaction between the intelligent agent and the environment to obtain experience and optimize the network. However, its disadvantage is that if a state that has not been encountered in the training process is encountered during the test, its output action is difficult to predict, and this process is often accompanied by instruction mutation.

[0060] The gait learning method based on reinforcement learning provided in the embodiment of the present invention is executed by a computer device, and accordingly, the gait learning system based on reinforcement learning runs in the computer device.

[0061] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention. Figure 1 The execution subject may be a gait learning system based on reinforcement learning. According to different requirements, the order of the steps in the flowchart may be changed, and some may be omitted.

[0062] like Figure 1 As shown, the method includes:

[0063] S1, collects various states generated during the execution of historical expert strategies as initialization states.

[0064] Determine the scope of data collection:

[0065] Clarify the definition and source of historical expert strategies. These strategies can be manually designed by domain experts based on experience, or they can be mature strategies obtained through extensive previous training.

[0066] Determine the type of environment in which data is collected, including different terrains such as flat ground, slopes, stairs, and rough roads, as well as environments with different interference factors, such as different frictions, external impacts, etc.

[0067] Consider the state data collection of a robot when performing different tasks, such as walking, running, turning, jumping, carrying objects, etc.

[0068] Select the state feature:

[0069] Body posture information: including the robot's roll angle, pitch angle and yaw angle, as well as the rate of change of these angles. This information can reflect the robot's overall posture and stability.

[0070] Joint angle and angular velocity: The angle and angular velocity of each joint of a quadruped robot, such as the hip joint and knee joint, are key indicators for describing the motion state of its legs.

[0071] Kinematic parameters: The robot’s linear velocity, angular velocity, and the position and velocity information of each foot end. These parameters can help understand the robot’s overall motion state.

[0072] Sensor data: Collect data from various sensors on the robot, such as accelerometers, gyroscopes, force sensors, etc., which can provide detailed information about the interaction between the robot and the environment.

[0073] Data collection and preprocessing:

[0074] Use appropriate recording devices and software tools to record various status data of the robot in real time during the execution of historical expert strategies.

[0075] The collected data is preprocessed, including data cleaning, normalization, filtering and other operations, to improve the quality and usability of the data.

[0076] The processed data is stored in a certain format and structure for subsequent reading and use.

[0077] Initialization state selection:

[0078] According to the specific learning tasks and requirements, select the appropriate state from the collected historical data as the initialization state. Random sampling, stratified sampling and other methods can be used to ensure the diversity and representativeness of the initialization state.

[0079] The selected initialization states are further screened and verified to ensure that these states can provide valuable information for the deep reinforcement learning model, helping the model to converge faster and learn effective strategies.

[0080] S2, using a pre-built deep reinforcement learning model, performs gait learning based on the initialization state to obtain a gait control strategy.

[0081] Gait learning is performed based on the initial state of step S1 using a deep Q network (DQN) or proximal policy optimization (PPO).

[0082] In an embodiment of the present invention, based on step S1, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.

[0083] S101. Collect the states generated during the execution of the historical expert strategy, wherein the states include the position, posture, angle and angular velocity of each connecting rod of the quadruped robot.

[0084] The states generated during the execution of other expert strategies are used as the initialization states. When other expert strategies are trained and tested, their states are collected and the positions, postures, speeds, and angular velocities of the robot's links are compiled into the robot's initial state tuple ( , , , ), the collection of all state elements is the expert strategy initial state library.

[0085] In a specific example, the data collection process is as follows:

[0086] Before starting to collect the status of other expert strategies during execution, a series of prerequisites and environment settings need to be clarified. First, ensure that the expert strategy training process is carried out in a relatively stable and representative environment. This environment should cover various typical scenarios that the robot may encounter, such as different terrains (such as flat ground, slopes, stairs, and rugged roads), different lighting conditions, and different interference factors (such as wind, vibration, etc.). At the same time, set the training goals of the expert strategy. These goals can be for the robot to complete specific tasks, such as walking a certain distance, crossing complex terrain, and completing a specific combination of actions.

[0087] When other expert strategies complete training and enter the testing phase, it is a critical period for state collection. During the test process, it is necessary to determine the appropriate state collection frequency. Too high a collection frequency may lead to data redundancy and increase the burden of data processing and storage; while too low a collection frequency may miss important state information. Generally speaking, the collection frequency can be determined based on the robot's motion characteristics and the complexity of the task. For example, for a relatively stable walking task, the state can be collected at a certain time interval (such as 0.1 seconds); for some rapidly changing actions (such as jumping, avoiding obstacles, etc.), more frequent collection is required, such as once every 0.01 seconds.

[0088] Each time the status is collected, the status information of each link of the robot needs to be collected comprehensively and carefully. Specifically, it includes the following aspects:

[0089] Position information of each link: The coordinate position of each link of the robot in three-dimensional space is obtained through accurate position sensors (such as encoders, laser rangefinders, etc.). This position information can accurately describe the distribution of the robot's limbs in space, which is crucial for understanding the overall shape and motion trajectory of the robot. For example, for each leg of a quadruped robot, the position coordinates of the hip joint, knee joint and other joints need to be recorded separately to determine the extension and contraction state of the leg.

[0090] Attitude information of each link: Use inertial measurement units (IMUs) and other devices to measure the attitude of each link, including roll angle, pitch angle, and yaw angle. These angle information reflects the rotation state of the link relative to the reference coordinate system, which can help understand the attitude change and stability of the robot during movement. For example, when the robot is walking, by monitoring the attitude changes of each link, it is possible to detect whether the robot is tilted or unbalanced in time.

[0091] Speed ​​information of each link: Based on the collected position information, the linear speed of each link is obtained by calculating the change in position at adjacent moments. The linear speed reflects how fast the position of the link changes in space, which is important for evaluating the robot's motion efficiency and coordination. For example, when the robot runs fast, the speed of each link needs to maintain a certain coordination to ensure the robot's stable progress.

[0092] Angular velocity information of each link: Combined with the posture information and time interval, the angular velocity of each link is calculated. Angular velocity describes the speed and direction of the link rotation, which is critical for analyzing the robot's steering, posture adjustment and other actions. For example, when the robot turns, the angular velocity of different links needs to be reasonably adjusted to achieve a smooth steering action.

[0093] The collected information on the position, posture, speed, and angular velocity of each link of the robot is sorted and compiled to form the robot's initial state tuple.

[0094] S102. Compile the state into an initial state tuple, and save the initial state tuple to an initial state library.

[0095] As the testing process continues, a large number of state tuples collected according to the above method continue to accumulate. All these state tuples are aggregated and stored to form a set, which is the expert strategy initial state library. In order to facilitate the management and use of this state library, a suitable data structure and storage method need to be adopted. A database (such as the relational database MySQL or the non-relational database MongoDB) can be used to store these state tuples, and the data can be classified and labeled, for example, according to different test environments, task types, robot actions, etc., so that these initialization states can be quickly and accurately retrieved and used in the subsequent deep reinforcement learning model training process.

[0096] S103. Cluster the initial state tuples in the initial state library.

[0097] The states in different scenes and tasks are often quite different, but the states at different times in the same task are often very different. For example, in a flip, the states in the take-off phase and the rolling phase are completely different. On the other hand, similar states can exist in different tasks. For example, the support phase of a jump is very similar to the support phase of normal walking, but the tasks are completely different. The above problems lead to extremely unbalanced samples in the expert strategy initial state library and it is difficult to determine the number of sample classes. If random sampling is used, it is difficult for the expert strategy to truly learn enough different initial states.

[0098] To solve this problem, the initial states in the initial state library are clustered, and representative states are evenly selected from each category after clustering as the initial states of the learning cycle. Specifically, the clustering method includes:

[0099] Predefined sample sets , neighborhood parameters .

[0100] 1) Initialize the core object collection , initialize the number of clusters , initialize the unvisited sample set ,Cluster Partition ;

[0101] 2) For , follow the steps below to find all core objects:

[0102] a) Find samples by distance measurement of -Neighborhood subsample set ;

[0103] b) If the number of samples in the subsample set satisfies , the sample Added core object sample collection: ;

[0104] 3) If the core object collection , the algorithm ends, otherwise it goes to step 4).

[0105] 4) In the core object collection In the example above, randomly select a core object , initialize the current cluster core object queue , Initialize the category number , initialize the current cluster sample set , Update the unvisited sample set ;

[0106] 5) If the current cluster core object queue , then the current cluster Generation completed, update cluster division , Update the core object collection , go to step 3);

[0107] 6) In the current cluster core object queue Extract a core object , through the neighborhood distance threshold Find all -Neighborhood subsample set ,make , Update the current cluster sample set , Update the unvisited sample set ,in, is the intersection of the neighborhood subsample set and the unvisited sample set; update , go to step 5);

[0108] The output result is: cluster division .

[0109] In order to further improve the rationality of clustering, considering the different importance of each link of the robot, simply using Euclidean distance to measure the robot's initial state ancestor cannot reasonably represent the differences between different states. Therefore, when calculating the distance between different states, momentum is introduced instead of speed to measure the distance between different states. That is, for state ( , , , ), when calculating the distance between this state and other states, use ( , , , ) as the calculation basis, where , represent the mass and moment of inertia of the corresponding connecting rod respectively.

[0110] For example, for the sample The state introduces momentum factor ( , , , ), and the momentum state is obtained ( , , , ),in, For location, For posture, For speed, is the angular velocity, For quality, is the moment of inertia; calculate the sample With sample The Euclidean distance of the momentum state, if the Euclidean distance does not exceed , the sample is classified into the -Neighborhood subsample set .

[0111] In addition, in order to remove the adverse effects of noise data on clustering results, a noise judgment method is further provided, comprising the following steps:

[0112] (1) In a given sample data set, each subsample set has been determined by distance metrics and neighborhood parameters. Each subsample set consists of a group of samples that are relatively close in space.

[0113] (2) For each sub-sample set, the potential difference between any two samples needs to be calculated. The calculation of the potential difference is based on the potential value of the samples.

[0114] The potential value calculation formula is:

[0115]

[0116] in, is the influence factor between samples, which can be 1; n is the number of samples in the subsample set.

[0117] Calculate the potential values ​​of two samples and then take the difference to get the potential difference between the two samples.

[0118] After calculating the potential differences between all pairs of samples in the subsample set, we perform statistical analysis on these potential differences and calculate their average values.

[0119] In the entire sample data set, there are some samples that are not included in any neighborhood subsample set. These samples are called isolated samples. Then the potential difference between the isolated sample and the core sample of each subsample set is calculated according to the above method. After calculating the potential difference between the isolated sample and the core sample of each subsample set, these potential differences are compared with the average potential difference of the corresponding subsample set. For each subsample set, if the potential difference between the isolated sample and the core sample of the subsample set does not exceed the average potential difference of the subsample set, then this subsample set is selected as the target subsample set.

[0120] If there is at least one target sub-sample set, it means that the isolated sample has a certain similarity with these sub-sample sets in potential difference, so the isolated sample can be classified into these target sub-sample sets. The purpose of this is to classify samples that are relatively close in the feature space into one category for better data analysis and model training.

[0121] If after comparison, there is no target sub-sample set, that is, the potential difference between the isolated sample and the core samples of all sub-sample sets exceeds the average potential difference of the corresponding sub-sample sets, then we determine that the isolated sample is noise. Noise samples usually have large differences in characteristics from other samples, which may be caused by measurement errors, abnormal data points, etc. Judging them as noise helps improve the accuracy of data analysis and the stability of the model.

[0122] S104. Update the initial state tuple in the initial state library.

[0123] The state data generated by the gait control strategy during the test are collected, and the state data are marked with the parameters of the quadruped robot and then imported into the initial state library.

[0124] In an embodiment of the present invention, based on step S2, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.

[0125] S201.Data preparation.

[0126] (1) Determine the number of selections: Determine the number of representative states to be selected from each category based on the amount of data and training requirements for each cluster category. You can use a fixed number approach, that is, select the same number of states from each category; or you can determine the number of selections based on the proportion of the category data volume, selecting more states for categories with larger data volumes to ensure that the proportions of each category in the training data are relatively balanced.

[0127] (2) Selection method: In order to achieve uniform selection, stratified sampling is used. Stratified sampling is to further divide each category into several layers based on certain characteristics of the data (such as the distribution and importance of the state), and then extract a certain number of states from each layer. This ensures that the selected states are representative in all feature dimensions.

[0128] S202. Execute training.

[0129] The selected representative states are organized into training data sets for the training of expert strategies. During the training process, these representative states are used as the initial input for each training cycle to help the expert strategy learn the optimal decision under different states, thereby improving the performance and adaptability of the strategy.

[0130] Training process: The selected representative states are input into the expert strategy model, and the model parameters are adjusted according to the model's training algorithm and objective function so that the model can produce the best decision output under these states. The training process usually involves multiple iterations. In each iteration, the model calculates the output results based on the current parameters and compares them with the expected optimal results. The errors are calculated and the parameters are updated through algorithms such as back propagation, gradually improving the performance of the model.

[0131] Effect evaluation: After training is completed, the expert strategy needs to be evaluated on new test data to verify whether the use of selected representative states for training has effectively improved the performance of the strategy. Evaluation indicators can include the execution efficiency, accuracy, stability and other aspects of the strategy. By comparing with the performance of the strategy before training, the effect of training can be evaluated, and the expert strategy can be further adjusted and optimized based on the evaluation results.

[0132] The following is the process of a training cycle of a deep reinforcement learning model, including:

[0133] 1. Environment and Task Definition

[0134] Environmental modeling: Use a physical simulation engine (such as PyBullet, MuJoCo, etc.) to build a simulation environment for the quadruped robot. In the environment, it is necessary to accurately simulate the dynamic characteristics of the quadruped robot, including the robot's mass, inertia, friction and damping of the joints, etc. At the same time, set different terrain conditions (such as flat ground, slopes, and rugged roads) and task goals (such as moving forward, turning, and crossing obstacles).

[0135] State space definition: Determine the state representation of the quadruped robot, which usually includes the robot's joint angles, joint angular velocities, body position and posture (such as roll angle, pitch angle, yaw angle), and body linear velocity and angular velocity. These state information constitute the state space of the PPO algorithm input.

[0136] Action space definition: Define the action space of the quadruped robot, which is generally the control instructions for each joint of the robot, such as the target angle of the joint, the target angular velocity or the applied torque. The design of the action space should take into account the physical limitations of the robot to ensure that the output actions are feasible.

[0137] 2. Reward Function Design

[0138] Basic reward: Design a basic reward function that can reflect the robot's task completion and movement efficiency. For example, for the forward task, you can give the robot a positive reward for moving in the target direction, weighted according to the distance and speed of the movement; for the task of maintaining balance, you can give rewards based on the body's attitude stability (such as the deviation of the pitch angle and roll angle), and the smaller the deviation, the higher the reward.

[0139] Penalty term: Penalty terms are introduced to constrain the robot's bad behavior. For example, if the robot's joint angle exceeds the safe range, or the robot falls (the body posture exceeds a certain threshold), a large negative reward is given. In addition, in order to encourage the robot to save energy, penalty terms can be set based on the motor's output torque or power consumption.

[0140] Task-specific rewards: Add specific rewards based on specific task objectives. For example, in an obstacle crossing task, give extra rewards when the robot successfully crosses an obstacle; in a turning task, give rewards based on the accuracy and fluency of the turn.

[0141] 3. PPO algorithm implementation

[0142] Policy network: Build a neural network as a policy network (actor) to generate actions based on the input state. The structure of the policy network can be a multi-layer perceptron (MLP) or a convolutional neural network (CNN, if the state contains image information). The input is the state of the robot, and the output is the probability distribution over the action space (for discrete action space) or the mean and standard deviation (for continuous action space).

[0143] Value network: Build another neural network as the value network (critic) to estimate the value of the state. The input of the value network is also the state of the robot, and the output is a scalar representing the expected cumulative reward of the state.

[0144] Advantage Estimation: During training, the generalized advantage estimation (GAE) is used to calculate the advantage function. The advantage function represents the advantage of taking an action under the current strategy relative to the average strategy, and is used to guide the update of the policy network.

[0145] Policy update: Use a truncated importance sampling objective function to update the policy network. In each update, a batch of data is sampled from the experience replay buffer, the action probability ratio of the new and old policies is calculated, and then the truncated objective function is calculated based on the action probability ratio. The parameters of the policy network are updated by maximizing this objective function.

[0146] Value network update: At the same time, the value network is updated using the mean squared error (MSE) loss function, with the goal of minimizing the error between the value estimate and the actual accumulated reward.

[0147] 4. Training Process

[0148] Initialization: Initialize the parameters of the policy network and the value network, and set the training hyperparameters, such as learning rate, discount factor, GAE coefficient, truncation parameter, number of training rounds, number of steps per training round, etc.

[0149] Data collection: In the simulation environment, let the quadruped robot perform actions according to the current policy network, interact with the environment, collect information such as state, action, reward, and next state, and store this data in the experience replay buffer.

[0150] Training iteration: In each round of training, multiple batches of data are sampled from the experience replay buffer, and the policy network and value network are updated multiple times. After each update, the new policy network is used to generate actions, continue to interact with the environment, collect new data, and repeat the above process until the training reaches the predetermined number of rounds or meets the stopping condition.

[0151] 5. Evaluation and Optimization

[0152] Evaluation indicators: During the training process, the validation set is used regularly to evaluate the trained strategy. The evaluation indicators include the robot's walking speed, stability (such as fluctuations in the body posture), energy consumption, task completion rate, etc.

[0153] Strategy Adjustment: Based on the evaluation results, adjust the hyperparameters during training, such as learning rate, discount factor, etc., to optimize the performance of the strategy. In addition, if it is found that the robot performs poorly in certain specific situations, the reward function can be further adjusted to strengthen the learning of these situations.

[0154] Migrate to a real robot: After training in a simulation environment to obtain a satisfactory strategy, migrate the strategy to a real quadruped robot for testing. Since there are certain differences between the simulation environment and the real environment (such as model error, sensor noise, etc.), it may be necessary to fine-tune the strategy to adapt to the characteristics of the real environment.

[0155] In some embodiments, the reinforcement learning-based gait learning system may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the reinforcement learning-based gait learning system may be stored in a memory of a computer device and executed by at least one processor to perform (see Figure 1 Description) Functionality of gait learning based on reinforcement learning.

[0156] In this embodiment, the gait learning system based on reinforcement learning can be divided into multiple functional modules according to the functions it performs, such as Figure 2 As shown. The functional modules of the system 200 may include: a data collection module 210, a gait learning module 220. The module referred to in the present invention refers to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, which are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0157] The data collection module is used to collect various states generated during the execution of historical expert strategies as initialization states;

[0158] The gait learning module is used to utilize a pre-built deep reinforcement learning model to perform gait learning based on the initialization state to obtain a gait control strategy.

[0159] Optionally, as an embodiment of the present invention, the data collection module includes:

[0160] A collecting unit, used for collecting states generated during the execution of historical expert strategies, wherein the states include positions, postures, angles and angular velocities of each link of the quadruped robot;

[0161] The storage unit is used to compile the state into an initial state tuple and save the initial state tuple to an initial state library.

[0162] Figure 3The gait learning method based on reinforcement learning provided for the embodiment of the present application can be applied to equipment. It will be appreciated by those skilled in the art that the equipment structure involved in the embodiment of the present invention does not constitute a limitation on the equipment, and the equipment may include more or less components than shown, or combine certain components, or arrange different components. In an embodiment of the present invention, the equipment includes but is not limited to laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The equipment may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples, and are not intended to limit the implementation of the embodiments of the present application described and / or required herein.

[0163] The device 300 may include: a processor 310, a memory 320 and a communication unit 330. These components communicate via one or more buses. Those skilled in the art will appreciate that the server structure shown in the figure does not limit the present invention, and it may be a bus structure or a star structure, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.

[0164] The memory 320 may be used to store the execution instructions of the processor 310, and the memory 320 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. When the execution instructions in the memory 320 are executed by the processor 310, the device 300 is enabled to perform some or all of the steps in the following method embodiments.

[0165] The processor 310 is the control center of the storage device, and uses various interfaces and lines to connect various parts of the entire electronic device. It runs or executes software programs and / or modules stored in the memory 320, and calls data stored in the memory to perform various functions of the electronic device and / or process data. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of a plurality of packaged ICs with the same or different functions. For example, the processor 310 can include only a central processing unit (CPU). In an embodiment of the present invention, the CPU can be a single computing core or multiple computing cores.

[0166] The communication unit 330 is used to establish a communication channel so that the storage device can communicate with other devices, receive user data sent by other devices or send user data to other devices.

[0167] The present invention also provides a computer storage medium, wherein the computer storage medium may store a program, and when the program is executed, the program may include some or all of the steps in each embodiment provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).

[0168] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes, including several instructions for enabling a computer device (which can be a personal computer, a server, or a second device, a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.

[0169] In this specification, the same or similar parts between the various embodiments can be referred to each other. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiment.

[0170] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or modules, which can be electrical, mechanical or other forms.

[0171] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0172] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0173] Although the present invention has been described in detail with reference to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, a person of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions shall be within the scope of the present invention. Any person of ordinary skill in the art may easily think of changes or substitutions within the technical scope disclosed by the present invention, and these shall be within the scope of protection of the present invention.

Claims

1. A gait learning method based on reinforcement learning, characterized in that: include: Collect various states generated during the execution of historical expert strategies as initialization states; Using a pre-built deep reinforcement learning model, gait learning is performed based on the initialization state to obtain a gait control strategy; Collect various states generated during the execution of historical expert strategies as initialization states, including: Collecting states generated during the execution of historical expert strategies, wherein the states include positions, postures, angles, and angular velocities of each link of the quadruped robot; Compiling the state into an initial state tuple, and saving the initial state tuple to an initial state repository; The initial state tuples in the initial state library are clustered, and representative states are evenly selected from each category after clustering as the initialization state of the learning cycle; in the clustering process, when calculating the distance between different states, momentum is introduced instead of speed to measure the distance between different states, that is, for state ( , , , ), when calculating the distance between this state and other states, use ( , , , ) as the calculation basis, where For location, For posture, For speed, is the angular velocity, For quality, is the moment of inertia.

2. The method according to claim 1, characterized in that: The method further comprises: Defining a sample set , neighborhood parameters ; Initialize the core object collection , initialize the number of clusters , initialize the unvisited sample set ,Cluster Partition ; Find samples by distance measurement of -Neighborhood subsample set ; Confirm that the number of samples in the subsample set meets , the sample Add core object sample collection ; In the core object collection In the example above, randomly select a core object , initialize the current cluster core object queue , Initialize the category number , initialize the current cluster sample set , Update the unvisited sample set ; In the current cluster core object queue Extract a core object , through the neighborhood distance threshold Find all -Neighborhood subsample set ,make , Update the current cluster sample set , Update the unvisited sample set ,renew ;in, is the neighborhood subsample set The intersection with the set of unvisited samples; Confirm the current cluster core object queue , then the current cluster Generation completed, update cluster division , Update the core object collection , and confirm the core object set , the processing is determined to be complete, the cluster division results are output, and samples are evenly selected from multiple clusters as the initial state of each learning cycle of the deep reinforcement learning model.

3. The method according to claim 2, characterized in that Find samples by distance measurement of -Neighborhood subsample set ,include: Calculation Sample With sample The Euclidean distance of the momentum state, if the Euclidean distance does not exceed , the sample is classified into the -Neighborhood subsample set .

4. The method according to claim 2, characterized in that: The method further comprises: Calculate the potential difference between any two samples in the subsample set and calculate the average value of the potential difference; For isolated samples that are not included in the neighborhood subsample set, the potential difference between the isolated sample and the core sample of each subsample set is calculated, and the target subsample set whose potential difference does not exceed the corresponding potential difference average value is selected; If there is a target sub-sample set, the isolated sample is classified into the target sub-sample set; if there is no target sub-sample set, the isolated sample is determined to be noise.

5. The method according to claim 1, characterized in that The method further comprises: The state data generated by the gait control strategy during the test are collected, and the state data are marked with the parameters of the quadruped robot and then imported into the initial state library.

6. A gait learning system based on reinforcement learning, characterized in that: include: The data collection module is used to collect various states generated during the execution of historical expert strategies as initialization states; A gait learning module, configured to use a pre-built deep reinforcement learning model to perform gait learning based on the initialization state to obtain a gait control strategy; The data collection module includes: A collecting unit, used for collecting states generated during the execution of historical expert strategies, wherein the states include positions, postures, angles and angular velocities of each link of the quadruped robot; A storage unit, used for compiling the state into an initial state tuple, and saving the initial state tuple to an initial state library; The initial state tuples in the initial state library are clustered, and representative states are evenly selected from each category after clustering as the initialization state of the learning cycle; in the clustering process, when calculating the distance between different states, momentum is introduced instead of speed to measure the distance between different states, that is, for state ( , , , ), when calculating the distance between this state and other states, use ( , , , ) as the calculation basis, where For location, For posture, For speed, is the angular velocity, For quality, is the moment of inertia.

7. A device, characterized in that: include: A memory, used for storing a gait learning program based on reinforcement learning; A processor, used to implement the steps of the gait learning method based on reinforcement learning as described in any one of claims 1 to 5 when executing the gait learning program based on reinforcement learning.

8. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores a gait learning program based on reinforcement learning, and when the gait learning program based on reinforcement learning is executed by a processor, the steps of the gait learning method based on reinforcement learning as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Self-adaptive gait planning method, system and device for hexapod robot and medium

    CN114326722A

  • Supply chain management method and system based on volatility clustering and block chain

    CN115689443A

  • Security assessment and application method for reinforcement learning type control strategy in high-dimensional space

    CN118444659A