Robot walking control method, device, equipment, storage medium and product
By establishing a mapping relationship between the robot's walking state and control strategy through reinforcement learning, and combining it with reference gait patterns, the problem of insufficient flexibility and adaptability of the robot in complex environments is solved, and the robot can walk stably in complex environments.
Patent Information
- Application Number
- CN202411730365.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2044-11-28
AI Technical Summary
In existing technologies, robot walking control methods have poor flexibility and adaptability in complex and changing environments, making it difficult to cope with environmental changes.
By using reinforcement learning, the robot can autonomously explore its environment, learn optimal behavioral strategies, establish a mapping relationship between walking states and control strategies, and combine reference gait patterns to guide walking postures, thereby improving the robot's walking flexibility and adaptability in complex environments.
It improves the robot's walking control flexibility and adaptability in complex environments, and enhances the robot's stability and efficiency in practical applications.
Smart Images

Figure CN119610091B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, specifically to a control method, device, equipment, storage medium, and product for robot walking. Background Technology
[0002] With the continuous advancement and development of computer technology, robots, as intelligent machines capable of semi-autonomous or fully autonomous operation, are increasingly being widely used because they can perform a variety of tasks through programming and automated control. For example, wheel-legged robots, which combine wheeled and legged locomotion, consist of multiple joints, motors, sensors, legs, wheels, etc., and can mimic the walking posture of humans or animals, moving through different terrains and environments.
[0003] Currently, robot movement can be controlled using rule-based control methods. This involves using a predefined set of rules; when the robot encounters a specific situation, it executes corresponding actions according to these rules. This approach decomposes the robot's behavior into a series of rules that define the robot's actions and reactions in different situations. While this method is simple to implement and easy to understand, it requires redefining a large number of rules when the environment changes, resulting in poor flexibility and adaptability, making it difficult to cope with complex and changing environments. Summary of the Invention
[0004] This application provides a robot walking control method, device, equipment, storage medium, and product, which helps to improve the flexibility and adaptability of robot walking control, enabling the robot to walk stably in complex and ever-changing environments.
[0005] In a first aspect, embodiments of this application provide a control method for robot walking, including:
[0006] Obtain the robot's current walking status and the robot's accumulated walking time;
[0007] Based on the walking time and the robot's gait planning information, the robot's gait information at the current moment is determined, and the gait information and the walking state at the current moment are taken as the target walking state at the current moment;
[0008] Based on the first mapping relationship between the target walking state and the control strategy, the target control strategy mapped to the target walking state at the current moment is determined; wherein, the first mapping relationship is determined based on the target walking states at multiple moments prior to the current moment.
[0009] The robot's movement is controlled based on the target control strategy.
[0010] Secondly, embodiments of this application provide a control device for robot walking, comprising:
[0011] The acquisition unit is used to acquire the robot's current walking status and the robot's walking time.
[0012] The determining unit is configured to determine the robot's gait information at the current moment based on the walking time and the robot's gait planning information, and to use the gait information and the walking state at the current moment as the target walking state at the current moment;
[0013] The determining unit is configured to determine the target control strategy mapped to the target walking state at the current moment based on a first mapping relationship between the target walking state and the control strategy; wherein the first mapping relationship is determined based on the target walking states at multiple moments prior to the current moment.
[0014] A control unit is used to control the robot's movement based on the target control strategy.
[0015] In one possible implementation, the robot includes a first set of motion components and a second set of motion components; the gait planning information includes a reference stride period and a reference gait offset; the determining unit is used to determine the gait information at the current moment based on the walking time and the robot's gait planning information, specifically for:
[0016] The gait phase at the current moment is determined based on the ratio of the walking time to the reference stride period.
[0017] Based on the gait phase at the current moment and the first gait offset associated with the first set of moving parts in the reference gait offset, the first gait information is determined;
[0018] Based on the gait phase at the current moment and the second gait offset associated with the second set of motion components in the reference gait offset, the second gait information is determined; the gait information at the current moment includes the first gait information and the second gait information.
[0019] In one possible implementation, the first mapping relationship is obtained through reinforcement learning, and the determining unit is further configured to determine the state feedback at each time step during the reinforcement learning process based on the target walking state at each time step among the plurality of time steps.
[0020] The determining unit is further configured to determine the state value mapped to the target walking state at each time step based on the second mapping relationship between the target walking state and the state value.
[0021] The determining unit is further configured to determine the advantage function at each moment in the reinforcement learning process based on the state feedback at each moment and the state value at each moment;
[0022] The update unit is also used to update the historical mapping relationship of the reinforcement learning based on the advantage function at each time point to obtain the first mapping relationship.
[0023] In one possible implementation, the target walking state at each moment includes gait information associated with the robot's motion components; the determining unit is configured to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each moment among the plurality of moments, specifically for:
[0024] Based on the set gait information masking rules, the gait information associated with the moving parts is converted into a gait mask to obtain the gait mask at each moment.
[0025] The gait indication information of the moving component at each time point is obtained, and the gait indication information is used to indicate whether the moving component is in a specific gait.
[0026] Based on the degree of matching between the gait mask and the gait indication information at each time point, the state feedback at each time point is determined.
[0027] In one possible implementation, the target walking state at each moment includes the linear velocity of the robot in the target direction and other directions, wherein the other directions are perpendicular to the target direction; the determining unit is configured to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each of the plurality of moments, specifically configured to:
[0028] The linear velocity reward at each time point is determined based on the linear velocity in the target direction, the set linear velocity in the target direction, and the reward parameter associated with the target direction.
[0029] The linear velocity penalty at each time moment is determined based on the linear velocity in the other directions and the penalty parameters associated with the other directions;
[0030] The state feedback at each time step is determined based on the sum of the linear velocity reward and the linear velocity penalty at each time step.
[0031] In one possible implementation, the target walking state at each moment includes the robot's angular velocity in a specified direction and the robot's pose projection in a specific direction; the determining unit is configured to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each of the plurality of moments, specifically configured to:
[0032] The angular velocity penalty at each moment is determined based on the angular velocity in the specified direction and the penalty parameter associated with the angular velocity;
[0033] The attitude penalty at each time moment is determined based on the attitude projection in the specific direction and the penalty parameter associated with the attitude projection.
[0034] Based on the angular velocity penalty and the attitude penalty at each time step, the state feedback at each time step is determined.
[0035] In one possible implementation, the target walking state at each moment includes the angular position of each joint among multiple joints of the robot; the determining unit is used to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each moment among the multiple moments, specifically for:
[0036] Based on the angular position of each joint, the angular position limit value corresponding to each joint, and the penalty parameter associated with the angular position, the joint position penalty corresponding to each joint at each time is determined;
[0037] Based on the sum of the joint position penalties corresponding to the plurality of joints at each time, the joint angle position penalty at each time is determined, and based on the joint angle position penalties at each time, the state feedback at each time is determined.
[0038] In one possible implementation, the target walking state at each moment includes the angular velocity of each joint among multiple joints of the robot; the determining unit is configured to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each moment among the multiple moments, specifically for:
[0039] Based on the angular velocity of each joint and the angular velocity of each joint at the previous moment, determine the acceleration of each joint at each moment.
[0040] Based on the accelerations of the multiple joints at each time point and the penalty parameters associated with the accelerations, the joint acceleration penalties at each time point are determined, and based on the joint acceleration penalties at each time point, the state feedback at each time point is determined.
[0041] In one possible implementation, the acquisition unit is further configured to acquire the joint torque of each joint of the robot at each time.
[0042] The determining unit is further configured to determine the joint torque penalty at each time based on the maximum joint torque value corresponding to each joint, the joint torque of each joint at each time, and the penalty parameter associated with the joint torque.
[0043] The determining unit is configured to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each of the plurality of moments, specifically for:
[0044] Based on the target walking state at each time point and the joint torque penalty at each time point, the state feedback at each time point is determined.
[0045] In one possible implementation, the acquisition unit is further configured to acquire the control strategies at each time point;
[0046] The calculation unit is used to perform differential calculations based on the control strategies at each time point and the control strategies at adjacent time points to obtain the policy change rate corresponding to each time point.
[0047] The determining unit is used to determine the policy smoothness penalty at each time based on the control policy at each time, the policy change rate at each time, and the penalty parameter associated with the policy change rate.
[0048] The determining unit is configured to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each of the plurality of moments, specifically for:
[0049] Based on the target walking state at each time step and the policy smoothness penalty at each time step, the state feedback at each time step is determined.
[0050] In one possible implementation, the determining unit is configured to determine the target control strategy mapped to the current target walking state based on a first mapping relationship between the target walking state and the control strategy, specifically configured to:
[0051] The target walking state at the current moment is input into the policy network in the reinforcement learning model to obtain the target control policy output by the policy network; wherein, the policy network is used to indicate the first mapping relationship.
[0052] In one possible implementation, the updating unit is used to update the historical mapping relationship of the reinforcement learning based on the advantage function at each time step to obtain the first mapping relationship, specifically for:
[0053] Based on the advantage function at each time point, the target walking state at each time point, and the control strategy at each time point, the strategy loss data is determined.
[0054] The model parameters of the historical policy network are updated based on the policy loss data to obtain the policy network, wherein the historical policy network is used to indicate the historical mapping relationship.
[0055] In one possible implementation, the reinforcement learning model further includes a value network, and the determining unit is used to determine value loss data based on the advantage function and the state value at each time step.
[0056] The updating unit is further configured to update the model parameters of the historical value network based on the value loss data to obtain the value network; wherein the historical value network is used to indicate the second mapping relationship.
[0057] In one possible implementation, the current target walking state includes the angular positions of each joint of the robot, and the target control strategy is used to indicate the target angular positions of each joint; the control unit is used to control the robot's walking based on the target control strategy, specifically for:
[0058] Based on the target angle position of each joint, the angle position of each joint, and the default angle position of each joint, the angle position error of each joint is determined.
[0059] Proportional-derivative control processing is performed based on the angular position error of each joint to determine the control torque of each joint.
[0060] The robot moves by controlling each joint based on the control torque of each joint.
[0061] In one possible implementation, the acquisition unit is further configured to acquire the current angular position of each joint at a set time interval, and determine the current angular position error of each joint based on the current angular position of each joint and the default angular position of each joint.
[0062] The processing unit is used to perform proportional-derivative control processing based on the current angular position error of each joint to determine the current control torque of each joint.
[0063] The control unit is also used to control each joint again based on the current control torque of each joint until the number of control operations reaches a set number.
[0064] The determining unit is further configured to determine the walking state of the robot after the number of control operations reaches the set number as the walking state at the next moment.
[0065] Thirdly, embodiments of this application provide an electronic device, which includes one or more processors and a memory for storing one or more computer programs. When the one or more computer programs are executed by the one or more processors, the electronic device enables the robot walking control method described in the first aspect.
[0066] Fourthly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the robot walking control method described in the first aspect.
[0067] Fifthly, embodiments of this application provide a computer program product, which includes a computer program or computer instructions. When the computer program or computer instructions are executed by a processor, they implement the robot walking control method as described in the first aspect.
[0068] In some embodiments of this application, the technical solutions first acquire the robot's current walking state and the duration of its walking. Then, based on the duration of walking and the robot's gait planning information, the robot's gait information at the current moment is determined, and the gait information and the current walking state are used as the target walking state at the current moment. Next, based on the mapping relationship between the target walking state and the control strategy, the target control strategy mapped to the target walking state at the current moment is determined. This mapping relationship is determined based on the target walking states at multiple moments prior to the current moment. Finally, the robot's walking can be controlled based on the target control strategy. Therefore, by obtaining the mapping relationship through reinforcement learning through the interaction between the robot and the environment to determine the current control strategy, the robot can learn and optimize autonomously to cope with complex and changing environments, which is beneficial to improving the flexibility and adaptability of robot walking control. Furthermore, determining the control strategy through the mapping relationship helps improve the real-time performance of robot walking. Moreover, by introducing a reference gait pattern as the initial walking pattern, the robot can be guided to learn walking postures, thereby improving the efficiency of the robot's autonomous learning. Attached Figure Description
[0069] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0070] Figure 1 This is a schematic diagram of the architecture of a robot walking control system provided in an embodiment of this application;
[0071] Figure 2 This is a schematic diagram of the architecture of another robot walking control system provided in an embodiment of this application;
[0072] Figure 3 This is a schematic diagram of a simulation environment built based on Isaac Gym, provided in an embodiment of this application;
[0073] Figure 4 This is a schematic diagram of a quadrupedal humanoid robot provided in an embodiment of this application;
[0074] Figure 5 This is another schematic diagram of a quadrupedal humanoid robot provided in an embodiment of this application;
[0075] Figure 6 This is a flowchart illustrating a robot walking control method provided in an embodiment of this application;
[0076] Figure 7 This is a schematic diagram of the gait phase trajectory of a robot provided in an embodiment of this application;
[0077] Figure 8 This is a schematic diagram illustrating the principle of robot walking control provided in an embodiment of this application;
[0078] Figure 9 This is a schematic diagram of a robot walking in a simulation environment, provided in an embodiment of this application.
[0079] Figure 10 This is a schematic diagram of the robot's action sequence during walking, provided in an embodiment of this application;
[0080] Figure 11 This is a schematic diagram of the structure of a robot walking control device provided in an embodiment of this application;
[0081] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0082] It should be noted in advance that, in order to enable those skilled in the art to better understand the technical solutions proposed in the embodiments of this application, the embodiments of this application will be described clearly and completely in conjunction with one or more accompanying drawings. Furthermore, the various drawings shown in the embodiments of this application are merely illustrative examples; for example, the execution order of each step in the drawings can be adaptively adjusted according to the actual application scenario. In addition, in the embodiments of this application, the block diagrams shown in the various drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, or in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0083] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0084] It should be noted that "multiple" in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0085] With the rapid development of computer technology and artificial intelligence, robots, capable of performing various tasks through programming and automated control and adapting to diverse environments, are increasingly being widely applied in industrial manufacturing, medical, and service sectors. For example, wheeled-legged robots, which combine wheeled and legged locomotion, can consist of multiple joints, motors, sensors, legs, wheels, etc., and can mimic human or animal walking postures to move in different environments. Due to the highly complex interaction between the robot and the ground, robot locomotion control is an important and challenging problem in robotics research. Currently, to achieve stable robot locomotion, control strategies can be determined in two ways:
[0086] One approach is rule-based control. This method uses a predefined set of rules to guide the robot's actions in specific situations. This approach breaks down the robot's behavior into a series of rules that define the robot's actions and reactions in different circumstances.
[0087] Another approach is Model Predictive Control (MPC). MPC describes the robot's motion by establishing an accurate dynamic model and predicting the robot's future motion trend based on the current state. Based on this, it determines the optimal control input to achieve precise control of the robot's movement.
[0088] While the former is simple to implement and easy to understand, it requires redefining a large number of rules when the environment changes, resulting in poor flexibility and adaptability, making it difficult to cope with complex and ever-changing environments. The latter, compared to the former, can improve the stability and adaptability of robot walking to some extent, but its computational complexity is higher, its real-time performance is poorer, and its adaptability is lower, making it difficult to widely promote in practical applications.
[0089] Based on this, this application provides a robot walking control scheme that employs reinforcement learning. Reinforcement learning enables the robot to autonomously explore its environment and learn optimal behavioral strategies without requiring labeled samples. By learning from past walking states in the environment, the robot determines the mapping relationship between its walking states and control strategies, which can then be used to generate control strategies for the robot. This allows the robot to walk in different environments, improving its walking flexibility. Furthermore, by using gait patterns as guidance, the robot's posture during walking is determined, which to some extent improves the efficiency of learning optimal behavioral strategies. This enables the robot to walk more efficiently and stably in its environment, improving its adaptability and better meeting the needs of robot walking in practical applications.
[0090] To better understand the solutions of the embodiments of this application, the relevant terms and concepts that may be involved in the embodiments of this application will be introduced below.
[0091] 1. Agent: An agent is an entity that can perceive the environment through sensors and act on the environment through actuators. A reinforcement learning agent is a computational entity based on reinforcement learning algorithms, such as the robot in this application embodiment.
[0092] 2. Reinforcement Learning (RL): A crucial branch of machine learning, RL enables agents to learn policies through exploration within their environment without requiring training on labeled samples. The process of reinforcement learning can be understood as a training process. It involves mapping environmental states to an action space, typically described using a Markov Decision Process (MDP). An agent exists in an environment where each state represents the agent's perception of the current environment. The agent can influence the environment through actions, causing the environment to transition to another state with a certain probability. Simultaneously, the environment provides a reward to the agent. In other words, reinforcement learning learns an optimal policy that allows the agent to take actions based on its current state within a specific environment to maximize its reward.
[0093] In the field of reinforcement learning, proximal policy optimization (PPO) is a commonly used algorithm. Its core idea lies in using a loss function that includes a clipping term. This limit the amount of change in the policy during each update, preventing large fluctuations during policy updates, thus stabilizing training and improving learning efficiency.
[0094] In reinforcement learning, the state space, observation space, and action space define the basic framework for the interaction between the agent and the environment.
[0095] A state space is the set of all possible states used to describe a specific situation in an environment. A state space can be discrete, such as each position on a chessboard in a game of chess, or continuous, such as the coordinates of a robot's position in an environment.
[0096] The observation space refers to the set of all possible observations that an agent can receive during its interaction with the environment, and can be used by the agent to make decisions. The observation space is only the set of environmental information that can be perceived; it may contain all information about the state or only some information.
[0097] Action space refers to the set of all possible actions an agent can perform given a state. In the PPO algorithm, the action space can be continuous or discrete. A continuous action space means that the actions taken by the agent are continuous values within the real number range; for example, in robot control tasks, the robot's joint angles are continuous variables. A discrete action space means that the agent can only take a priority number of discrete actions; for example, the agent can only choose a limited number of actions such as "go right" or "go left".
[0098] In the embodiments of this application, the robot is an intelligent agent. The robot can exist in a real environment or a simulated environment. Each state of the robot can be a perception of the current environment at different times, which can be the robot's walking state perceived by configured sensors. At the current moment, the robot can determine a control strategy to execute the actions indicated by the strategy, enabling the robot to acquire perception of the current environment at the next moment and obtain its state for that next moment. Furthermore, based on the robot's state, corresponding feedback (reward or punishment) can be determined, and this feedback can be used to guide the robot to learn the optimal control strategy, enabling the robot to walk stably in the environment.
[0099] Based on the above description, please refer to Figure 1 , Figure 1 This is a schematic diagram of the architecture of a robot walking control system provided in an embodiment of this application, as shown below. Figure 1 As shown, the robot's walking control system includes a robot 101 and a robot control device 102. The robot 101 and the robot control device 102 can be connected directly or indirectly via wired or wireless means. It should be noted that... Figure 1 The number and configuration of devices shown are for illustrative purposes only and do not constitute a limitation on the embodiments of this application. In some embodiments, there may be multiple robots 101. In some embodiments, there may be multiple robot control devices 102 connected to the robot (such as robot 101). In some embodiments, robot 101 and robot control device 102 may be the same electronic device, which is not limited in this application.
[0100] Robot 101 is a mechatronic electronic device capable of automatically performing tasks. It can be defined as an autonomous or semi-autonomous system with the ability to perceive its environment, process information, and perform actions. Robots typically possess sensors (such as cameras, lidar, accelerometers, and inertial sensors) to perceive their surroundings. For example, a robot can determine its current walking state by acquiring its linear velocity, angular velocity, projected gravity, joint angle positions, and joint angular velocities. Robots also typically have processors (such as microcontrollers, central processing units (CPUs), and graphics processing units (GPUs)) to process data and make decisions, and actuators (such as motors and hydraulic cylinders) to perform physical actions, such as walking. Robot 101 can have various different forms, for example... Figure 1 The humanoid robot shown includes four-wheeled robots, which can be called wheeled robots, humanoid robots, quadruped robots, quadruped humanoid robots, etc., to meet the needs of different application scenarios.
[0101] Specifically, the robot control device 102 can be a terminal device or a server. The terminal device can include, but is not limited to, smartphones (such as Android phones, iOS phones, etc.), tablet computers, portable personal computers, mobile internet devices (MIDs), smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, wearable devices, etc. This application embodiment does not limit this. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. This application embodiment does not limit this.
[0102] The general flow of the robot walking control method provided in this application is as follows:
[0103] The robot control device 102 can acquire the current walking state and the accumulated walking time of the robot 101. For example, the robot control device 102 can receive the walking state sent by the robot 101, and can also calculate the accumulated walking time of the robot 101. Furthermore, based on gait planning information and the accumulated walking time, the robot control device 102 can determine the gait information of the robot 101 at the current moment, and use the gait information and the current walking state as the target walking state of the robot 101 at the current moment. Then, based on the mapping relationship between the target walking state and the control strategy, the robot control device 102 can determine the target control strategy mapped to the target walking state at the current moment. Finally, the robot control device 102 can control the robot 101 to walk based on the target control strategy.
[0104] The mapping relationship can be obtained by the robot control device 102 through reinforcement learning based on the target walking states of the robot 101 at multiple time points prior to the current time. The robot control device 102 can determine the state feedback at each time point during the reinforcement learning process based on the target walking states of the robot 101 at each of the multiple time points, and then determine the state value mapped to the target walking state at each time point based on the mapping relationship between the target walking states and state values. Subsequently, the robot control device 102 can determine the dominance function at each time point during the reinforcement learning process based on the state feedback and state values at each time point. Finally, the robot control device 102 can update the historical mapping relationship based on the dominance function at each time point to obtain the first mapping relationship.
[0105] Please refer to the following: Figure 2 , Figure 2 This is a schematic diagram of the architecture of another robot walking control system provided in an embodiment of this application, as shown below. Figure 2 As shown, the robot's walking control system includes only the robot simulation device 201. The robot simulation device 201 provides a walking environment for performing the robot's walking tasks, specifically, the robot can perform actions such as... Figure 2 The walking task shown is performed in a flat environment, which simulates robot walking. This environment can be called a simulation environment. In this simulation environment, the robot can learn the optimal control strategy through reinforcement learning to achieve stable walking in the simulation environment.
[0106] The robot simulation device 201 can be a terminal device or a server. Terminal devices can include, but are not limited to, smartphones (such as Android phones, iOS phones, etc.), tablets, portable personal computers, MIDs, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, wearable devices, etc. This application embodiment does not limit this. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. This application embodiment does not limit this. It should be noted that, since the simulation environment has certain requirements on the video memory and computing power of electronic devices such as GPUs, the robot simulation device 201 can be an electronic device with specific hardware to support the operation of the simulation environment.
[0107] The general flow of the robot walking control method provided in this application is as follows:
[0108] In the simulation environment, the robot simulation device 201 can determine the mapping relationship between the robot's target walking state and the control strategy based on the robot's target walking states at multiple time points prior to the current time. It then acquires the robot's walking state and the duration of its walking at the current time, and determines the robot's gait information at the current time based on a reference gait pattern and the duration of walking. This gait information and the current walking state are used as the robot's target walking state at the current time. Therefore, based on this mapping relationship, the robot simulation device 201 can determine the target control strategy mapped to the target walking state at the current time, and subsequently control the robot's movement based on the target control strategy.
[0109] For ease of description, this application uses the simulation environment built on IsaacGym as an example to illustrate the robot simulation device 201. The robot simulation device 201 can also run other simulation environments, and this application does not limit this.
[0110] Isaac Gym is a physics simulation framework specifically developed for robotics and reinforcement learning tasks, providing a simulator for building simulation environments. It includes an efficient physics engine, powerful collision handling capabilities, and flexible modeling tools, enabling fast and accurate physical simulations. This simulation environment also supports large-scale parallel simulations, suitable for complex robot control tasks, such as simulating 100 robots performing a walking task in parallel. Furthermore, Isaac Gym integrates a deep learning library (PyTorch), enabling the implementation of various reinforcement learning algorithms, such as the PPO algorithm in this embodiment. Since the robot's walking task in this embodiment involves a large amount of physical collision and dynamics calculations, the task scenario of the robot walking in the environment can be built within Isaac Gym.
[0111] Specifically, a simulation environment can be built in Isaac Gym, which can include the ground, walls, obstacles, etc., to simulate the robot's real-world movement scenarios. Furthermore, to improve the robot's adaptability and robustness in various environments, the scenario built in the simulation environment can be modified, for example, through domain randomization. Domain randomization refers to randomizing various parameters in the simulation environment, allowing the robot to encounter changes in environmental conditions during the learning process. This enables the robot to learn the optimal control strategy under different external conditions, thereby enhancing its robustness and adaptability in practical applications.
[0112] For example, in a simulation environment, the coefficient of friction of the ground can be set to a random number within the range of 0.5 to 1.25 to simulate the friction characteristics under different ground conditions. As another example, the simulation environment can be programmed to subject the robot to a random thrust at predetermined intervals, such as 15 seconds, to simulate external disturbances the robot might encounter in the real world. This helps improve the robot's stability and recovery ability under such external disturbances. The maximum inference speed can be set, for example, to 1.0 m / s, to test the robot's response and adaptability under different inference conditions. Thus, through domain randomization techniques, the robot can encounter various environmental conditions in the simulation environment during reinforcement learning, thereby enhancing its robustness and adaptability in practical applications.
[0113] After setting up the simulation environment in Isaac Gym, you can import or create robot models. For example, you can write a configuration file, which can be a robot model file. The configuration file can define the robot's geometry, components, actuators, and other attributes for physical modeling within the simulation environment. Please refer to [link / reference]. Figure 3 , Figure 3This is a schematic diagram of a simulation environment built based on Isaac Gym, provided in an embodiment of this application. The simulation environment includes a robot body, which is a four-legged humanoid robot. Each leg can be equipped with a leg and a wheel; it can also be called a four-wheeled humanoid robot, or XMAN. The simulation environment also includes the ground environment on which the robot stands, such as a flat ground environment composed of black and white grids. Optionally, the simulation environment may also include information about the robot's component structure, such as... Figure 3 As shown in the upper left corner, you can select different components to obtain their current detailed information. For example, by selecting a robot joint, you can obtain the robot's joint angle position information and angular velocity information.
[0114] After the simulation environment and robot model are deployed, the simulation can be run. During training, the robot can iterate multiple times in the simulation environment built on Isaac Gym. In each iteration, the robot can acquire state information through interaction with the environment and generate corresponding actions based on the learned control strategy. During this process, Isaac Gym calculates the robot's state changes based on the deployed physical parameters and robot model, and updates the simulation environment in real time. Thus, the robot can continuously optimize its strategy in the simulation environment built on Isaac Gym to gradually learn to walk stably.
[0115] In one implementation, the aforementioned walking state, reference gait pattern, gait information, and the robot's target walking state and target control strategy at various moments can all be stored in the blockchain, preventing this information from being tampered with. Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, it is a decentralized database, a chain of data blocks linked using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block.
[0116] It is understood that the embodiments described in this application are as follows: Figure 1 and Figure 2 The robot walking control system shown is for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0117] In this application embodiment, the robot can refer to a humanoid robot with multiple legs, each leg including a mechanical leg and a mechanical wheel. Therefore, the robot can also be called a foot-wheel hybrid robot or a wheel-leg robot. The mechanical legs and mechanical wheels in the robot can be used to perform walking tasks. For ease of description, this application embodiment uses a humanoid robot with four wheels and legs as an example for explanation, and this application does not limit the shape of the robot.
[0118] Please refer to the following: Figure 4 and Figure 5 , Figure 4 and Figure 5 These are all schematic diagrams of a quadrupedal humanoid robot provided in the embodiments of this application, such as... Figure 4 As shown, the robot 400 in this embodiment of the application may consist of a head 401, upper limbs 402, torso 403, waist 404, and motion components. The motion components may include legs (mechanical legs) and wheels (mechanical wheels). Taking a quadrupedal humanoid robot as an example, the robot 400 may have four mechanical legs: two outer mechanical legs 405 located on the outside, which can be simply referred to as outer feet, and two inner mechanical legs 406 located on the inside, which can be simply referred to as inner feet. Mechanical wheels 407 are installed at the ends of the mechanical legs, and each mechanical wheel can be driven independently.
[0119] Figure 4 Taking the robot 400 with mechanical wheels 407 connected to the ends of its outer mechanical leg 405 and inner mechanical leg 406 as an example, Figure 5 It is also shown that mechanical feet 408 are mounted on the outer side of the mechanical wheels 407 connected to the ends of the outer mechanical legs 405 and the inner mechanical legs 406. Each mechanical wheel can be driven independently, and each mechanical foot can rotate independently. The mechanical feet can be located on the left side of the mechanical wheels, or on the right side of the mechanical wheels, or in a hollowed-out style at the root hollowed-out area of the mechanical feet; this application does not limit this.
[0120] The robot 400 can stand on the ground using either its outer mechanical leg 405 or its inner mechanical leg 406, thus achieving a bipedal standing position. The leg used for standing can be referred to as the supporting leg. The robot 400 can also stand on both its outer and inner mechanical legs simultaneously, achieving a quadrupedal standing position. The robot 400 can glide on the ground using the mechanical wheels on its outer or inner mechanical legs 405, and can also move (walk) by controlling the alternating swinging of its outer and inner mechanical legs 405.
[0121] The outer mechanical leg 405 and the inner mechanical leg 406 of robot 400 can respectively follow along Figure 4 and Figure 5The extension and retraction are performed independently in the directions shown by the double-headed arrows. The outer mechanical leg 405 can rotate around the hip rotation center 409, and the two outer mechanical legs remain linked. Similarly, the inner mechanical leg 406 can also rotate around the hip rotation center 410 (also called the pitch rotation center 410), and the two inner mechanical legs also remain linked. The hip rotation centers of the outer mechanical leg 405 and the inner mechanical leg 406 are driven independently, and the two hip rotation centers are located in the same vertical plane. The outer mechanical leg 405 and the inner mechanical leg 406 of the robot 400 can be symmetrically distributed on both sides of the robot's central axis (i.e., the sagittal plane 411).
[0122] In some embodiments, the hip rotation center of the outer mechanical leg 405 and the hip rotation center of the inner mechanical leg 406 can be coaxial, that is, the two hip rotation centers are located on the same straight line, or they can be non-coaxial, which is not limited in this application. Figure 4 and Figure 5 The drawing and explanation will be based on the example of two hip rotation centers being coaxial.
[0123] The upper end of the robotic legs (outer robotic leg 405 and inner robotic leg 406) is connected to one end of the robot's waist 404, and the other end of the waist 404 is connected to one end of the robot's torso 403. The waist 404 includes two rotation centers (also called drive joints): a pitch rotation center 410 (i.e., hip rotation center 410) that enables the torso 403 to pitch, and a lateral rotation center that enables the torso 403 to yaw. The pitch rotation center 410 and the lateral rotation center can be designed in series and are located above the pitch rotation center 410, connected to the robot's torso 403.
[0124] The robot 400 has a head 401 connected to the upper end of its body 403, and a multi-degree-of-freedom upper limb 402 connected to each side of the body 403. In some embodiments, the upper limb 210 is equipped with an end effector, such as a gripper or a suction cup. Data acquisition devices can be deployed in the head 209 to perceive the real or simulated environment. These data acquisition devices may include, but are not limited to, image acquisition devices, video recording devices, and inertial measurement units (IMUs).
[0125] In some embodiments, the IMU may be placed at the geometric center of the body 403, the center of hip rotation, etc., and can be used to measure the acceleration, attitude angular velocity, projected gravity in the x-axis, y-axis and z-axis directions of the body 403.
[0126] In the embodiments of this application, the waist 404, motion components (outer mechanical leg 405 and inner mechanical leg 406), mechanical wheel 407, body 403 (including IMU), and multiple drive joints (such as 12 drive joints) used for walking in the robot 400 are essential hardware for the control algorithm, while the rest are non-essential hardware.
[0127] Based on the aforementioned robot walking control system, this application provides a robot walking control method. The robot walking control method described in this application can be executed by an electronic device, which can be... Figure 1 The robot control device 101 in the robot walking control system shown can also be... Figure 2 The robot simulation device 201 is shown in the robot walking control system. Please refer to [link / reference]. Figure 6 , Figure 6 This is a flowchart illustrating a robot walking control method provided in an embodiment of this application. The robot walking control method includes the following steps S601-S604:
[0128] S601. Obtain the robot's current walking status and the robot's accumulated walking time.
[0129] In this application embodiment, a robot can refer to a mechatronic electronic device capable of automatically performing tasks. It can be defined as an autonomous or semi-autonomous system with the ability to perceive the environment, process information, and perform actions. For specific robot configurations, please refer to... Figure 4 and Figure 5 As shown, because it is equipped with multiple mechanical legs and wheels for movement, this type of robot can also be called a wheeled-legged robot. It can be equipped with sensors (such as cameras, lidar, accelerometers, IMUs, etc.) to perceive the surrounding environment, processors (such as microcontrollers, CPUs, GPUs, etc.) to process data and make decisions, and actuators (such as motors) to perform physical actions, such as walking.
[0130] The robot referred to in this application can be an intelligent machine operating semi-autonomously or fully autonomously in a real environment, or a robot model in a simulation environment, which also includes the aforementioned sensors, processors, and actuators. The simulation environment can simulate the real-world scenario in which the robot operates, such as the ground or walls where it walks. Physical parameters related to the environment and the robot can be deployed in the simulation environment. Based on the robot model and these physical parameters, the robot's state changes in the simulation environment can be calculated, and the simulation environment can be updated in real time. For ease of description, this application uses a robot model in a simulation environment as an example to illustrate the robot's ability to perform walking tasks in a specific environment. This specific environment can be, for example, a flat ground environment; this application does not limit this.
[0131] In some embodiments, configuration files (such as robot model files) can be written and imported into the simulation environment to import such... Figure 4 or Figure 5 The robot model shown can also be created in a simulation environment, and this application does not limit this. After the robot is added to the simulation environment, the robot's initial state can be configured, such as the initial angle and position of each joint in the robot, as well as the initial state of the environment, such as the ground friction coefficient being 0, or the position of obstacles on the ground, and the robot can also be controlled to start walking.
[0132] It is understandable that during a robot's walking process, different moments correspond to different walking states. These walking states can include multiple features of the robot, which collectively describe its state. These features can be the set of all possible observations the robot can receive during its interaction with the simulation environment, forming the robot's observation space. This space can serve as all the information in the robot's state during reinforcement learning while walking, or it can serve as only a portion of the information in the robot's state during reinforcement learning while walking.
[0133] Specifically, reinforcement learning during robot walking refers to the robot having different walking states at each moment during its walking process. Based on these states, a walking behavior (action) can be determined, which is based on the robot's current walking strategy and walking state. After the robot walks according to the determined behavior, the corresponding feedback (reward or penalty) can be calculated. Thus, the robot's state, actions, and feedback at each moment over a period of time can be collected, and the walking strategy used to determine the behavior can be optimized based on the feedback, with the goal of maximizing the reward in the feedback. Furthermore, the robot can acquire the walking state again and determine the current walking behavior based on the optimized walking strategy and walking state to recalculate the corresponding feedback. Afterward, the robot's state, actions, and feedback at each moment over a period of time based on the optimized walking strategy can be collected again to further optimize the optimized walking strategy. Therefore, the robot can train while walking, autonomously learning and optimizing its control strategy (i.e., walking strategy) to achieve stable walking in complex environments.
[0134] In this embodiment, the current moment can refer to any moment during the robot's walking process. Since the robot's behavioral state at each moment during its walking process can be understood as data at a point in time in time series analysis, the moment in this embodiment can also be understood as a time step, i.e., the number of times the robot interacts with the environment during reinforcement learning. In each time step, the robot makes a decision based on its current walking state, executing the selected action (i.e., choosing a behavior to walk) to generate the robot's next walking state. Feedback (reward or penalty) can also be calculated based on the robot's interaction with the environment during this time step. The interval between each time step (each moment) is the same. For example, if the current robot walking behavior is determined at a frequency of 50 Hz, then the interval between each time step is 0.02 seconds (s), which can be obtained from 1 / 50, that is, each time step is 0.02 seconds. Since the robot's walking behavior can be understood as a control strategy for controlling the robot's walking, the 50 Hz can also be understood as the strategy update cycle.
[0135] In this embodiment, the current walking state can refer to features that collectively describe the robot's state at the current moment. These features are the set of all possible observations the robot can receive during its interaction with the simulation environment. Specifically, the walking state can include the robot's exhibited state, such as the robot's base linear velocity, base angular velocity, projected gravity, joint angular position, and joint angular velocity.
[0136] The basic linear velocity can be used to represent the robot's linear velocity in space. It is a three-dimensional vector that includes the robot's velocity components in the x, y, and z directions. The robot's velocity components in the x, y, and z directions (i.e., the basic linear velocities in the x, y, and z directions) can be expressed as follows: and
[0137] The fundamental angular velocity can be used to represent the robot's angular velocity in space. It is also a three-dimensional vector that includes the robot's rotational velocities about the x, y, and z axes. The robot's rotational velocities about the x, y, and z axes (i.e., the fundamental angular velocities of the x, y, and z axes) can be expressed as: and
[0138] Projected gravity can be used to represent the projection of gravity onto the robot's coordinate system, helping to determine the robot's posture and tilt angle in space. It is also a three-dimensional vector, including the robot's posture projections on the x, y, and z axes. The robot's posture projections on the x, y, and z axes can be represented as follows: and
[0139] The angular position of a robot's joints can be used to represent the current angle of each joint. For ease of description, this application takes a robot with 12 drive joints as an example. The angle of each joint is a 12-dimensional vector, including the current angle of each of the 12 drive joints, which can be represented as follows: or Where i represents the i-th driven joint of the robot.
[0140] The angular velocity of a robot's joints can be used to represent the current velocity of each joint. Taking a robot with 12 driven joints as an example, the angular velocity of each joint is a 12-dimensional vector, including the current angular velocity of each of the 12 driven joints, which can be represented as follows: or Here, i represents the i-th driven joint of the robot. It is understandable that in a simulation environment built on Isaac Gym, the angular position and angular velocity of each joint of the robot can be directly obtained. Users can also select specific joints to obtain detailed information (angular position and angular velocity) for those joints.
[0141] It should be noted that the robot's basic linear velocity, basic angular velocity, projected gravity, joint angular positions, and joint angular velocities mentioned above are features included in the robot's observation space. The robot's observation space can also include features from its motion space. Taking a robot with 12 driven joints as an example, this motion space can include a 12-dimensional vector corresponding to the angular positions of each joint. That is, the robot can determine the angular positions of the 12 joints at each instant (time step). In other words, the robot's observation space can also include the action (control strategy) selected by the robot at the current time step, which is a 12-dimensional vector including the angular positions of the 12 joints. This action is generated by a reinforcement learning algorithm and used to control the robot's movement.
[0142] In this embodiment, the robot's accumulated walking time is the total time the robot has accumulated so far. For example, it is the total time from when the robot model is built (imported or created) in the simulation environment to the present, or the total time after the robot starts walking in the real environment. Since the robot starts timing when it is built in the simulation environment, the robot's accumulated walking time in the simulation environment can also be understood as the current time (current duration).
[0143] Therefore, the robot's target walking state at the current moment can be determined based on the robot's walking time and walking state.
[0144] S602. Based on the walking time and the robot's gait planning information, determine the gait information at the current moment, and use the gait information and the walking state at the current moment as the target walking state at the current moment.
[0145] In this embodiment, the robot's gait planning information refers to a pre-determined gait plan, i.e., an initial walking pattern provided by the developers, used to provide preliminary gait guidance. By using the set gait plan as input, initial gait information is provided, which can guide the robot to learn walking postures more quickly, making the reinforcement learning process more efficient. Gait information can refer to information describing the robot's walking state and can be represented in different forms, such as... Figure 4 and Figure 5The wheeled robot shown includes four mechanical legs. To better guide the robot and train its walking posture during reinforcement learning, the gait information can be a four-dimensional vector, each corresponding to the gait of one of the four mechanical legs. That is, the state space includes four-dimensional gait information, which can correspond to the gait of each of the robot's four mechanical legs. This four-dimensional gait information is used in the reinforcement learning training process to guide the robot in learning its walking posture, thereby achieving better training results. The target walking state can be composed of the machine's current walking state and gait information, which can be understood as the robot's state at a time step (the current moment) during the reinforcement learning process.
[0146] In one possible implementation, the robot may include a first set of motion parts and a second set of motion parts, and the gait planning information may include a reference step length period and a reference gait offset. The motion parts included in the first set of motion parts and the second set of motion parts are different, for example... Figure 4 and Figure 5 In the robot, the first set of moving parts may include two external mechanical legs (which may be called external legs / external feet) or two internal mechanical legs (which may be called internal legs / internal feet), and the second set of moving parts may include two external mechanical legs or two internal mechanical legs, which are not limited in this application.
[0147] The reference step length period refers to the pre-planned duration of one step the robot takes, such as 1 second per step. The reference gait offset can refer to the time interval during which the robot supports and oscillates within one reference step length period, such as the time interval between the first and second sets of moving parts, during its movement. Figure 4 and Figure 5 For wheeled and legged robots, this reference gait offset can be used to indicate when, within a reference step length period (e.g., 1 second), the bipedal (inner or outer foot) support phase and the quadrupedal support phase.
[0148] Therefore, in determining the gait information at the current moment based on the already walked time and the robot's gait planning information, the gait phase at the current moment can be determined first based on the ratio of the already walked time to the reference step length period. This gait phase can be understood as the robot's gait phase angle at the current moment. Since a robot's walking can be viewed as the robot's supporting leg and swinging leg alternating with each step in a continuous cycle, the sine and cosine values of the gait phase can be calculated using sine and cosine functions after determining the gait phase. These two values represent the gait phase, respectively. The projections onto the unit circle can be represented by these two values, which can be used to represent the first gait information corresponding to the first set of moving parts of the robot and the second gait information corresponding to the second set of moving parts.
[0149] In other words, gait information is used to describe the gait phase of a robot. For example, the calculated sine value can be used to describe the gait phase of the robot's outer foot, and the calculated cosine value can be used to describe the gait phase of the robot's inner foot.
[0150] Specifically, the first gait information can be determined based on the current gait phase and the first gait offset associated with the first set of moving parts in the reference gait offset. The second gait information is determined based on the current gait phase and the second gait offset associated with the second set of moving parts in the reference gait offset. Here, the first and second gait offsets can refer to the offsets ε added when calculating the sine and cosine values, respectively. These offsets indicate the time period during which the first and second sets of moving parts support the gait, i.e., the time period during which the moving parts act as the supporting leg, ensuring the accuracy of the gait phase and allowing the gait information to more accurately represent the gait phase. It is understood that if no offset is set for the inner or outer foot (i.e., offset ε is 0), then the first or second gait offset can be 0.
[0151] Please refer to the following: Figure 7 , Figure 7 This is a schematic diagram of the gait phase trajectory of a robot provided in an embodiment of this application, as shown below. Figure 7 As shown, since the robot's outer and inner feet alternate in support, similar to a bipedal robot, this gait phase trajectory can also be called the bipedal robot gait regulation. Figure 7 Let's take a robot's one-step time as an example and draw it. Figure 7 The sine wave in Figure 7 The solid line (shown) can be used to represent the gait phase trajectory of each step of the outer foot (one sine wave represents one step). Figure 7 cosine wave in Figure 7 (As shown by the dashed line) can be used to represent the gait phase trajectory of each step of the inner foot (one cosine wave represents one step).
[0152] like Figure 7 As shown, each step can include a two-legged support phase (as shown by the light gray square) and a four-legged support phase (as shown by the dark gray square). The two-legged support phase can be understood as the phase in which only the inner foot or the outer foot is supported, such as the inner foot supporting while the outer foot swings, or the outer foot supporting while the inner foot swings. The four-legged support phase can be understood as the phase in which both the inner foot and the outer foot are supported. Figure 7Taking the example of no offset for the outer foot (i.e., no offset for the sine wave) and an offset for the inner foot (i.e., offset for the cosine wave) as an example, that is, as... Figure 7 As shown, in the reference gait offset, the first gait offset is 0, while the second gait offset is not 0. (From...) Figure 7 It can be seen that when the sine value is greater than or equal to zero, it indicates that the outer foot is in a standing position, and when the cosine value is greater than or equal to zero, it indicates that the inner foot is in a standing position. Furthermore, when both the sine and cosine values are less than or equal to zero, it indicates that both the outer and inner feet are in a standing position.
[0153] For example, taking a reference step size period of 1 second, meaning the period of a sine wave and a cosine wave is 1 second. The proportion of the four-legged support phase in the reference step size period (1 second per step) can be set to, for example, 0.3, or 30%, then the duration of the four-legged support phase is 1 × 0.3 = 0.3 seconds. Taking the first step state offset as 0 as an example, the second step state offset can be half of 0.3 multiplied by 2π and then subtracted by π / 2, i.e., 0.3 × 1 / 2 × 2π - π / 2. Here, 1 / 2 represents half of the four-legged support phase time, because the four-legged support phase is usually symmetrical; 2π is used to convert the time proportion into radians, i.e., the projection on the unit circle; and π / 2 can adjust the phase of the cosine waveform to align with the sine waveform.
[0154] Furthermore, these calculation results can be stored in a tensor as the gait information at the current moment, resulting in 4-dimensional gait information. This gait information includes first-step gait information and second-step gait information. The first-step gait information can, for example, include the sine value of the gait phase. The second gait information may include, for example, the cosine value of the gait phase: The offset ε can be 0. In this way, the robot's gait phase can be accurately obtained, providing more comprehensive and accurate gait information for reinforcement learning. This helps to optimize the robot's walking strategy and improve the robot's adaptability and stability in complex environments.
[0155] Therefore, the robot's observation space can be composed of 49 key features, including basic linear velocity (3-dimensional vector), basic angular velocity (3-dimensional vector), projected gravity (3-dimensional vector), joint angular position (12-dimensional vector), joint angular velocity (12-dimensional vector), the robot's selected action (control strategy), which is a 12-dimensional vector corresponding to the angular position of the 12 joints, and gait information (4-dimensional vector).
[0156] Furthermore, gait information and the current walking state can be used as the robot's target walking state at the current moment. That is, the target walking state can include the current baseline linear velocity, baseline angular velocity, projected gravity, joint angular positions, joint angular velocities, and gait information. In reinforcement learning, the stage of determining the target walking state can be understood as the sampling stage. After determining the target walking state at the current moment, the control strategy for the current moment can be determined to control the robot's movement.
[0157] S603. Based on the first mapping relationship between the target walking state and the control strategy, determine the target control strategy mapped by the target walking state at the current moment.
[0158] In this embodiment, the control strategy refers to the behavior determined by the robot based on the target walking state during its walking process; that is, the actions selected by the robot, which may include the angular positions of 12 joints to achieve stable walking. The first mapping relationship can be understood as the rules or methods by which the robot selects actions in a given state (target walking state). In reinforcement learning, this first mapping relationship can refer to the mapping relationship between the independent variable (target walking state) and the dependent variable (control strategy) in the policy function. This first mapping relationship can also refer to the mapping relationship indicated by the policy network, which can be a neural network model that outputs a control strategy based on the input target walking state.
[0159] Therefore, based on this first mapping relationship and the robot's target walking state at the current moment, the target control strategy mapped by the target walking state at the current moment can be determined, such as the angular positions of the 12 joints selected by the robot at the current moment.
[0160] In this embodiment, the first mapping relationship, indicated by the policy network, is used as an example for explanation. Based on this first mapping relationship, determining the target control policy mapped to the target walking state at the current moment can refer to inputting the target walking state at the current moment into the policy network in the reinforcement learning model to obtain the target control policy output by the policy network. This policy network can be a policy network within a custom reinforcement learning framework (policy structure) designed by the developers. In other words, during the robot's walking process, it can interact with the environment to obtain the target walking state and generate corresponding actions (control policies) based on the policy network in the custom reinforcement learning framework.
[0161] In one possible implementation, this embodiment of the application uses an actor and critic structure as an example of a policy structure. The policy network in this reinforcement learning model is the "actor," which selects actions based on the current state. Its goal is to learn a policy to enable the robot to walk. The reinforcement learning model also includes a value network (Critic), which acts as the "critic" and evaluates the effectiveness of the actor's actions. Its goal is to accurately predict future rewards to guide the actor's decisions. It is understood that the process of determining the control policy based on the target walking state involves only the application of the policy network.
[0162] In some embodiments, both the policy network and the value network may include multiple hidden layers (such as three hidden layers), such as three fully connected layers (Dense Layers) with dimensions of 512, 256 and 128 respectively. The policy network and the value network also deploy activation functions, which may be Exponential Linear Units (ELUs) to introduce nonlinear characteristics.
[0163] In one possible implementation, the first mapping relationship (such as a policy network) can be determined based on the target walking state at multiple time points prior to the current time, which can be obtained through reinforcement learning. Specifically, based on the target walking state at each time point prior to the current time point, the state feedback at each time point during the reinforcement learning process can be determined. Then, based on the second mapping relationship between the target walking state and the state value, the state value mapped to the target walking state at each time point can be determined. Subsequently, the first mapping relationship can be determined based on the state values at each time point.
[0164] In reinforcement learning, state feedback at each time step refers to the sum of rewards and penalties during the learning process. It is used to evaluate the merits of the action chosen at that time step within the given state. Through continuous trial and error, the robot learns to select behavioral strategies that yield high rewards in different states, such as achieving stable walking in various flat terrain scenarios. This state feedback provides immediate feedback on the robot's behavior and can be determined based on predefined feedback rules and the target walking state. The state value refers to the expected reward obtained by using the chosen action at each time step within the target walking state. It can be used to update the first mapping relationship (such as the model parameters of the policy network) to optimize the behavioral strategy.
[0165] In reinforcement learning, this second mapping relationship can refer to the mapping relationship between the independent variable (target walking state) and the dependent variable (state value) in the value function. Alternatively, it can refer to the mapping relationship indicated by a value network (such as the value network (Critic) mentioned above), which can be a neural network model that outputs a predicted state value based on the input target walking state. Therefore, based on this second mapping relationship and the robot's target walking state at each time step, the state value mapped to that target walking state at each time step can be determined.
[0166] Furthermore, based on the state feedback and state value at each time step, the dominance function at each time step in the reinforcement learning process can be determined. The dominance function measures the relative advantage of performing a certain action (control policy) in a given state (target walking state) compared to the average. Finally, the historical mapping relationship in the reinforcement learning process can be updated based on the dominance function at each time step to obtain the first mapping relationship. In other words, before the current time step, the control policy used by the robot is based on the target walking state and the historical mapping relationship. The dominance function at each time step can update this historical mapping relationship, such as updating the model parameters of the policy network, thereby obtaining an optimized policy network. The optimized policy network is used to indicate the first mapping relationship.
[0167] Specifically, taking the historical mapping relationship as the mapping relationship indicated by the historical policy network as an example, in the process of updating the historical mapping relationship of the reinforcement learning based on the advantage function at each time step to obtain the first mapping relationship, policy loss data (actor loss) can be determined first based on the advantage function at each time step, the target walking state at each time step, and the control policy at each time step. Then, the model parameters of the historical policy network are updated based on the policy loss data to obtain the policy network used to indicate the first mapping relationship. That is, policy loss data can be constructed based on the state (target walking state), action (control policy), and reward (state feedback) at multiple time steps before the current time step to construct policy loss data, and the historical policy network is trained (model parameters are adjusted) based on the policy loss data to obtain the policy network.
[0168] Furthermore, the model parameters of the value network in the reinforcement learning model can be updated. Specifically, this can be done by first determining the critical loss data based on the advantage function and state value at each time step, and then updating the model parameters of the historical value network based on the critical loss data to obtain the final value network. During the process of updating the reinforcement learning model parameters, the model parameters of the policy network can be updated based on the policy loss data, and the model parameters of the value network can be updated based on the critical loss data. The second mapping relationship for determining the state value during the determination of the policy loss data and the critical loss data can be indicated by the historical value network before the model parameters are updated.
[0169] Reinforcement learning includes various algorithms, and the PPO algorithm is one of the advanced reinforcement learning algorithms, widely used due to its superior performance in handling continuous action spaces and stability. Therefore, in the training environment of this application embodiment, the PPO algorithm is used as an example to explain how to optimize the robot's walking strategy. During the reinforcement learning (training) process, the PPO algorithm continuously updates the parameters of the policy network and value network through interaction with the environment. The specific process includes a sampling phase (i.e., determining the target walking state), calculating the dominance function, updating the policy network, and updating the value network. Thus, through continuous iteration, the PPO algorithm can effectively optimize the walking strategy of the wheeled robot, enabling it to walk stably in environments such as flat ground, which is beneficial to improving training efficiency and enhancing the robot's adaptability in complex environments.
[0170] In the PPO algorithm, based on the advantage function, the target walking state, and the control strategy at each time step, the determined policy loss data can be shown in Equation 1. Based on the advantage function and the state value at each time step, the determined value loss data can be shown in Equation 2.
[0171]
[0172] In Formula 1, L CLIP(θ) E represents the policy loss data in the PPO algorithm. t The expected value can be obtained by summing the values at each time step (t) and then averaging them. `min` represents the minimum value. (r) t (θ) represents the strategy ratio, which can be obtained by comparing the log probability of taking the action at each time (t) under the old strategy with the target walking state at that time, and the log probability of taking the action at that time under the current strategy. It can be used to measure the degree of change of the new strategy relative to the old strategy. Let represent the dominance function at time t (the t-th time step), clip represent the clipping function, ∈ represent the clipping range, and is a constant, for example, set to 0.2.
[0173] L VF (φ)=E t [(V φ (st)-R t ) 2 ] Formula 2
[0174] In Formula 2, L VF (φ) represents the value loss data in the PPO algorithm, E t The expected value, V, can be obtained by summing the values at each time step (t) and then averaging them. φ (st) represents the state value estimated by the value network based on the target walking state at each time step. R t The present value and expected return of all rewards (state feedback) from the first time point to the last time point can be calculated using the advantage function. To improve the computational efficiency of the advantage function, a generalized advantage estimation (GAE) method can be used to approximate its estimation.
[0175] Therefore, the simulation environment built on IsaacGym (such as...) Figure 3 Using a flat terrain environment as the training environment, a custom reinforcement learning framework is employed for training. During the robot's movement, it autonomously explores and updates the policy network based on policy loss data and the value network based on value loss data—essentially, it performs autonomous optimization using reinforcement learning algorithms. Through continuous policy optimization, the robot gradually learns to walk stably in environments such as flat terrain, meeting the needs of practical applications.
[0176] The status feedback involved in the embodiments of this application will be described in detail below.
[0177] In reinforcement learning applications, state feedback can be understood as rewards or reward functions. The design of this state feedback is crucial, directly impacting the algorithm's learning direction and effectiveness. It defines the feedback the robot should receive after each step in an environment (such as a simulation environment). The simulation environment generates a scalar signal as reward feedback based on the current state and the robot's actions (control strategy), used to measure the quality of the robot's actions in that round. Each round refers to each moment (time step).
[0178] It's important to note that if the reward design is too sparse—for example, only rewarding the step of successfully scooping up the target object with a spoon or only rewarding the step of successfully clamping the target object into the target container with tongs—the agent may fail to learn an effective strategy even after millions or tens of millions of explorations due to the large environmental space (too many rounds). Therefore, designing reasonable state feedback to guide the robot step by step in completing tasks such as walking on flat ground is crucial.
[0179] In this application embodiment, state feedback can be divided into three categories: gait task feedback, motion stability and trajectory tracking feedback, and energy efficiency feedback. Gait task feedback refers to state feedback designed to determine the robot's contact with the ground based on gait information. This is because ensuring contact between the mechanical wheels and mechanical feet and the ground is crucial in the walking control of wheeled robots. Motion stability and trajectory tracking feedback refers to state feedback designed to ensure that the robot maintains a stable motion state during walking and tracks the predetermined trajectory as accurately as possible. Energy efficiency feedback refers to state feedback designed to ensure that the robot uses energy as efficiently as possible during walking, while maintaining the health and stability of mechanical components.
[0180] Specifically, the target walking state at each time point prior to the current time can include gait information associated with the robot's moving parts, such as gait information associated with the robot's lateral foot. Gait information associated with the inner foot Based on the target walking state at various time points, the specific process for determining the state feedback at each time point during reinforcement learning can be as follows: First, based on a predefined gait information masking rule, the gait information associated with the moving parts is converted into a gait mask, resulting in the gait mask for each time point. Next, the gait indication information of the moving parts at each time point is obtained; this gait indication information indicates whether the moving parts are in a specific gait. Finally, based on the degree of matching between the gait mask and the gait indication information at each time point, the state feedback at each time point is determined.
[0181] The gait information masking rule can be a rule for converting gait information associated with moving parts into a gait mask. Specifically, a default Boolean mask can be initialized first. The Boolean mask value includes 0 and 1, with each bit corresponding to a mechanical foot (or wheel) of the robot. For example, 0 indicates that the robot's mechanical foot (or wheel) is suspended in the air, and 1 indicates that the robot's mechanical foot (or wheel) is in contact with the ground (i.e., on the ground). Figure 4 and Figure 5Taking the quadruped robot shown as an example, the initial Boolean mask can be, for example, {0,0,0,0}, which by default represents all values being floating. Furthermore, by judging the relationship between the sine and cosine values in this gait information and 0 (e.g., whether the sine and cosine values are greater than or equal to 0), it is determined whether the mechanical leg (or mechanical wheel) is in a standing state, i.e., whether it is in a standing position. Figure 7 The support phase is shown. Since this Boolean mask indicates whether the mechanical foot (or wheel) is in a standing state, it can also be called the standing mask. For example, when the sine value... When the value is greater than or equal to zero, it indicates that the outer foot (or the mechanical wheel connected to the outer foot) is in a standing state; when the cosine value is greater than or equal to zero, it indicates that the outer foot is in a standing state. When the value is greater than or equal to zero, it indicates that the inner foot (or the mechanical wheel connected to the inner foot) is in a standing position.
[0182] Furthermore, the gait information masking rule includes a specific conditional judgment to further optimize the standing mask. Specifically, if the sine and cosine values have the same sign (i.e., both greater than or equal to 0 or both less than or equal to 0), it indicates that the current gait phase is in a specific stage, meaning all mechanical feet (or wheels) are supporting legs, i.e., all mechanical feet (or wheels) are in a standing state. In this case, the value of the standing mask is set to 1, i.e., {1,1,1,1}. Therefore, based on this gait information masking rule, gait information can be converted into a gait mask (i.e., a standing mask), thus obtaining the gait mask at each moment. This allows for accurate determination of the standing state of each mechanical foot (or leg) in the robot based on gait information, providing more comprehensive and accurate gait information for reinforcement learning algorithms. This helps optimize the robot's walking strategy and improve its adaptability and stability in complex environments.
[0183] The moving parts may include the robot's mechanical feet and mechanical wheels, and the gait information associated with the moving parts may include the gait information of the mechanical feet and the gait information of the mechanical wheels. It is understood that the gait information is the same for each mechanical foot and mechanical wheel, differing only in whether it is an outside foot or an inside foot. For example, the gait information applied to a pair of mechanical wheels for conversion into a gait mask can be represented as follows:
[0184] The gait information applied to mechanical foot pairs and converted into a gait mask can be represented as: It should be noted that in these two representations, stance_mask wheel This refers to the gait information used by mechanical wheelsets to convert into a gait mask, stance_mask. feet This refers to the gait information that is applied to mechanical foot pairs and converted into a gait mask.
[0185] Furthermore, based on the relationship between the sine and cosine values and 0 in the gait information masking rules, and whether the signs of the sine and cosine values are the same, a gait mask containing only 0 and 1 can be obtained. For example, the gait mask for a mechanical wheel is: stance_mask whee1 ={0,1,1,1}, the gait mask of the mechanical foot is: stance_mask feet ={1,1,1,1}.
[0186] Furthermore, the gait indication information of the moving parts at various times refers to the indication information used at each time to indicate whether the robot's moving parts are in a specific gait. A specific gait refers to a gait in contact with the ground. Gait indication information can be represented by 0 and 1 to indicate contact with the ground and suspension, for example, 0 indicates suspension and 1 indicates contact with the ground.
[0187] Specifically, contact indicators for the robot's four mechanical wheels and mechanical feet can be defined, for example: Contact whee1 ={I 01 ,I 02 ,I 03 ,I 04} and Contact feet ={I 01 ,I 02 ,I 03 ,I 04}. I i It is used to indicate, such as Figure 4 and Figure 5 The i-th mechanical leg or i-th mechanical wheel of the robot shown is used. The contact indicator I is updated by detecting whether each mechanical wheel and mechanical leg is in contact with the ground. i This allows the updated contact indicator to be used as gait indication information. In the simulation environment, it's possible to directly determine whether the contact force between the mechanical wheel and the mechanical foot meets the set contact conditions, such as whether it exceeds 5 Newtons. If the contact force is greater than 5 Newtons, it's determined that the mechanical foot is in contact with the ground; otherwise, it's determined that it's not in contact. When both the mechanical foot and the mechanical wheel are detected to be in contact with the ground, the contact indicator could be, for example, a Contact indicator. wheel ={1,1,1,1} and Contact feet ={1,1,1,1}, to obtain gait indication information.
[0188] Furthermore, the gait indication information and gait mask can be compared to determine the state feedback at each time step based on the degree of matching between the gait mask and the gait indication information at each time step. Two reward functions can be defined: one to reward the mechanical wheel for not being suspended in the air, and the other to reward the mechanical foot for not being suspended in the air. The degree of matching refers to whether the values of the corresponding mechanical foot or mechanical wheel in the gait mask and gait indication information at each time step are the same. For example, if the values of the corresponding mechanical foot or mechanical wheel in the gait mask and gait indication information are the same, a positive reward is given, such as a reward value of 0.25; otherwise, a negative reward (i.e., a penalty) is given, such as -0.1. Thus, the gait indication information and the gait mask at each time step can be compared to calculate the rewards for the mechanical foot and mechanical wheel, thereby obtaining the state feedback at each time step.
[0189] In some embodiments, it can also be checked whether all mechanical wheels or mechanical feet are suspended in the air. If all mechanical wheels or all mechanical feet are suspended, an additional negative reward, such as -1.0, is given to strongly penalize this situation. Thus, the reward for mechanical wheels and mechanical feet is a combined result of contact reward and suspension penalty, thereby obtaining state feedback.
[0190] For example, gait indication information includes Contact wheel ={1,1,1,1} and Contact feet ={1,1,1,1}, the gait mask includes stance_mask wheel ={0,1,1,1} and stance_mask feet Taking {1,1,1,1} as an example, the rewards for the mechanical wheel and mechanical feet can be as follows:
[0191] r gait-wheel =sum{-0.1,0.25,0.25,0.25}=0.65
[0192] r gait-feet =sum{0.25,0.25,0.25,0.25}=1
[0193] Therefore, the contact reward at each moment can be based on r gait-wheel and r gait-feetThe sum is calculated by adding the rewards from all mechanical wheels and mechanical feet. Optionally, the sum of the rewards from all mechanical wheels and mechanical feet can be multiplied by a coefficient, such as 1.2, to obtain the contact reward. It should be noted that if all mechanical wheels or all mechanical feet are suspended in the air, the contact reward obtained after multiplying by the coefficient can be reduced by 1.0; if all mechanical wheels and all mechanical feet are suspended in the air, the contact reward obtained after multiplying by the coefficient can be reduced by 2.0, thus obtaining the final contact reward.
[0194] Therefore, the contact reward at each moment can be used as the state feedback at that moment, or the contact reward can be the sum of the contact reward and other feedback (reward or penalty) to determine the contact reward at that moment. Training based on this contact reward can ensure that the robot maintains stable ground contact during walking, achieves coordinated gait, and avoids the dangerous situation of completely losing ground contact, thereby improving walking stability and safety. Based on state feedback, the effectiveness of guiding the robot to achieve stability in complex environments can be improved to a certain extent.
[0195] In one possible implementation, the target walking state at each time step includes the robot's linear velocity in the target direction and in other directions perpendicular to the target direction. In determining the state feedback at each time step during reinforcement learning based on the target walking state at multiple time steps, specifically, the linear velocity reward at each time step can be determined based on the linear velocity in the target direction, a set linear velocity in the target direction, and a reward parameter associated with the target direction. Then, the linear velocity penalty at each time step is determined based on the linear velocity in other directions and penalty parameters associated with those directions. Finally, the state feedback at each time step is determined based on the sum of the linear velocity rewards and penalties at each time step.
[0196] Wherein, the linear velocity in the target direction is the basic linear velocity in the x-axis direction. Linear velocity in other directions can refer to the linear velocity in the y-axis and z-axis directions. and Both the y-axis and z-axis are perpendicular to the x-axis. The set linear velocity in the target direction can refer to a predetermined linear velocity value, such as 0.3 m / s. Therefore, a reward function can be set to encourage the robot to maintain a stable forward speed based on how close its base linear velocity in the x-axis direction is to this set linear velocity. Specifically, the calculation method for the linear velocity reward is shown in Formula 3:
[0197]
[0198] In formula 3, r lin_x_tracking This represents the linear velocity reward at a given moment, and max represents the maximum value function. This represents the robot's basic linear velocity in the x-axis direction. This represents the robot's set linear velocity in the target direction (i.e., the x-axis direction), such as 0.3 m / s. Formula 3 uses... Taking the reward parameter associated with the target direction (i.e., the x-axis direction) as an example, this explanation is based on the fact that the reward parameter can also be other values, and this application does not limit this. Therefore, the robot's base linear velocity in the x-axis direction can be compared with a set linear velocity, and then multiplied by the reward parameter. Receive linear velocity bonus.
[0199] Furthermore, by setting another reward function to penalize the robot's linear velocity in other directions (y-axis and z-axis), the robot can be encouraged to maintain straight-line movement and reduce unnecessary lateral and vertical motion. Specifically, the calculation method for the linear velocity penalty can be found in Equation 4:
[0200]
[0201] In formula 4, r lin_yz_tracking This represents the linear velocity penalty at a given moment. and Formula 4 uses -5 as an example to illustrate the linear velocities of the robot in the y-axis and z-axis directions, i.e., the linear velocities in other directions. This example uses a penalty parameter of -5 associated with the other directions (y-axis and z-axis). This penalty parameter can also be other values, and this application does not limit its application to such values. In other words, it calculates the linear velocities of the robot in the y and z directions. and The sum of squares, multiplied by a negative coefficient such as -5, yields the linear velocity penalty. Understandably, the computer's summation of squares ensures that the summation result is positive.
[0202] Furthermore, the sum of the linear velocity reward and the linear velocity penalty at each time step can be determined as the state feedback at that time step. In some embodiments, the sum of the linear velocity reward, the linear velocity penalty at each time step, and the contact reward calculated based on the gait information at each time step can be determined as the state feedback at each time step.
[0203] In one possible implementation, the target walking state at each time step includes the robot's angular velocity in a specified direction and the robot's pose projection in a specific direction. During the process of determining the state feedback at each time step in the reinforcement learning process based on the target walking state at each of the plurality of time steps, an angular velocity penalty can also be determined at each time step based on the angular velocity in the specified direction and a penalty parameter associated with the angular velocity, and an attitude penalty can be determined at each time step based on the pose projection in the specific direction and a penalty parameter associated with the attitude projection. Then, the state feedback at each time step is determined based on the angular velocity penalty and the attitude penalty at each time step.
[0204] The angular velocity in the specified direction can include the robot's rotational velocities about the x, y, and z axes (i.e., the base angular velocities of the x, y, and z axes). and An attitude projection in a specific direction can include the robot's attitude projections on the x-axis and y-axis, respectively. and Therefore, a reward function can be set to penalize the robot's angular velocity, encouraging it to maintain a stable rotational posture and reducing unnecessary rotational motion. Specifically, the calculation method for the angular velocity penalty can be found in Formula 5:
[0205]
[0206] In formula 5, r ang_base This represents the angular velocity penalty at a given moment. and This represents the robot's angular velocity in a specified direction. Formula 5 uses -0.02 as an example, with the penalty parameter associated with the angular velocity as an example. This penalty parameter can also be other values, and this application does not limit it. In other words, it calculates the robot's rotational velocities around the x, y, and z axes. and The sum of squares, multiplied by a negative coefficient such as -0.02, yields the angular velocity penalty. It's understandable that using the sum of squares ensures the sum is positive.
[0207] Furthermore, by setting another reward function, the degree of deviation from the robot's posture can be penalized, ensuring that the robot can maintain a stable posture and reducing the risk of tilting and rolling. Specifically, the calculation method for the linear velocity penalty can be found in Formula 6:
[0208]
[0209] In formula 6, r orientation This indicates a posture penalty at a particular moment. and This represents the robot's pose projection in a specific direction, specifically the robot's pose projection along the x-axis and y-axis. In other words, it involves calculating the robot's pose projection along the x-axis and y-axis. and The sum of squares, multiplied by a negative coefficient such as -1.0, yields the attitude penalty. Understandably, calculating the sum of squares ensures that the summation result is positive.
[0210] Furthermore, the sum of the angular velocity penalty and the attitude penalty at each time step can be determined as the state feedback at that time step. In some embodiments, the sum of the linear velocity reward, linear velocity penalty, angular velocity penalty, and attitude penalty at each time step can be determined as the motion stability and trajectory tracking feedback, which can be used as the state feedback at each time step.
[0211] In some embodiments, the sum of linear velocity reward, linear velocity penalty, angular velocity penalty, posture penalty, and contact reward calculated based on gait information at each time step can be determined as the state feedback at each time step.
[0212] In one possible implementation, the target walking state at each time step includes the angular positions of each joint in a robot's multiple joints, such as the angular positions of each of the 12 driven joints. In determining the state feedback at each time step during reinforcement learning based on the target walking state at multiple time steps, specifically, the joint position penalty at each time step can be determined based on the angular positions of each joint, the corresponding angular position constraints for each joint, and the penalty parameters associated with the angular positions. Then, based on the sum of the joint position penalties for multiple joints at each time step, the joint angular position penalty at each time step is determined, and based on the joint angular position penalties at each time step, the state feedback at each time step is determined.
[0213] The angular position limit values for each joint refer to the positional limitations of each joint of the robot, with each joint having its own maximum value. and minimum value i represents the i-th joint. The joint position penalty for each joint at each time step refers to the penalty imposed on each joint at each time step. The sum of the joint position penalties for all joints at each time step is determined as the joint angle position penalty for that time step. Based on this penalty, the state feedback for each time step can be determined. Therefore, a reward function can be set to penalize the robot for joint angle positions exceeding limits during walking, ensuring that the joint positions remain within a safe range and avoiding mechanical damage. Specifically, the calculation method for the joint angle position penalty is shown in Formula 7:
[0214]
[0215] In Formula 7, r dof_pos_limit This represents the penalty for the joint angle position at a given moment. This represents the angular position of the i-th joint during the robot's walking process. This represents the minimum angular position constraint value corresponding to the i-th joint. This represents the maximum angular position limit value corresponding to the i-th joint. `sum` represents the summation function, and `clip` represents the clipping function. This means that if the value inside the parentheses is greater than 0, the result is 0; otherwise, the result is the value inside the parentheses. This means that if the value inside the parentheses is less than 0, the result is 0; otherwise, the result is the value inside the parentheses. Formula 7 uses 0.95 to calculate the angular position limit value corresponding to each joint as an example, which means that the angular position limit value can be 95% of the maximum and minimum values of the corresponding angular position. This represents the joint position penalty corresponding to the i-th joint at that moment. The joint position penalties corresponding to multiple (e.g., 12) joints at that moment are summed to obtain the joint angle position penalty.
[0216] Formula 7 uses a penalty parameter of -10 associated with the angular position as an example for explanation. This penalty parameter can also be other values, and this application does not limit its application to such values. In other words, it calculates that all joints of the robot exceed the joint position limits individually. The condition is then multiplied by a negative coefficient, such as -10.0, to obtain the joint angle position penalty.
[0217] In some embodiments, the joint angle position penalty at each moment can be determined as the state feedback at each moment. Alternatively, the sum of the linear velocity reward, linear velocity penalty, angular velocity penalty, posture penalty, contact reward calculated based on gait information at each moment, and joint angle position penalty at each moment can be determined as the state feedback at each moment.
[0218] In one possible implementation, the target walking state at each moment includes the angular velocities of each joint in a robot's multiple joints, such as the angular velocities of each of the 12 driven joints. In determining the state feedback at each moment during reinforcement learning based on the target walking state at multiple moments, specifically, the acceleration corresponding to each joint at each moment can be determined based on the angular velocities of each joint and the angular velocities of each joint at the previous moment. Then, based on the accelerations corresponding to multiple joints at each moment, and the penalty parameters associated with the accelerations, the joint acceleration penalty at each moment can be determined, and the state feedback at each moment can be determined based on the joint acceleration penalties at each moment.
[0219] The acceleration of each joint at each moment can be calculated based on the ratio of the angular velocity of each joint and the angular velocity of each joint at the previous moment to the duration between two adjacent moments (i.e., a time step, which is 0.02 seconds for 50Hz). Therefore, another reward function can be set to penalize the robot's joint acceleration during walking, encouraging the robot to maintain smooth joint movement and reduce mechanical wear. Specifically, the calculation method for joint acceleration penalty can be found in Formula 8.
[0220]
[0221] In Formula 8, r dof_acc This represents the joint acceleration penalty at a given moment. This represents the angular velocity of the i-th joint at time t. Let Δt represent the angular velocity of the i-th joint at time t-1 (the previous time step), where Δt can represent the duration between two time steps, such as 0.02 seconds. The fractional term represents the acceleration of the i-th joint at time t. `sum` is a summation function, representing the sum of the squares of the accelerations of multiple joints at each time step. Formula 8 uses a penalty parameter associated with acceleration of -0.0000003 as an example; this penalty parameter can also be other values, which are not limited in this application. In other words, it calculates the angular velocities of all joints of the robot at each time step. angular velocity compared to the previous moment The acceleration is calculated and multiplied by a negative coefficient, such as -0.0000003, to obtain the joint acceleration penalty.
[0222] In some embodiments, the joint acceleration penalty at each moment can be determined as the state feedback at each moment. Alternatively, the linear velocity reward, linear velocity penalty, angular velocity penalty, attitude penalty, contact reward, joint angle position penalty, and the sum of the joint acceleration penalties at each moment can be determined as the state feedback at each moment.
[0223] In one possible implementation, energy efficiency feedback may include joint acceleration penalties and joint angle position penalties at various time points, as well as joint torque penalties. The joint torque penalty refers to the penalty imposed on each driven joint of the robot for the torque τ it uses. i This encourages robots to minimize energy consumption during movement. Joint torque refers to the torque exerted on the joints after each control strategy is determined, used to adjust the joint angle and position. Specifically, the joint torque of each joint of the robot at each moment can be obtained. Then, based on the maximum joint torque value of each joint, the joint torque of each joint at each moment, and the penalty parameters associated with the joint torque, the joint torque penalty at each moment can be determined.
[0224] The joint torques of each joint of the robot at each moment can be obtained directly from the simulation environment or through corresponding sensors. The maximum joint torque for each joint refers to the maximum torque of the joint. Therefore, a reward function can be set to penalize the torque τ used by each driven joint. i This allows the robot to be penalized for excessive energy consumption during walking. Specifically, the calculation method for joint acceleration penalty can be found in Formula 9:
[0225]
[0226] In Formula 9, r torque_limit τ represents the joint torque penalty at a given moment. i This represents the joint torque of the i-th joint at that moment. Let represent the maximum joint torque corresponding to the i-th joint, and let abs represent the absolute value of the ratio between the joint torque and the maximum joint torque, as well as the absolute value of the difference between the joint torque and 95% of the maximum joint torque. sum is the summation function, representing the absolute values. and The sum, Equation 9, is explained using penalty parameters of -0.8 and -0.05 related to joint torque as examples. That is, on the one hand, the absolute value of the torque used by the robot is calculated, and then normalized (i.e.,...) After that, multiply by a negative coefficient, such as -0.8, to obtain the first part of the joint torque penalty. On the other hand, it calculates whether the robot exceeds the torque limit. In the case of exceeding torque limits, the torque penalty is calculated by multiplying the result by a negative coefficient, such as -0.005, to obtain another portion of the joint torque penalty. The sum of the two portions is taken as the joint torque penalty. This allows the robot to avoid excessive torque usage during movement by penalizing instances where torque exceeds limits, thus protecting the health of mechanical components.
[0227] In some embodiments, the state feedback at each time step during reinforcement learning is determined based on the target walking state at each time step across multiple time steps. Specifically, the state feedback at each time step can be determined based on the target walking state at each time step and the joint torque penalty at each time step. For example, the joint acceleration penalty at each time step can be determined as the state feedback at each time step. Alternatively, the sum of the linear velocity reward, linear velocity penalty, angular velocity penalty, posture penalty, contact reward, joint angle position penalty, joint acceleration penalty, and joint torque penalty at each time step can be determined as the state feedback at each time step.
[0228] In one possible implementation, energy efficiency feedback may further include a policy smoothing penalty to encourage the smoothness of the robot's chosen actions (control policies, such as a 12-dimensional vector in the action space) to reduce unnecessary jarring movements. Specifically, the control policies at each time step can be obtained, and then differential calculations can be performed based on the control policies at each time step and the control policies at adjacent time steps to obtain the policy change rate at each time step. Then, based on the control policies at each time step, the policy change rate at each time step, and the penalty parameter associated with the policy change rate, the policy smoothness penalty at each time step is determined.
[0229] The robot's control strategies at each moment can be stored after being determined. During training, the stored data can be retrieved to calculate the policy smoothness penalty. Adjacent moments refer to the time intervals between each moment and the two moments preceding it. The difference calculation includes first-order and second-order differences. The first-order difference is the difference between two consecutive values in the sequence, such as the difference between the control strategy at each moment and the time interval preceding it. The second-order difference is the difference between three consecutive values in the value sequence, such as the difference between the control strategy at each moment and the two moments preceding it. Therefore, a reward function can be defined to penalize the smoothness of control strategy changes. Specifically, the calculation method for the policy smoothness penalty is shown in Formula 10.
[0230] r action_smoothness =-0.02*((a) t-1 -a t ) 2 +(a t +at-2 -2a t-1 ) 2 +(a t ) 2 ) Formula 10
[0231] In Formula 10, r action_smoothness Let a represent the policy smoothness penalty at a given moment. t This represents the action (control strategy) at that moment, a t-1 This represents the action (control policy) of the previous time step adjacent to this time step, a t-2 This represents the actions (control strategy) at two points in time at this given moment. t-1 -a t This represents the difference in control strategy between the current time step and the previous time step, calculated using first-order difference. t +a t-2 -2a t-1 This represents the difference in control strategy between the current time point and the previous time point, as well as the time point before that, calculated using second-order difference (a). t -a t-1 )-(a t-1 -a t-2 The sum of squares is used to ensure that all terms are positive. Formula 10 is explained using an example where the penalty parameter associated with the rate of change of the policy is -0.02. That is, by calculating the action (control policy) at time a... t At the previous moment a t-1 and the moment before that a t-2 The smoothness of the change is calculated and multiplied by a negative coefficient, such as -0.02, to obtain the motion smoothness penalty.
[0232] In some embodiments, the determination of state feedback at each time step during reinforcement learning is based on the target walking state at each time step across multiple time steps. Specifically, the state feedback at each time step can be determined based on the target walking state at each time step and the policy smoothness penalty at each time step. For example, the joint acceleration penalty at each time step can be determined as the state feedback at each time step. Alternatively, the sum of the linear velocity reward, linear velocity penalty, angular velocity penalty, posture penalty, contact reward, joint angle position penalty, joint acceleration penalty, joint torque penalty, and policy smoothness penalty at each time step can be determined as the state feedback at each time step.
[0233] Therefore, by training mapping relationships (such as policy networks and value networks) based on these comprehensive reward and punishment mechanisms, robots can be effectively guided to achieve efficient, stable, and safe walking in complex environments. These state feedbacks encourage robots to reduce energy consumption and protect mechanical components through rewards, and ensure their mechanical health and stability during walking by punishing unnecessary violent movements and exceeding limits.
[0234] S604. Control the robot to walk based on the above target control strategy.
[0235] In this embodiment, the target control strategy refers to a control strategy derived from mapping the robot's target walking state at the current moment. It can be used to indicate the target angular positions of each joint of the robot. For example, the target control strategy is a multi-dimensional vector, such as a 12-dimensional vector, corresponding to the target angular positions of the 12 driven joints of the robot. It can be understood that the robot's target walking state at the current moment includes the angular positions of each joint of the robot, which is also a multi-dimensional vector, such as a 12-dimensional vector, corresponding to the angular positions of the 12 driven joints of the robot. Controlling robot walking based on the target control strategy means adjusting the angular positions of each joint of the robot at the current moment according to the target angular positions indicated by the target control strategy, thereby achieving robot walking control.
[0236] Specifically, the angular position error of each joint can be determined based on its target angular position, its current angular position, and its default angular position. Then, proportional-derivative control is performed based on these angular position errors to determine the control torque for each joint. Finally, the joints are controlled based on their control torques to enable robot movement. The default angular position of each joint can be understood as its initial angular position; that is, the default angular position of each joint is not zero. The angular position error of each joint can be obtained by subtracting the target angular position from the current angular position and the default angular position of each joint, i.e., angular position error = target angular position - current angular position - default angular position.
[0237] The proportional-derivative (PD) control process includes proportional control and derivative control. Proportional control means the control torque of each joint is proportional to the angular position error of the corresponding joint. For example, Kp * angular position error of each joint, where Kp is the proportional coefficient, is used to calculate the control torque for proportional control. Derivative control means the control torque is proportional to the rate of change of the angular position error of each joint. In the simulation environment, the rate of change of the angular position error of each joint can be obtained and then multiplied by Kd to obtain the control torque for derivative control, where Kd is the derivative coefficient. Therefore, the sum of the control torques of each joint in the proportional control process and the control torques of the corresponding joints in the derivative control process is taken as the control torque of each joint. Furthermore, the robot's joints can be controlled based on this control torque. For example, in the simulation environment, the control torque of each joint can be sent to each joint to enable the robot to move.
[0238] In some embodiments, proportional-derivative control processing is performed based on the angular position error of each joint to determine the control torque of each joint. This can be a proportional-derivative (PD) controller in the robot, such as a low-level proportional-derivative controller (LLPDC), or other controllers. This application does not limit the specific controllers used.
[0239] In one possible implementation, the robot's joints are controlled based on their control torques. After the robot begins to move, the current angular position of each joint is acquired at set time intervals. Based on the current angular position and the default angular position of each joint, the current angular position error of each joint is determined. Then, proportional-derivative control is performed based on these errors to determine the current control torque for each joint. The joints are then controlled again based on their current control torques until a set number of control operations is reached. The robot's walking state after this set number of control operations is then identified as its walking state at the next moment.
[0240] After performing a PD control process based on the target control strategy, the robot can acquire the current angular position of each joint again at a set time interval. This set time interval can be calculated based on the set PD control frequency (e.g., 500 Hz), such as 1 / 500 = 0.002 seconds. That is, if the PD control frequency (500 Hz) is greater than the robot's control strategy update frequency (e.g., 50 Hz), then after determining a control strategy once, multiple PD control processes are required. For example, if the robot's control strategy update frequency is 50 Hz (determining the control strategy every 0.02 seconds), and the PD control frequency is 500 Hz (determining the control strategy every 0.002 seconds), then after determining a control strategy once, the robot will perform 10 more PD control processes.
[0241] During each PD control process, the current angular position of each joint can be acquired at a set time interval (e.g., 0.002 seconds), and the difference between this and the default angular position of each joint is used to obtain the current angular position error of each joint. Then, based on this current angular position error, the control torque for proportional control and the control torque for derivative control are calculated respectively. The sum of the two calculated control torques is then used as the current control torque for each joint. The joints can then be controlled again based on this current control torque, and this process is repeated until the set number of control cycles is reached, such as 10 cycles as mentioned above, at which point the next control strategy is determined. At this point, the robot's walking state after the set number of control cycles is determined as the walking state for the next moment.
[0242] Furthermore, the target walking state of the robot at the next moment can be determined based on the walking state at the next moment and the walking time of the robot at the next moment, and the control strategy of the target walking state mapping at the next moment can be determined based on the first mapping relationship to further control the robot's walking.
[0243] Please refer to the following: Figure 8 , Figure 8 This is a schematic diagram illustrating the principle of robot walking control provided in an embodiment of this application, as shown below. Figure 8 As shown, robot walking control can include a reinforcement learning model, which is a neural network and can include a policy network (actor) and a value network (critic). During robot walking control, the robot's base linear velocity (V) can be obtained. base ), fundamental angular velocity (w) base ), Projected Gravity (IMU) proj ), the angular position q of the robot's joints actThe robot's joint angular velocities and gait information, such as those calculated based on gait phase and trigonometric functions, can be obtained. This includes gait information for the outer foot (phase(sin,cos)) and the inner foot (phase(sin,cos)). This acquired data is input into the policy network of the reinforcement learning model as the robot's target walking state at the current moment, yielding the target control policy for that moment. Furthermore, a low-level proportional-differential controller can instruct the target angular positions of each joint based on the target control policy, determining the control torque τ for each joint to control it.
[0244] In reinforcement learning, the control policy and the current target walking state can be used as training data to train the policy network and value network in the reinforcement learning model, thereby updating the model parameters of the policy network and the value network. Thus, by utilizing reinforcement learning algorithms, the robot can autonomously explore and optimize its gait to adapt to different environments, such as different flat terrain environments. Furthermore, reinforcement learning methods also incorporate gait planning (i.e., determining gait information based on gait planning), which to some extent improves the stability and adaptability of the control policy, enabling the robot to walk efficiently and stably in environments such as flat terrain.
[0245] Furthermore, after training is complete, you can observe the robot's movement sequence as it walks in the environment; please refer to the following: Figure 9 , Figure 9 This is a schematic diagram of a robot walking in a simulation environment, provided in an embodiment of this application. Figure 9 As shown, in a simulation environment (simulator) built on Isaac Gym, the robot model's walking sequence in a flat environment, from left to right, is as follows: The robot first uses its outer mechanical leg as a supporting leg, then raises its inner leg and begins to swing it in front of the outer mechanical leg (i.e., in front of the standing leg). Then, the robot's inner mechanical leg and mechanical wheel simultaneously fall down. Afterward, the robot's body moves forward, using the inner mechanical leg as a supporting leg and the outer mechanical leg as a swinging leg. This action is repeated, thus enabling the robot to walk stably in a flat environment.
[0246] For a better demonstration of the robot's motion sequence, please refer to the following: Figure 10 , Figure 10 This is a schematic diagram of the robot's action sequence during walking, as provided in the embodiments of this application. Figure 10 As shown, during the robot's walking process, it first uses the outer mechanical leg as a supporting leg and the inner mechanical leg as a swinging leg. The robot lifts the inner mechanical leg and swings it in front of the outer mechanical leg, thus making contact with the ground. Figure 10The robots in the first and second rows are shown. Then, the robot switches to using its inner mechanical leg as a supporting leg and its outer mechanical leg as a swinging leg. At this point, the robot can lift its outer mechanical leg and swing it until it is in front of the inner mechanical leg, as shown... Figure 10 The robots in the second and third rows are shown. Afterward, the robot can switch back to using its outer mechanical leg as a support leg and swinging its inner mechanical leg, repeating this cycle to achieve walking.
[0247] Therefore, by combining pre-determined gait and reinforcement learning, robots can be controlled to walk efficiently and stably in complex and ever-changing environments, thereby improving the flexibility of robot walking control. Since robot walking in environments such as flat ground is a crucial indicator of a robot's environmental adaptability and an important function of service robots in human-inhabited environments, the robot walking control method provided in this application also improves the adaptability of robot control.
[0248] In some embodiments, the pre-determined gait, i.e., gait planning information, can be changed. For example, the quadrupedal support phase (i.e., the phase in which four wheels and four feet support the robot simultaneously) can be changed within the cycle of one step, thereby enhancing the robot's robustness.
[0249] In some embodiments of this application, the technical solutions first acquire the robot's current walking state and the duration of its walking. Then, based on the duration of walking and the robot's gait planning information, the robot's gait information at the current moment is determined, and the gait information and the current walking state are used as the target walking state at the current moment. Next, based on the mapping relationship between the target walking state and the control strategy, the target control strategy mapped to the target walking state at the current moment is determined. This mapping relationship is determined based on the target walking states at multiple moments prior to the current moment. Finally, the robot's walking can be controlled based on the target control strategy. Therefore, by obtaining the mapping relationship through reinforcement learning through the interaction between the robot and the environment to determine the current control strategy, the robot can learn and optimize autonomously to cope with complex and changing environments, which is beneficial to improving the flexibility and adaptability of robot walking control. Furthermore, determining the control strategy through the mapping relationship helps improve the real-time performance of robot walking. Moreover, by introducing a reference gait pattern as the initial walking pattern, the robot can be guided to learn walking postures, thereby improving the efficiency of the robot's autonomous learning.
[0250] The methods of the embodiments of this application have been described in detail above. In order to facilitate better implementation of the above solutions of the embodiments of this application, the apparatus of the embodiments of this application is provided below.
[0251] Please see Figure 11 , Figure 11This is a schematic diagram of a robot walking control device provided in an embodiment of this application. The robot walking control device 110 can be used to perform... Figure 6 The corresponding steps in the robot walking control method shown are illustrated. The robot walking control device 110 includes the following units:
[0252] The acquisition unit 1101 is used to acquire the robot's current walking status and the robot's walking time at the current moment;
[0253] The determining unit 1102 is used to determine the gait information of the robot at the current moment based on the walking time and the robot's gait planning information, and to take the gait information and the walking state at the current moment as the target walking state at the current moment;
[0254] The determining unit 1102 is used to determine the target control strategy mapped to the target walking state at the current moment based on a first mapping relationship between the target walking state and the control strategy; wherein, the first mapping relationship is determined based on the target walking states at multiple moments prior to the current moment.
[0255] The control unit 1103 is used to control the robot's movement based on the target control strategy.
[0256] In one possible implementation, the robot includes a first set of motion components and a second set of motion components; the gait planning information includes a reference stride period and a reference gait offset; the determining unit 1102 is used to determine the gait information at the current moment based on the walking time and the robot's gait planning information, specifically for:
[0257] The gait phase at the current moment is determined based on the ratio of the walking time to the reference stride period.
[0258] Based on the gait phase at the current moment and the first gait offset associated with the first set of moving parts in the reference gait offset, the first gait information is determined;
[0259] Based on the gait phase at the current moment and the second gait offset associated with the second set of motion components in the reference gait offset, the second gait information is determined; the gait information at the current moment includes the first gait information and the second gait information.
[0260] In one possible implementation, the first mapping relationship is obtained through reinforcement learning, and the determining unit 1102 is further configured to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each moment among the plurality of moments.
[0261] The determining unit 1102 is further configured to determine the state value mapped to the target walking state at each time moment based on the second mapping relationship between the target walking state and the state value.
[0262] The determining unit 1102 is further configured to determine the advantage function at each moment in the reinforcement learning process based on the state feedback at each moment and the state value at each moment;
[0263] The update unit 1104 is also used to update the historical mapping relationship of the reinforcement learning based on the advantage function at each time point to obtain the first mapping relationship.
[0264] In one possible implementation, the target walking state at each moment includes gait information associated with the robot's motion components; the determining unit 1102 is configured to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each moment among the plurality of moments, specifically for:
[0265] Based on the set gait information masking rules, the gait information associated with the moving parts is converted into a gait mask to obtain the gait mask at each moment.
[0266] The gait indication information of the moving component at each time point is obtained, and the gait indication information is used to indicate whether the moving component is in a specific gait.
[0267] Based on the degree of matching between the gait mask and the gait indication information at each time point, the state feedback at each time point is determined.
[0268] In one possible implementation, the target walking state at each moment includes the linear velocity of the robot in the target direction and other directions, wherein the other directions are perpendicular to the target direction; the determining unit 1102 is used to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each moment among the plurality of moments, specifically for:
[0269] The linear velocity reward at each time point is determined based on the linear velocity in the target direction, the set linear velocity in the target direction, and the reward parameter associated with the target direction.
[0270] The linear velocity penalty at each time moment is determined based on the linear velocity in the other directions and the penalty parameters associated with the other directions;
[0271] The state feedback at each time step is determined based on the sum of the linear velocity reward and the linear velocity penalty at each time step.
[0272] In one possible implementation, the target walking state at each moment includes the robot's angular velocity in a specified direction and the robot's pose projection in a specific direction; the determining unit 1102 is used to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each moment among the plurality of moments, specifically for:
[0273] The angular velocity penalty at each moment is determined based on the angular velocity in the specified direction and the penalty parameter associated with the angular velocity.
[0274] The attitude penalty at each time moment is determined based on the attitude projection in the specific direction and the penalty parameter associated with the attitude projection.
[0275] Based on the angular velocity penalty and the attitude penalty at each time step, the state feedback at each time step is determined.
[0276] In one possible implementation, the target walking state at each moment includes the angular position of each joint among multiple joints of the robot; the determining unit 1102 is used to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each moment among the multiple moments, specifically for:
[0277] Based on the angular position of each joint, the angular position limit value corresponding to each joint, and the penalty parameter associated with the angular position, the joint position penalty corresponding to each joint at each time is determined;
[0278] Based on the sum of the joint position penalties corresponding to the multiple joints at each time, the joint angle position penalty at each time is determined, and based on the joint angle position penalties at each time, the state feedback at each time is determined.
[0279] In one possible implementation, the target walking state at each moment includes the angular velocity of each joint among the robot's multiple joints; the determining unit 1102 is used to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each moment among the multiple moments, specifically for:
[0280] Based on the angular velocity of each joint and the angular velocity of each joint at the previous moment, determine the acceleration of each joint at each moment.
[0281] Based on the accelerations of the multiple joints at each time point and the penalty parameters associated with the accelerations, the joint acceleration penalties at each time point are determined, and based on the joint acceleration penalties at each time point, the state feedback at each time point is determined.
[0282] In one possible implementation, the acquisition unit 1101 is further configured to acquire the joint torque of each joint of the robot at each time.
[0283] The determining unit 1102 is further configured to determine the joint torque penalty at each time based on the maximum joint torque value corresponding to each joint, the joint torque of each joint at each time, and the penalty parameter associated with the joint torque.
[0284] The determining unit 1102 is configured to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each of the plurality of moments, specifically for:
[0285] Based on the target walking state at each time point and the joint torque penalty at each time point, the state feedback at each time point is determined.
[0286] In one possible implementation, the acquisition unit 1101 is further configured to acquire the control strategy at each time point;
[0287] The calculation unit 1105 is used to perform differential calculation based on the control strategy at each time and the control strategy at adjacent times to obtain the strategy change rate corresponding to each time.
[0288] The determining unit 1102 is used to determine the policy smoothness penalty at each time based on the control policy at each time, the policy change rate at each time, and the penalty parameter associated with the policy change rate.
[0289] The determining unit 1102 is configured to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each of the plurality of moments, specifically for:
[0290] Based on the target walking state at each time step and the policy smoothness penalty at each time step, the state feedback at each time step is determined.
[0291] In one possible implementation, the determining unit 1102 is configured to determine the target control strategy mapped to the current target walking state based on a first mapping relationship between the target walking state and the control strategy, specifically for:
[0292] The target walking state at the current moment is input into the policy network in the reinforcement learning model to obtain the target control policy output by the policy network; wherein, the policy network is used to indicate the first mapping relationship.
[0293] In one possible implementation, the update unit 1104 is used to update the historical mapping relationship of the reinforcement learning based on the advantage function at each time point to obtain the first mapping relationship, specifically for:
[0294] Based on the advantage function at each time point, the target walking state at each time point, and the control strategy at each time point, the strategy loss data is determined.
[0295] The model parameters of the historical policy network are updated based on the policy loss data to obtain the policy network, wherein the historical policy network is used to indicate the historical mapping relationship.
[0296] In one possible implementation, the reinforcement learning model further includes a value network, and the determining unit 1102 is used to determine value loss data based on the advantage function and the state value at each time step.
[0297] The update unit 1104 is further configured to update the model parameters of the historical value network based on the value loss data to obtain the value network; wherein the historical value network is used to indicate the second mapping relationship.
[0298] In one possible implementation, the current target walking state includes the angular positions of each joint of the robot, and the target control strategy is used to indicate the target angular positions of each joint; the control unit 1103 is used to control the robot's walking based on the target control strategy, specifically for:
[0299] Based on the target angle position of each joint, the angle position of each joint, and the default angle position of each joint, the angle position error of each joint is determined.
[0300] Proportional-derivative control processing is performed based on the angular position error of each joint to determine the control torque of each joint.
[0301] The robot moves by controlling each joint based on the control torque of each joint.
[0302] In one possible implementation, the acquisition unit 1101 is further configured to acquire the current angular position of each joint at a set time interval, and determine the current angular position error of each joint based on the current angular position of each joint and the default angular position of each joint.
[0303] The processing unit 1106 is used to perform proportional-derivative control processing based on the current angular position error of each joint to determine the current control torque of each joint.
[0304] The control unit 1103 is also used to control each joint again based on the current control torque of each joint until the number of control operations reaches a set number.
[0305] The determining unit 1102 is further configured to determine the walking state of the robot after the number of control operations reaches the set number as the walking state at the next moment.
[0306] According to one embodiment of this application, Figure 6 The steps involved in the method shown can all be derived from... Figure 11 The actions are performed by individual units in the control device for the robot's movement, as shown. For example, Figure 6 The step S601 shown is by Figure 11 The acquisition unit 1101 shown is used to execute steps S602 and S603. Figure 11 The determination unit 1102 shown is responsible for executing step 604. Figure 11 The control unit 1103 shown is used to perform this operation.
[0307] According to one embodiment of this application, Figure 11 The various units in the robot walking control device 110 shown can be individually or entirely combined into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The above-mentioned units are based on logical function division. In practical applications, the function of one unit can also be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the robot walking control device 110 may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by multiple units working together.
[0308] According to another embodiment of this application, a general-purpose computing device capable of performing operations such as those described above can be run on a general-purpose computer including processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). Figure 6 The computer program (including program code) for each step involved in the corresponding method shown, to construct such... Figure 11 The robot walking control device 110 shown herein, and the robot walking control method for implementing the embodiments of this application, are also described. The computer program may be recorded on, for example, a computer-readable storage medium, and loaded onto, via the computer-readable storage medium. Figure 1 The robot control device in the control system of the robot walking shown, and Figure 2The simulation device for the robot's walking control system is shown, and it operates within it.
[0309] Based on the description of the above-described robot walking control method embodiments, this application also discloses an electronic device; please refer to [link to relevant documentation]. Figure 12 The electronic device 120 may include at least a processor 1201, an input device 1202, an output device 1203, and a memory 1204. The processor 1201, input device 1202, output device 1203, and memory 1204 within the electronic device 120 may be connected via a bus or other means.
[0310] The aforementioned memory 1204 is a memory device in the robot walking control device 120, used to store programs and data. It is understood that the memory 1204 here can include the built-in storage medium of the robot walking control device, or it can include extended storage media supported by the robot walking control device 120. The memory 1204 provides storage space for storing the operating system of the robot walking control device 120. Furthermore, the computer program (including program code) is also stored in this storage space. It should be noted that the computer storage medium here can be a high-speed RAM memory; optionally, it can also be at least one computer storage medium located away from the aforementioned processor, which can be called a Central Processing Unit (CPU), the core and control center of the robot walking control device, used to run the computer program stored in the aforementioned memory 1204.
[0311] In one embodiment, the processor 1201 can load and execute the computer program stored in the memory 1204 to implement the corresponding steps of the method in the above-described embodiments of the robot walking control method; specifically, the processor 1201 loads and executes the computer program stored in the memory 1204 for:
[0312] Obtain the robot's current walking status and the robot's accumulated walking time;
[0313] Based on the walking time and the robot's gait planning information, the robot's gait information at the current moment is determined, and the gait information and the walking state at the current moment are taken as the target walking state at the current moment;
[0314] Based on the first mapping relationship between the target walking state and the control strategy, the target control strategy mapped to the target walking state at the current moment is determined; wherein, the first mapping relationship is determined based on the target walking states at multiple moments prior to the current moment.
[0315] The robot's movement is controlled based on the target control strategy.
[0316] In one possible implementation, the robot includes a first set of motion components and a second set of motion components; the gait planning information includes a reference stride period and a reference gait offset; the processor 1201 loads and executes a computer program stored in the memory 1204 to determine the gait information at the current moment based on the walking time and the robot's gait planning information, specifically for:
[0317] The gait phase at the current moment is determined based on the ratio of the walking time to the reference stride period.
[0318] Based on the gait phase at the current moment and the first gait offset associated with the first set of moving parts in the reference gait offset, the first gait information is determined;
[0319] Based on the gait phase at the current moment and the second gait offset associated with the second set of motion components in the reference gait offset, the second gait information is determined; the gait information at the current moment includes the first gait information and the second gait information.
[0320] In one possible implementation, the first mapping relationship is obtained through reinforcement learning, and the processor 1201 loads and executes the computer program stored in the memory 1204, and is further used for:
[0321] Based on the target walking state at each of the multiple time points, determine the state feedback at each of the time points during the reinforcement learning process;
[0322] Based on the second mapping relationship between the target walking state and the state value, the state value mapped to the target walking state at each time moment is determined.
[0323] Based on the state feedback and state value at each time step, the dominance function at each time step in the reinforcement learning process is determined.
[0324] The historical mapping relationship of the reinforcement learning is updated based on the advantage function at each time point to obtain the first mapping relationship.
[0325] In one possible implementation, the target walking state at each moment includes gait information associated with the robot's moving parts; the processor 1201 loads and executes a computer program stored in the memory 1204, used to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each moment among the plurality of moments, specifically used for:
[0326] Based on the set gait information masking rules, the gait information associated with the moving parts is converted into a gait mask to obtain the gait mask at each moment.
[0327] The gait indication information of the moving component at each time point is obtained, and the gait indication information is used to indicate whether the moving component is in a specific gait.
[0328] Based on the degree of matching between the gait mask and the gait indication information at each time point, the state feedback at each time point is determined.
[0329] In one possible implementation, the target walking state at each moment includes the linear velocity of the robot in the target direction and other directions, wherein the other directions are perpendicular to the target direction; the processor 1201 loads and executes the computer program stored in the memory 1204, which is used to determine the state feedback at each moment in the reinforcement learning process based on the target walking state at each moment among the plurality of moments, specifically for:
[0330] The linear velocity reward at each time point is determined based on the linear velocity in the target direction, the set linear velocity in the target direction, and the reward parameter associated with the target direction.
[0331] The linear velocity penalty at each time moment is determined based on the linear velocity in the other directions and the penalty parameters associated with the other directions;
[0332] The state feedback at each time step is determined based on the sum of the linear velocity reward and the linear velocity penalty at each time step.
[0333] In one possible implementation, the target walking state at each moment includes the robot's angular velocity in a specified direction and the robot's pose projection in a specific direction; the processor 1201 loads and executes a computer program stored in the memory 1204, used to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each of the plurality of moments, specifically used for:
[0334] The angular velocity penalty at each moment is determined based on the angular velocity in the specified direction and the penalty parameter associated with the angular velocity;
[0335] The attitude penalty at each time moment is determined based on the attitude projection in the specific direction and the penalty parameter associated with the attitude projection.
[0336] Based on the angular velocity penalty and the attitude penalty at each time step, the state feedback at each time step is determined.
[0337] In one possible implementation, the target walking state at each moment includes the angular position of each joint among the robot's multiple joints; the processor 1201 loads and executes the computer program stored in the memory 1204, for determining the state feedback at each moment during the reinforcement learning process based on the target walking state at each moment among the multiple moments, specifically for:
[0338] Based on the angular position of each joint, the angular position limit value corresponding to each joint, and the penalty parameter associated with the angular position, the joint position penalty corresponding to each joint at each time is determined;
[0339] Based on the sum of the joint position penalties corresponding to the plurality of joints at each time, the joint angle position penalty at each time is determined, and based on the joint angle position penalties at each time, the state feedback at each time is determined.
[0340] In one possible implementation, the target walking state at each moment includes the angular velocity of each joint among the robot's multiple joints; the processor 1201 loads and executes a computer program stored in the memory 1204, used to determine the state feedback at each moment during the reinforcement learning process based on the target walking state at each moment among the multiple moments, specifically for:
[0341] Based on the angular velocity of each joint and the angular velocity of each joint at the previous moment, determine the acceleration of each joint at each moment.
[0342] Based on the accelerations of the multiple joints at each time point and the penalty parameters associated with the accelerations, the joint acceleration penalties at each time point are determined, and based on the joint acceleration penalties at each time point, the state feedback at each time point is determined.
[0343] In one possible implementation, the processor 1201 loads and executes a computer program stored in the memory 1204, and is further configured to:
[0344] Obtain the joint torque of each joint of the robot at each time point;
[0345] Based on the maximum joint torque value corresponding to each joint, the joint torque of each joint at each time, and the penalty parameter associated with the joint torque, the joint torque penalty at each time is determined.
[0346] The processor 1201 loads and executes the computer program stored in the memory 1204, which is used to determine the state feedback at each moment in the reinforcement learning process based on the target walking state at each of the plurality of moments, specifically for:
[0347] Based on the target walking state at each time point and the joint torque penalty at each time point, the state feedback at each time point is determined.
[0348] In one possible implementation, the processor 1201 loads and executes a computer program stored in the memory 1204, and is further configured to:
[0349] Obtain the control strategies at each of the aforementioned time points;
[0350] Based on the control strategies at each time point and the control strategies at adjacent time points, differential calculation is performed to obtain the policy change rate corresponding to each time point.
[0351] Based on the control policy at each time point, the policy change rate at each time point, and the penalty parameter associated with the policy change rate, the policy smoothness penalty at each time point is determined.
[0352] The processor 1201 loads and executes the computer program stored in the memory 1204, which is used to determine the state feedback at each moment in the reinforcement learning process based on the target walking state at each of the plurality of moments, specifically for:
[0353] Based on the target walking state at each time step and the policy smoothness penalty at each time step, the state feedback at each time step is determined.
[0354] In one possible implementation, the processor 1201 loads and executes a computer program stored in the memory 1204, used to determine the target control strategy mapped to the current target walking state based on a first mapping relationship between the target walking state and the control strategy, specifically used for:
[0355] The target walking state at the current moment is input into the policy network in the reinforcement learning model to obtain the target control policy output by the policy network; wherein, the policy network is used to indicate the first mapping relationship.
[0356] In one possible implementation, the processor 1201 loads and executes a computer program stored in the memory 1204 to update the historical mapping relationship of the reinforcement learning based on the advantage function at each time point, thereby obtaining the first mapping relationship. Specifically, this is used for:
[0357] Based on the advantage function at each time point, the target walking state at each time point, and the control strategy at each time point, the strategy loss data is determined.
[0358] The model parameters of the historical policy network are updated based on the policy loss data to obtain the policy network, wherein the historical policy network is used to indicate the historical mapping relationship.
[0359] In one possible implementation, the reinforcement learning model further includes a value network, and the processor 1201 loads and executes a computer program stored in the memory 1204, and is further configured to:
[0360] Based on the advantage function and the state value at each time point, the value loss data is determined.
[0361] The model parameters of the historical value network are updated based on the value loss data to obtain the value network; wherein the historical value network is used to indicate the second mapping relationship.
[0362] In one possible implementation, the current target walking state includes the angular positions of each joint of the robot, and the target control strategy is used to indicate the target angular positions of each joint; the processor 1201 loads and executes the computer program stored in the memory 1204 to control the robot's walking based on the target control strategy, specifically for:
[0363] Based on the target angle position of each joint, the angle position of each joint, and the default angle position of each joint, the angle position error of each joint is determined.
[0364] Proportional-derivative control processing is performed based on the angular position error of each joint to determine the control torque of each joint.
[0365] The robot moves by controlling each joint based on the control torque of each joint.
[0366] In one possible implementation, the processor 1201 loads and executes a computer program stored in the memory 1204, and is further configured to:
[0367] The current angular position of each joint is obtained at a set time interval, and the current angular position error of each joint is determined based on the current angular position of each joint and the default angular position of each joint.
[0368] Proportional-derivative control is performed based on the current angular position error of each joint to determine the current control torque of each joint.
[0369] Based on the current control torque of each joint, the joints are controlled again until the set number of control operations is reached.
[0370] The walking state of the robot after the set number of control cycles is determined as the walking state at the next moment.
[0371] It should be understood that, in the embodiments of this application, the processor 1201 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0372] This application provides a computer-readable storage medium storing a computer program, which includes program instructions. When the program instructions are executed by a processor, they can perform the steps described in all the above embodiments.
[0373] This application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. When the computer instructions are executed by the processor of a computer device, they perform the methods described in all the above embodiments.
[0374] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0375] The above description discloses only one preferred embodiment of the present invention, and should not be construed as limiting the scope of the present invention. Those skilled in the art will understand that all or part of the processes of the above embodiments can be implemented, and equivalent changes made in accordance with the claims of the present invention are still within the scope of the invention.
[0376] It should also be noted that when the above embodiments of this application are applied to specific products or technologies, if it is necessary to obtain user data, the user's permission or consent must be obtained, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
Claims
1. A control method of robot walking, characterized by, The method comprises: obtaining a walking state of a robot at a current time and a walked duration of the robot; determining gait information of the robot at the current time based on the walked duration and gait planning information of the robot, and taking the gait information and the walking state at the current time as a target walking state at the current time; determining a target control strategy mapped by the target walking state at the current time based on a first mapping relationship between a target walking state and a control strategy, wherein the first mapping relationship is determined based on target walking states at multiple times before the current time; controlling the robot to walk based on the target control strategy.
2. The method of claim 1, wherein, The robot comprises a first set of motion components and a second set of motion components; the gait planning information comprises a reference step cycle and a reference gait offset; and the determination of the gait information at the current time based on the walked duration and the gait planning information of the robot comprises: determining a gait phase at the current time based on a ratio of the walked duration to the reference step cycle; determining first gait information based on the gait phase at the current time and a first gait offset in the reference gait offset associated with the first set of motion components; determining second gait information based on the gait phase at the current time and a second gait offset in the reference gait offset associated with the second set of motion components; and the gait information at the current time comprises the first gait information and the second gait information.
3. The method of claim 1, wherein, The first mapping relationship is obtained through reinforcement learning, and the method further comprises: determining state feedback at each of the multiple times in the reinforcement learning process based on a target walking state at the each of the multiple times; determining a state value mapped by the target walking state at the each of the multiple times based on a second mapping relationship between a target walking state and a state value; determining an advantage function at the each of the multiple times in the reinforcement learning process based on the state feedback at the each of the multiple times and the state value at the each of the multiple times; updating a historical mapping relationship of the reinforcement learning based on the advantage function at the each of the multiple times to obtain the first mapping relationship.
4. The method of claim 3, wherein, The target walking state at the each of the multiple times comprises gait information associated with a motion component of the robot; and the determination of the state feedback at the each of the multiple times in the reinforcement learning process based on a target walking state at the each of the multiple times comprises: converting the gait information associated with the motion component into a gait mask based on a set gait information mask rule to obtain a gait mask at the each of the multiple times; obtaining gait indication information of the motion component at the each of the multiple times, the gait indication information being used to indicate whether the motion component is in a specific gait; determining the state feedback at the each of the multiple times based on a matching degree of the gait mask at the each of the multiple times and the gait indication information at the each of the multiple times.
5. The method of claim 3, wherein, The target walking state at each time point includes linear velocities of the robot in a target direction and other directions perpendicular to the target direction; and determining the state feedback at each time point in the reinforcement learning process based on the target walking state at each time point includes: determining a linear velocity reward at each time point based on the linear velocity in the target direction, a set linear velocity in the target direction, and a reward parameter associated with the target direction; determining a linear velocity penalty at each time point based on the linear velocity in the other direction and a penalty parameter associated with the other direction; and determining the state feedback at each time point based on a sum of the linear velocity reward at each time point and the linear velocity penalty at each time point.
6. The method of claim 3, wherein, The target walking state at each time point includes an angular velocity of the robot in a specified direction and a pose projection of the robot in a specific direction; and determining the state feedback at each time point in the reinforcement learning process based on the target walking state at each time point includes: determining an angular velocity penalty at each time point based on the angular velocity in the specified direction and a penalty parameter associated with the angular velocity; determining a pose penalty at each time point based on the pose projection in the specific direction and a penalty parameter associated with the pose projection; and determining the state feedback at each time point based on the angular velocity penalty at each time point and the pose penalty at each time point.
7. The method of claim 3, wherein, The target walking state at each time point includes an angular position of each joint of a plurality of joints of the robot; and determining the state feedback at each time point in the reinforcement learning process based on the target walking state at each time point includes: determining a joint position penalty of each joint at each time point based on the angular position of the joint, an angular position limit value corresponding to the joint, and a penalty parameter associated with the angular position; and determining a joint angular position penalty at each time point based on a sum of the joint position penalties of the plurality of joints at each time point, and determining the state feedback at each time point based on the joint angular position penalty at each time point.
8. The method of claim 3, wherein, The target walking state at each time point includes an angular velocity of each joint of a plurality of joints of the robot; The target walking state at each time point includes an angular position of each joint of a plurality of joints of the robot; and determining the state feedback at each time point in the reinforcement learning process based on the target walking state at each time point includes: determining an acceleration of each joint at each time point based on the angular velocity of the joint and an angular velocity of the joint at a previous time point; determining a joint acceleration penalty at each time point based on a sum of the accelerations of the plurality of joints at each time point and a penalty parameter associated with the acceleration, and determining the state feedback at each time point based on the joint acceleration penalty at each time point.
9. The method of claim 3, wherein, The method further includes: obtaining joint torques of each joint of the robot at each time point; determine a joint torque penalty of the each time point based on the maximum joint torque of the each joint, the joint torque of the each joint at the each time point, and a penalty parameter associated with the joint torque; the determining the state feedback of the each time point in the reinforcement learning process based on the target walking state of the each time point, comprises: the determining the state feedback of the each time point based on the target walking state of the each time point and the joint torque penalty of the each time point.
10. The method of claim 3, wherein, the method further comprises: obtaining the control strategy of the each time point; performing differential calculation based on the control strategy of the each time point and the control strategy of the adjacent time point of the each time point to obtain a strategy change rate corresponding to the each time point; determining a strategy smoothness penalty of the each time point based on the control strategy of the each time point, the strategy change rate corresponding to the each time point, and a penalty parameter associated with the strategy change rate; the determining the state feedback of the each time point in the reinforcement learning process based on the target walking state of the each time point, comprises: the determining the state feedback of the each time point based on the target walking state of the each time point and the strategy smoothness penalty of the each time point.
11. The method of claim 3, wherein, the determining the target control strategy mapped by the target walking state of the current time point based on the first mapping relationship between the target walking state and the control strategy, comprises: inputting the target walking state of the current time point into a policy network in a reinforcement learning model to obtain the target control strategy output by the policy network; wherein the policy network is used to indicate the first mapping relationship.
12. The method of claim 11, wherein, the updating the historical mapping relationship of the reinforcement learning based on the advantage function of the each time point to obtain the first mapping relationship, comprises: determining policy loss data based on the advantage function of the each time point, the target walking state of the each time point, and the control strategy of the each time point; updating model parameters of a historical policy network based on the policy loss data to obtain the policy network, wherein the historical policy network is used to indicate the historical mapping relationship.
13. The method of claim 11, wherein, the reinforcement learning model further comprises a value network, and the method further comprises: determining value loss data based on the advantage function of the each time point and the state value of the each time point; updating model parameters of a historical value network based on the value loss data to obtain the value network; wherein the historical value network is used to indicate the second mapping relationship.
14. The method of claim 1, wherein, the target walking state of the current time point comprises angle positions of each joint of the robot, and the target control strategy is used to indicate target angle positions of the each joint; the controlling the robot to walk based on the target control strategy, comprises: determining an angle position error of the each joint based on the target angle position of the each joint, the angle position of the each joint, and a default angle position of the each joint; performing proportional and differential control processing based on the angle position error of the each joint to determine a control torque of the each joint; control the joints based on the control torques of the joints to make the robot walk.
15. The method of claim 14, wherein, After the control of the joints based on the control torques of the joints to make the robot walk, the method further comprises: acquire current angle positions of the joints at a set time interval, and determine current angle position errors of the joints based on the current angle positions of the joints and default angle positions of the joints; perform proportional and differential control processing based on the current angle position errors of the joints to determine current control torques of the joints; control the joints again based on the current control torques of the joints until a control number reaches a set number; confirm a walking state of the robot after the control number reaches the set number as a walking state at a next time.
16. A control device for robot walking, characterized by comprise: an acquisition unit configured to acquire a walking state of a robot at a current time and a walked time length of the robot; a determination unit configured to determine gait information of the robot at the current time based on the walked time length and gait planning information of the robot, and determine the gait information and the walking state at the current time as a target walking state at the current time; the determination unit is further configured to determine a target control strategy mapped by the target walking state at the current time based on a first mapping relationship between target walking states and control strategies, wherein the first mapping relationship is determined based on target walking states at a plurality of times before the current time; a control unit configured to control the robot to walk based on the target control strategy.
17. An electronic device, comprising: comprise: one or more processors; a memory configured to store one or more computer programs, which, when executed by the one or more processors, cause the electronic device to implement the control method of robot walking according to any one of claims 1-15.
18. A computer readable medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the control method of robot walking according to any one of claims 1-15.
19. A computer program product, characterised in that, The computer program product comprises a computer program stored in a computer readable storage medium, and the processor of the electronic device reads and executes the computer program from the computer readable storage medium, so that the electronic device executes the control method of robot walking according to any one of claims 1-15.
Citation Information
Patent Citations
Robot motion planning method and device and robot
CN116985142A
Humanoid gait control method, device and storage medium of humanoid robots
US20220184807A1