Humanoid robot foot lifting movement training method and device, electronic equipment and medium

By combining a motor position control network and a joint response controller with a fast style transfer and two-dimensional noise terrain generation method, the problem of poor environmental adaptability in humanoid robot foot-lifting movement training was solved, improving movement stability and motor control performance, and reducing damage rate.

CN119115940BActive Publication Date: 2026-03-03SHANGHAI QI ZHI INSTITUTE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-13
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing technologies for training humanoid robots to lift their feet to move suffer from high data collection costs, wasted computing resources, and poor environmental adaptability, resulting in poor movement stability and a high damage rate.

Method used

By employing a motor position control network and a joint response controller, simulated elevation terrain information is generated by acquiring the current observation sample information of the humanoid robot. The target joint position set is generated using an action network and a policy network, and the robot is controlled to perform foot lift-off movement training on the simulated elevation terrain. Combined with fast style transfer and two-dimensional noise terrain generation methods, the environmental adaptability and stability are improved.

Benefits of technology

It improves the mobility stability of humanoid robots and the performance of motor position control networks, reduces damage rates, and enhances adaptability to various environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119115940B_ABST
    Figure CN119115940B_ABST
Patent Text Reader

Abstract

This disclosure presents embodiments of a method, apparatus, electronic device, and medium for training a humanoid robot to lift its foot and move. One specific implementation of the method includes: acquiring current observation sample information; generating simulated elevation terrain information; inputting the current observation sample information into an action network to obtain a set of action probability distribution values; inputting the current observation sample information into a policy network to obtain a set of action evaluation values; generating a target joint position set based on the action probability distribution value set and the action evaluation value set; inputting the target joint position set into a joint response controller to generate a set of joint torques; controlling the humanoid robot to perform foot-lifting movement training; and ending the training of the motor position control network. This implementation, through the motor position control network and the joint response controller, controls the humanoid robot to perform foot-lifting movement training on rough ground, which can improve the movement stability of the humanoid robot and the performance and targeting of the motor position control network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to a method, apparatus, electronic device, and medium for training humanoid robots to lift their feet and move. Background Technology

[0002] Humanoid robots are robots that mimic human anatomy to perform similar movements. They possess complex leg joint structures, enabling greater flexibility and environmental adaptability, allowing for stable walking in various complex environments. With the rapid development of computer technology, training humanoid robots to lift their feet off the ground is receiving increasing attention. The common approach to training robot foot-lifting movement involves constructing a motion reference trajectory and gait database for the humanoid robot. Then, this database is used as a gait prior guidance network to learn similar gaits, thus completing the training for foot-lifting movement.

[0003] However, in practice, it has been found that when training humanoid robots to lift their feet off the ground using the above method, the following technical problems often occur: First, the need to collect motion reference trajectories and gait libraries increases the additional data collection cost and wastes storage and computing resources. Second, the strategy model using motion reference trajectories and gait libraries learns specific gaits for specific environments, resulting in poor environmental adaptability of the humanoid robot, leading to poor stability of the humanoid robot's movement and a high damage rate.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the present disclosure concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure provide methods, apparatus, electronic devices, and media for training humanoid robots to lift their feet and move, in order to solve one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide a training method for humanoid robot to lift its feet and move, comprising: acquiring current observation sample information of the humanoid robot, wherein the humanoid robot is a bipedal humanoid robot with clog-shaped feet, and the current observation sample information includes at least one of the following: joint position, joint velocity, joint angle, and ground coupling pressure information of the humanoid robot; generating simulated elevation and terrain information for the humanoid robot; inputting the current observation sample information into an action network included in a motor position control network to obtain a set of action probability distribution values; and inputting the current observation sample information into the motor position control network. The network includes a policy network, which obtains the action evaluation value set corresponding to the action probability distribution value set; based on the action probability distribution value set and the action evaluation value set, a target joint position set for the humanoid robot is generated; the target joint position set is input to a joint response controller to generate a joint torque set for the humanoid robot; based on the joint torque set, the humanoid robot is controlled to perform foot lift-off movement training on the simulated elevation terrain corresponding to the simulated elevation terrain information; in response to determining that the action probability distribution value set and the action evaluation value set satisfy a preset loss condition, the training of the motor position control network is terminated.

[0008] Secondly, some embodiments of this disclosure provide a humanoid robot foot-lifting movement training device, comprising: an acquisition unit configured to acquire current observation sample information of the humanoid robot, wherein the humanoid robot is a bipedal humanoid robot with clog-shaped feet, and the current observation sample information includes at least one of the following: joint position, joint velocity, joint angle, and ground coupling pressure information of the humanoid robot's feet; a first generation unit configured to generate simulated elevation terrain information for the humanoid robot; a first input unit configured to input the current observation sample information into an action network included in a motor position control network to obtain a set of action probability distribution values; and a second input unit configured to input the current observation sample information into a motor position control network. The network includes a policy network that obtains the action evaluation value set corresponding to the action probability distribution value set; a second generation unit configured to generate the target joint position set of the humanoid robot based on the action probability distribution value set and the action evaluation value set; a third input unit configured to input the target joint position set to a joint response controller to generate a joint torque set for the humanoid robot; a control unit configured to control the humanoid robot to perform foot lift-off movement training on the simulated elevation terrain corresponding to the simulated elevation terrain information based on the joint torque set; and a termination unit configured to terminate the training of the motor position control network in response to determining that the action probability distribution value set and the action evaluation value set meet a preset loss condition.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, such that when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method as described in any of the implementations of the first aspect.

[0011] The above embodiments of this disclosure have the following beneficial effects: The humanoid robot foot-lifting movement training method of some embodiments of this disclosure, through a motor position control network and a joint response controller, controls the humanoid robot to perform foot-lifting training on rough ground, which can improve the humanoid robot's movement stability, the performance of the motor position control network, and its generalization ability. Specifically, the reasons for the poor movement stability and high damage rate of the humanoid robot are: the need to collect motion reference trajectories and gait libraries increases additional data collection costs and wastes storage and computing resources; and the strategy model using motion reference trajectories and gait libraries learns specific gaits for specific environments, resulting in poor environmental adaptability for the humanoid robot, leading to poor movement stability and a high damage rate. Based on this, the humanoid robot foot-lifting movement training method of some embodiments of this disclosure can first obtain the current observation sample information of the humanoid robot, wherein the humanoid robot is a bipedal humanoid robot with clog-shaped feet, and the current observation sample information includes at least one of the following: the humanoid robot's joint position, joint velocity, joint angle, and ground coupling pressure information on both feet. Here, the current observation sample information is used for the generation of subsequent model elevation and terrain information and for training the humanoid robot's foot-lift-off-ground movement control. Next, simulated elevation and terrain information is generated for the humanoid robot. This simulated elevation and terrain information can characterize any environmental information used for the humanoid robot's movement, improving its adaptability to various environments. Subsequently, the current observation sample information is input into the motion network included in the motor position control network to obtain a set of motion probability distribution values. Here, the motion network included in the motor position control network performs policy selection for the humanoid robot's actions, removing unnecessary trajectory references and improving the robot's decision-making ability for action selection. Next, the current observation sample information is input into the policy network included in the motor position control network to obtain a set of action evaluation values ​​corresponding to the motion probability distribution value set. Here, the policy network included in the motor position control network evaluates the cumulative reward of the humanoid robot's selected actions, improving the accuracy of the humanoid robot's action selection and enabling training for alternating foot support and foot-lift-off-ground movement during walking, thus improving the humanoid robot's steady movement. Subsequently, based on the aforementioned action probability distribution numerical set and action evaluation value set, a target joint position set for the humanoid robot is generated. Here, by using the action probability distribution numerical set and action evaluation value set to determine the action with the highest cumulative reward value, the accuracy of the motor position control network and the stability of the humanoid robot's action execution can be improved. Then, the aforementioned target joint position set is input to the joint response controller to generate a joint torque set for the humanoid robot.Here, the joint response controller enables precise tracking, high-precision positioning, and real-time dynamic response of the humanoid robot, improving its control stability. Then, based on the aforementioned joint torque set, the humanoid robot is controlled to perform foot-lifting movement training on simulated terrain corresponding to the simulated elevation terrain information. This avoids the humanoid robot using sliding friction for movement training, ensuring the robot produces a foot-lifting motion. Finally, in response to the determination that the aforementioned action probability distribution value set and the aforementioned action evaluation value set satisfy a preset loss condition, the training of the motor position control network ends. Here, iterative training of the motor position control network using the action probability distribution value set and the action evaluation value set improves the accuracy of the motor position control network in the humanoid robot's foot-lifting training scenario. Therefore, this humanoid robot foot-lifting movement training method, through the motor position control network and joint response controller, controls the humanoid robot to perform foot-lifting movement training on rough ground, improving the humanoid robot's movement stability, the performance of the motor position control network, and its generalization ability. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1 This is a flowchart of some embodiments of the humanoid robot foot-lifting and movement training method disclosed herein;

[0014] Figure 2 This is a schematic diagram of a humanoid robot with one degree of freedom, using a clog-shaped foot in a simulated environment, according to some embodiments of the humanoid robot foot-lifting and movement training device of this disclosure.

[0015] Figure 3 This is a schematic diagram of a humanoid robot with one degree of freedom, exhibiting a straight, flat foot in a real environment, according to some embodiments of the humanoid robot foot-lifting and movement training device disclosed herein.

[0016] Figure 4 This is a schematic diagram of the clog-shaped foot of a humanoid robot with two degrees of freedom in a simulated environment, according to some embodiments of the humanoid robot foot-lifting and movement training device of this disclosure.

[0017] Figure 5 This is a schematic diagram of the flat, plate-shaped foot of a humanoid robot with two degrees of freedom in a real environment, according to some embodiments of the humanoid robot foot-lifting and movement training device of this disclosure.

[0018] Figure 6 This is a schematic diagram of a humanoid robot with one degree of freedom, according to some embodiments of the humanoid robot foot-lifting and movement training device disclosed herein, in a simulated environment, where the clog-shaped foot forms a collision coupling with the rough ground of the simulated elevation terrain;

[0019] Figure 7 This is a schematic diagram of a humanoid robot with two degrees of freedom, according to some embodiments of the humanoid robot foot-lifting and movement training device disclosed herein, in a simulated environment, where the clog-shaped foot forms a collision coupling with the rough ground of the simulated elevation terrain;

[0020] Figure 8 These are schematic diagrams of some embodiments of the humanoid robot foot-lifting and movement training device according to the present disclosure;

[0021] Figure 9 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0022] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0023] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0024] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0025] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0026] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0027] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0028] Figure 1 A flow 100 of some embodiments of a humanoid robot foot-lifting movement training method according to the present disclosure is shown. This humanoid robot foot-lifting movement training method includes the following steps:

[0029] Step 101: Obtain the current observation sample information.

[0030] In some embodiments, the execution entity (e.g., an electronic device) of the above-described humanoid robot foot-lifting and movement training method can acquire current observation sample information via a wired or wireless connection. The humanoid robot is a bipedal humanoid robot with clog-shaped feet. The current observation sample information includes at least one of the following: joint positions, joint velocities, joint angles, and ground coupling pressure information on both feet. The humanoid robot can be a bipedal humanoid robot model constructed in a physical simulation environment, including thighs, calves, and clog-shaped feet. The humanoid robot can act as an intelligent agent and perform corresponding actions. For example, the corresponding actions can be, but are not limited to, at least one of the following: walking, bending over, running, and falling. The current observation sample information can be acquired through body sensor samples of the bipedal humanoid robot sample. The body sensor samples can be obtained through simulation and may include joint encoders and inertial measurement units (IMUs). The joint velocities can be the angular and linear velocities of each joint obtained by applying torque to the motors on each joint of the humanoid robot. The angular velocities of the aforementioned joints may include: knee joint angular velocity and ankle joint angular velocity. Linear velocity can be the linear velocity of the humanoid robot relative to each coordinate axis in the body coordinate system. Angular velocity can be the turning velocity of the humanoid robot relative to each coordinate axis in the body coordinate system. The ground coupling pressure information experienced by the aforementioned feet can be information on the pressure exerted by the terrain coupled with the terrain, obtained through sensors on the humanoid robot's feet. The ground coupling pressure information experienced by the aforementioned feet may include, but is not limited to, at least one of the following: the magnitude of the ground reaction force and the terrain height.

[0031] The aforementioned humanoid robot with one degree of freedom has clog-like feet, such as Figure 2 As shown, two capsule-shaped protrusions are added to the bottom of the humanoid robot's feet. The foot structure of a real robot does not necessarily follow the clog shape; it can be a flat, plate-like structure, such as... Figure 3 As shown. The clog-shaped feet of the aforementioned humanoid robot with two degrees of freedom are as follows. Figure 4 As shown, four circular collision shapes are added to the bottom of the humanoid robot's feet. The actual foot structure of a real robot is not necessarily manufactured in the shape of clogs; it can be a flat, plate-like structure, such as... Figure 5 As shown.

[0032] In some alternative implementations of some embodiments, the humanoid robot described above includes: thighs, lower legs, and feet, wherein:

[0033] The thigh and lower leg are connected by a locking structure at the knee joint. The knee joint consists of a ratchet and pawl assembly in conjunction with a four-bar linkage assembly. This locking structure controls the humanoid robot's knee extension and flexion postures. The ratchet and pawl assembly includes a pawl, a first spring, a ratchet, a reduction motor, a first bevel gear, and a second bevel gear. The four-bar linkage assembly includes a drive motor, a fixed connector, a long connecting rod, a short connecting rod, a stop, a synchronous pulley, a harmonic reducer, the thigh, and the lower leg. The locking structure at the knee joint restricts the knee extension state during flexion and extension, facilitating the support and swinging of the leg. One end of the short connecting rod in the four-bar linkage assembly extends out as a cantilever, with a stop on the cantilever. When the lower leg rotates to the required maximum angle, the lower leg and thigh are collinear. Driven by the lower leg, the long connecting rod drives the short connecting rod to swing, eventually bringing the short connecting rod to a vertical position. The stop on the cantilever of the short connecting rod precisely collides with the thigh surface, creating a limit for the knee joint in its extended state. When the thigh and lower leg are collinear, the short connecting rod, which is in a vertical position, is also aligned on the same straight line. All three are aligned, reaching a near-singularity; only by applying greater force or torque can they transition and continue rotating, thus limiting knee extension. The geared motor in the ratchet and pawl assembly uses reduced torque to drive the first bevel gear, which in turn drives the shaft, which is fitted with a ratchet. The shaft's rotation drives the ratchet's rotation. The pawl, mounted at the top of the cantilever, engages with the ratchet, hindering the pawl's movement and limiting knee flexion. The knee extension limit formed by the four-bar linkage and the knee flexion limit formed by the ratchet and pawl mechanism work together to form the knee joint's locking mechanism.

[0034] The aforementioned foot and lower leg are connected via the ankle joint. The foot consists of a target number of circular collision shapes added to the bottom of the foot. These circular collision shapes generate coupled collisions with the simulated elevation topography map corresponding to the simulated elevation topography information. The target number can be a pre-set value; for example, it could be 4. The coupled collisions can be controlled by the pressure on the humanoid robot's foot and the reaction force from the elevation map corresponding to the simulated elevation topography information, allowing the humanoid robot to lift its foot a certain distance before moving.

[0035] Step 102: Generate simulated elevation and terrain information for the humanoid robot.

[0036] In some embodiments, the executing entity can generate simulated elevation and terrain information for the humanoid robot. This simulated elevation and terrain information may be related to the continuous rough ground of the simulated environment in which the humanoid robot is located. The simulated elevation and terrain information may include, but is not limited to, at least one of the following: ground height and ground roughness. In practice, the executing entity can generate simulated elevation and terrain information for the humanoid robot using a digital elevation model.

[0037] In some optional implementations of certain embodiments, generating the simulated elevation terrain information for the humanoid robot described above may include the following steps:

[0038] The first step is to obtain a terrain planar mesh of a preset size, wherein each vertex of the terrain planar mesh includes a pseudo-random gradient vector. The preset terrain planar mesh can be a two-dimensional network with a pre-defined terrain size. This pre-defined terrain size can be set according to actual conditions and is not limited here. The pseudo-random gradient vector represents the pseudo-random positive or negative influence of any vertex in the terrain planar mesh relative to any point within the cell. Pseudo-randomness indicates that for any identical input value, the output result is the same. The positive influence can be that the gradient vector points in the direction of the force applied to any point. The negative influence can be that the gradient vector points in the opposite direction to the force applied to any point.

[0039] The second step involves determining the adjacent grid vertex groups in the aforementioned terrain plane network for each input random coordinate point in the input random coordinate point set. This results in an adjacent grid vertex group set, where the input random coordinate point is any randomly generated two-dimensional coordinate point located within the aforementioned terrain plane network. The adjacent grid vertex group can be four grid vertices located near the input random coordinate point. For example, if the input random coordinate point is a coordinate point within a cell grid, then the adjacent grid vertex group consists of four vertices within that cell grid. If the input random coordinate point is located at a grid vertex, then the adjacent grid vertex group consists of four grid vertices at a distance equal to the cell distance.

[0040] Third, for each adjacent grid vertex group in the above adjacent grid vertex group set, perform the following linear interpolation step:

[0041] Sub-step 1 involves determining the coordinate distance vector between each adjacent network vertex in the aforementioned adjacent network vertex group and its corresponding input random coordinate point, thus obtaining a coordinate distance vector group. This coordinate distance vector can be the difference between the coordinate vectors pointing from adjacent grid vertices to the input random coordinate point.

[0042] Sub-step 2 involves performing a vector dot product on each coordinate distance vector in the above coordinate distance vector group and its corresponding pseudo-random gradient vector to obtain a coordinate dot product vector group.

[0043] Sub-step 3 involves performing nonlinear interpolation on the aforementioned coordinate point-multiplication vector group to obtain a set of random noise values ​​for the input random coordinate points. The random noise values ​​in this set can be coordinate data newly inserted into the planar grid.

[0044] The third step involves generating initial noise fractal terrain information based on a preset transition curve and the obtained sets of random noise values. The preset transition curve can be a curve function of a coordinate dataset within a pre-defined smooth planar grid. The initial noise fractal terrain information can be information about the ground in the humanoid robot's environment.

[0045] The fourth step is to perform terrain rendering on the above-mentioned initial noise fractal terrain information to obtain the rendered initial noise fractal terrain information, which is used as the simulated elevation terrain information.

[0046] In addressing the first technical problem mentioned above, a second technical problem often arises: how to generate accurate simulated elevation terrain information that closely approximates real terrain to recreate the foot-lifting movement training of humanoid robots on different terrains. The conventional solution for this second technical problem is to generate simulated elevation terrain information using a two-dimensional Berlin noise algorithm. However, this conventional solution still suffers from the following issues: due to the limited terrain detail and distortion caused by the two-dimensional Berlin noise algorithm, it cannot realistically simulate real terrain information. This leads to significant differences between the humanoid robot's foot-lifting movement training and its training in a real environment, making it difficult to accurately grasp potential training problems in real-world environments, increasing the robot's damage rate, and reducing its safety. Considering the shortcomings of conventional solutions and the current state of humanoid robot movement training and deep learning technology at the inventor's research institute, the following solution has been adopted.

[0047] In some optional implementations of certain embodiments, generating the simulated elevation terrain information for the humanoid robot described above may include the following steps:

[0048] The first step is to acquire a terrain style image set. The terrain style images in this set can be images of a terrain feature captured from different shooting angles.

[0049] The second step is to determine the terrain location information of each terrain style image in the aforementioned terrain style image set, thereby obtaining a style two-dimensional terrain information set. This set contains the two-dimensional coordinate information of each location in the corresponding terrain style image. Specifically, the style two-dimensional terrain information in the aforementioned style two-dimensional terrain information set can be the two-dimensional information of each point on the terrain corresponding to the aforementioned terrain style image set.

[0050] The third step is to determine the average value of the above-mentioned style two-dimensional terrain information set to obtain the style two-dimensional average terrain information.

[0051] The fourth step involves constructing a noise fractal map from the aforementioned two-dimensional average terrain information to generate initial noise fractal terrain information. This initial fractal terrain information can be random terrain information.

[0052] The fifth step involves partitioning each terrain style image in the aforementioned terrain style image set into different regions to generate terrain style image patch groups, thus obtaining a terrain style image patch set. The partitioning process can be based on the different topographical features of the terrain.

[0053] Step 6: Input the aforementioned terrain style image patch set into the terrain style transfer model to obtain terrain style information. This terrain style information represents the detailed terrain information corresponding to the aforementioned terrain style image patch set. The terrain style transfer model can be a model including a transformation network and a loss network. The transformation network can be a deep residual network including 5 residual blocks. The loss network can be a loss network including terrain content loss and style loss. The terrain content loss represents the geomorphic information of the terrain. The style loss represents the similarity between the processed terrain and the terrain style image.

[0054] Step 7: Based on the above terrain style information, perform terrain style transformation on the above initial noise fractal terrain information to obtain style fractal terrain information.

[0055] The eighth step is to perform partitioned smoothing on the above style fractal terrain information to obtain smoothed style fractal terrain information, which is used as simulated elevation terrain information.

[0056] The above technical solution and its related content, combined with step "107," serve as an inventive point of this disclosure's embodiment, solving the second technical problem mentioned in the background: "Due to the limited terrain details and distorted landscape generated by the two-dimensional Berlin noise algorithm, it is impossible to realistically simulate real terrain information, resulting in a significant difference between foot-lift movement training and training in a real environment for humanoid robots. This makes it impossible to accurately grasp potential training problems in real environments, increasing the damage rate and reducing the safety of humanoid robots." The factors that cause this significant difference between foot-lift movement training and training in a real environment, making it impossible to accurately grasp potential training problems in real environments, increasing the damage rate, and reducing the safety of humanoid robots are often as follows: Due to the limited terrain details and distorted landscape generated by the two-dimensional Berlin noise algorithm, it is impossible to realistically simulate real terrain information. Solving these factors can reduce the training difference between foot-lift movement training and training in a real environment for humanoid robots, accurately grasp potential training problems in real environments, reduce the damage rate, and increase the safety of humanoid robots. To achieve this effect, this disclosure uses a generation method that combines fast style transfer with two-dimensional noise terrain generation. The basic terrain partitions generated by noise generation are subjected to fast style transfer, and then post-processing is used to ensure smooth transition of the blocks and add post-processing effects such as denoising. This can generate terrain with good details and specific style. While enriching the terrain details and improving the realism of the terrain, it ensures that the generation is fast and efficient and has good application value.

[0057] Step 103: Input the current observation sample information into the motion network included in the motor position control network to obtain the numerical set of motion probability distribution.

[0058] In some embodiments, the aforementioned execution entity can input the currently observed sample information into the motion network included in the motor position control network to obtain a set of motion probability distribution values. The motor position control network can be a deep reinforcement learning model trained on the input currently observed sample information to perform foot-lifting movements. The motor position control network can be a model employing an actor-critic neural network structure. The activation function of the motor position control network can be a reward function consisting of a target movement command and a power-related safety penalty term. The target command can be the target movement speed. The difference between the actual movement speed and the target movement speed is used as the reward term. Power is the product of the joint angular velocity and the motor torque output. The safety-related penalty term is the value exceeding a certain torque limit. The motion probability distribution values ​​in the aforementioned motion probability distribution set can be the normally distributed probability values ​​of each action chosen by the humanoid robot at the next moment due to the influence of environmental information such as simulated elevation terrain information. The motion network is trained through the following steps: inputting the robot's action posture set, which exists in the experience pool of the aforementioned motor position control network, into the motion network under the old and new strategies to obtain a first normal distribution and a second normal distribution of the humanoid robot's action probabilities under different strategies. Next, the stored robot pose set is input into a first normal distribution and a second normal distribution to obtain the first action probability value and the second action probability value corresponding to each robot pose. Then, the value of dividing the first action probability value by the second action probability value is determined as the importance weight. Finally, importance sampling is used to correct the difference between the two action distributions under the old and new strategies to obtain the action loss function value. Training is complete when the action loss function value is less than or equal to a preset action loss threshold. Retraining is performed when the action loss function value is greater than the preset action loss threshold. The action loss function can be a minimization of the advantage function. The advantage function can be the deviation between the predicted value function and the expected true cumulative return. The preset action loss threshold can be a pre-set loss threshold.

[0059] The multi-round random sampling actions output by the aforementioned action network are considered as a single action trajectory, starting from the robot's initial bipedal stance, interacting with simulated elevation terrain, and ending with a bipedal stance again after alternating bipedal walking. Within a single action trajectory, the robot determines its current action state, adopts an action posture according to the policy, receives a reward, and obtains the next state. The goal of policy optimization is to maximize the expected cumulative reward value after adopting an action posture in the current action state until the round ends.

[0060] The aforementioned movement trajectory can include two double support (DS) phases and two single support (SS) phases. The movement cycle, in chronological order, can include a left-foot support phase (DS1), a first double support phase (DD1), a right-foot support phase (DS2), and a second double support phase (DD2). DS1 and DD1 constitute the left-foot landing phase, and DS2 and DD2 constitute the right-foot landing phase. Both the left and right-foot landing phases account for 50% of the movement cycle T. The two double support phases together account for 20% of the movement cycle, and the two single support phases together account for 80%. For a single leg, there is only one swing cycle in the entire movement cycle, accounting for 40% of the movement cycle. Periodic movement rewards can penalize the speed of both feet during the two-footed support phase and incentivize the ground reaction force of both feet. During the right-footed support phase, the ground reaction force of the right foot and the speed of the left foot are incentivized, while the ground reaction force of the left foot and the speed of the right foot are penalized. Similarly, during the left-footed support phase, the ground reaction force of the left foot and the speed of the right foot are incentivized, while the ground reaction force of the right foot and the speed of the left foot are penalized. This helps the motor position control network learn the center of gravity reciprocation and alternating support movements of humanoid bipedal robot samples during walking.

[0061] In some optional implementations of certain embodiments, the action network includes: an action input layer, a first hidden action layer, a second hidden action layer, a third hidden action layer, an attention mechanism layer, and an action output layer. The first, second, and third hidden action layers can all be fully connected layers whose activation functions are all Rectified Linear Units (ReLU). The action input layer can be a network layer that receives information from the currently observed samples and passes it to the first hidden action layer. The action output layer can be a network layer that outputs a numerical set of action probability distributions.

[0062] Optionally, inputting the aforementioned current observation sample information into the motion network included in the motor position control network to obtain a numerical set of motion probability distributions may include the following steps:

[0063] The first step involves inputting the current observation sample information into the action input layer to obtain the first humanoid robot action feature vector. This first humanoid robot action feature vector characterizes the walking posture of the humanoid robot on the simulated elevation terrain corresponding to the simulated elevation terrain information.

[0064] The second step involves inputting the first humanoid robot motion feature vector into the first motion hidden layer to obtain the second humanoid robot motion feature vector. This second humanoid robot motion feature vector can be a feature vector obtained by linearly processing the current observed sample weight feature vector using an exponential linear activation function.

[0065] The third step is to input the second humanoid robot motion feature vector into the second motion hidden layer to obtain the third humanoid robot motion feature vector.

[0066] The fourth step is to input the aforementioned third humanoid robot motion feature vector into the aforementioned third motion hidden layer to obtain the fourth humanoid robot motion feature vector.

[0067] Fifth, the aforementioned fourth humanoid robot motion feature vector is input into the aforementioned attention mechanism layer to obtain the humanoid robot motion weight feature vector. This humanoid robot motion weight feature vector can represent feature vectors with different weights in different dimensions of the aforementioned fourth humanoid robot motion feature vector.

[0068] The sixth step involves inputting the aforementioned humanoid robot action weight feature vector into the action output layer to obtain the action mean and action variance. The action mean characterizes the average expected value of the humanoid robot's selected action under the current state of the simulated elevation terrain corresponding to the simulated elevation terrain information. The action variance characterizes the dispersion and uncertainty of the humanoid robot's different actions under the simulated elevation terrain.

[0069] Step 7: Construct a numerical set of action probability distributions based on the aforementioned action mean and variance. In practice, the executing entity can generate a normal probability distribution of actions using the aforementioned action mean and variance. Then, random sampling is performed on the aforementioned normal probability distribution to obtain the numerical set of action probability distributions. It should be noted that action networks can accelerate the learning of stable action strategies by humanoid robots.

[0070] Step 104: Input the current observation sample information into the strategy network included in the motor position control network to obtain the action evaluation value set corresponding to the action probability distribution numerical set.

[0071] In some embodiments, the execution entity can input the current observation sample information into a policy network included in the motor position control network to obtain an action evaluation value set corresponding to the action probability distribution value set. The policy network can be a network that evaluates the value of actions corresponding to the action probability distribution value set of the input current observation sample information under the simulated elevation terrain information. The action evaluation value in the action evaluation value set can be the reward value obtained by the humanoid robot after taking any action posture. The policy network can be trained through the following steps: First, the action trajectory is input into an initial policy network to obtain the state value corresponding to all states of the humanoid robot in an action trajectory. The state value can be expressed as:

[0072]

[0073] Wherein, V(s) t ) represents state s t The state value under the given conditions. 'a' represents an action trajectory. t This represents the action performed by the humanoid robot at time t. t This represents the state of the humanoid robot at time t. π(a t |s t ) represents the policy function corresponding to the policy network. t+1 Let r represent the state-action transition probability function at time t+. t Let s represent the reward function after the humanoid robot takes its next action at time t, given its current state. t+1 V(s) represents the state of the humanoid robot at time t+1. γ represents the discount factor. t+1 ) indicates that in state s t+1 The value of the state under the given conditions.

[0074] Secondly, the expected cumulative reward value is considered as the average of the expected cumulative rewards obtained by the humanoid robot after performing an action and reaching a certain state, and then taking different action postures. This yields the mean expected cumulative reward value. Furthermore, an advantage function is generated based on the mean expected cumulative reward value. The mean expected cumulative reward value can be expressed as:

[0075] G t =r t +γV(s t+1 ).

[0076] Among them, G t This represents the expected cumulative return value of the humanoid robot at time t.

[0077] The aforementioned advantage function can be expressed as:

[0078] Aπ (a t ,s t ) = G t -V(s t ).

[0079] Among them, A π (a t s t ) represents the dominance function.

[0080] Then, first-order time difference estimation is applied to the advantage function to obtain the differenced advantage function.

[0081] The advantage function after differencing can be expressed as:

[0082]

[0083] in, This represents the dominance function after differencing. δ t The deviation of the time difference, δ t =r t +γV(s t+1 )-V(s t λ represents the factor that accelerates the learning speed of the policy network. L represents the action value loss function. T represents the number of steps.

[0084] Finally, based on the differential advantage function, the action value loss function of the policy network is generated.

[0085] The action value loss function can be expressed as:

[0086]

[0087] Where N represents the number of samples included in the above-mentioned current observation sample information.

[0088] In some optional implementations of certain embodiments, the policy network includes: a value input layer, a first value hidden layer, a second value hidden layer, a third value hidden layer, a long short-term memory neural network, and a value output layer. The first, second, and third value hidden layers can be fully connected layers with an activation function of an exponential linear unit (ELU). The value input layer can be a network layer that receives information from the currently observed samples and passes it to the first value hidden layer. The value output layer can be a network layer that outputs a set of action evaluation values.

[0089] Optionally, inputting the aforementioned current observation sample information into the policy network included in the motor position control network to obtain the action evaluation value set corresponding to the aforementioned action probability distribution numerical set may include the following steps:

[0090] The first step involves inputting the currently observed sample information into the value input layer to obtain the first humanoid robot action value feature vector. This first humanoid robot action value feature vector represents the reward value obtained by the humanoid robot after selecting any action.

[0091] The second step is to input the first humanoid robot action value feature vector into the first value hidden layer to obtain the second humanoid robot action value feature vector.

[0092] The third step is to input the second humanoid robot action value feature vector into the second value hidden layer to obtain the third humanoid robot action value feature vector.

[0093] The fourth step is to input the aforementioned third humanoid robot action value feature vector into the aforementioned third value hidden layer to obtain the fourth humanoid robot action value feature vector.

[0094] Fifth, the motion value feature vector of the fourth humanoid robot is input into the long short-term memory neural network to obtain a walking posture temporal feature vector set. The posture temporal feature vectors in this set can represent the temporal information of the motion trajectory.

[0095] The sixth step is to input the above walking posture temporal feature vector set into the above value output layer to obtain the action evaluation value set corresponding to the above action probability distribution.

[0096] Step 105: Generate the target joint position set of the humanoid robot based on the numerical set of action probability distribution and the action evaluation value set.

[0097] In some embodiments, the execution entity can generate a target joint position set for the humanoid robot based on the numerical set of action probability distributions and the set of action evaluation values. The target joint positions in the target joint position set can be the locations of the joints corresponding to the desired actions of the humanoid robot. A schematic diagram illustrates a humanoid robot with one degree of freedom undergoing foot-lifting movement training on simulated elevation terrain corresponding to simulated elevation terrain information. This simulates the rough surface of the elevation terrain to create a certain coupling, preventing the humanoid robot from moving horizontally through ground friction. Figure 6 As shown in the diagram, a humanoid robot with two degrees of freedom is trained to lift its feet and move on simulated elevation terrain corresponding to simulated elevation terrain information. The rough surface of the simulated elevation terrain creates a certain coupling, preventing the humanoid robot from moving horizontally through ground friction. Figure 7 As shown.

[0098] As an example, the aforementioned executing entity can select the action evaluation value with the highest value from the set of action evaluation values ​​to obtain the target action evaluation value. Then, it can select the action probability distribution value corresponding to the target action evaluation value from the set of action probability distribution values ​​to obtain the target action probability distribution value. Finally, the set of positions of each joint of the humanoid robot corresponding to the target action probability distribution value is determined as the target joint position set.

[0099] Step 106: Input the target joint position set into the joint response controller to generate a joint torque set for the humanoid robot.

[0100] In some embodiments, the aforementioned execution entity can input the target joint position set to a joint response controller to generate a set of joint torques for the humanoid robot. These joint torques can be torques that instruct motors to generate, causing changes in the humanoid robot's posture. The joint response controller can be a controller that integrates a PD controller and a PID controller.

[0101] In addressing the first technical problem mentioned above, a third technical problem often arises: the joint torque output by the joint response controller is not precise or flexible enough, causing the humanoid robot to fail to respond and control to the expected joint position in a timely manner. The conventional solution to this third technical problem is to use traditional parameter tuning algorithms to dynamically improve the joint response controller to output accurate joint torque. However, this conventional solution still has the following problems: because traditional parameter tuning algorithms rely on certain expert experience and have simple calculation steps, their accuracy in adjusting the joint response controller in complex scenarios is low and limited, and they have a certain degree of subjectivity. This results in low accuracy of the joint response controller's output torque, making it difficult to effectively control the humanoid robot's leg-lifting movement training, increasing the damage rate and instability of the humanoid robot. Considering the shortcomings of conventional solutions and the current state of humanoid robot movement training and deep learning technology at the inventor's research institute, the following solution has been adopted.

[0102] In some optional implementations of certain embodiments, inputting the target joint position set to a joint response controller to generate a joint torque set for the humanoid robot may include the following steps:

[0103] The first step is to generate a joint control fitness function for the joint response controller, where the joint response controller is a function including a joint control proportional parameter, a joint control differential parameter, and a joint control differential order parameter. The joint control proportional parameter can be a parameter that rapidly responds to the error between the desired joint state and the actual joint state. This parameter generates the control output proportionally to the error, characterizing the joint response controller's sensitivity to error. The joint control differential parameter characterizes the adjustment of the control output based on the error rate of change, improving the dynamic performance of the joint response controller. The joint control differential order parameter characterizes the adjustment of the differential term. A higher order of the joint control differential order parameter results in better control output performance. The joint control fitness function can be a weighted sum of the overshoot of angle error, velocity error, control input, and step response.

[0104] The second step involves using a chaotic algorithm to generate an initial population for the target joint position set and the joint torque set. Each initial individual in the initial population is a parameter determination scheme that determines the parameter values ​​corresponding to the joint control proportional parameter, joint control differential parameter, and joint control differential order parameter.

[0105] The third step is to use a reverse learning algorithm to determine the initial reverse population of the initial population. The initial reverse individual in the initial reverse population can be the sum of the upper and lower bound values ​​of the initial individual, or the difference between the upper bound value and the corresponding value of the initial individual.

[0106] The fourth step is to input the initial population and the initial reverse population into the joint control fitness function to obtain the initial fitness value set and the initial reverse fitness set.

[0107] The fifth step involves filtering the initial fitness value set and the initial reverse fitness value set to obtain a filtered fitness value set. The filtered fitness value in the filtered fitness value set can be the largest value among the corresponding initial fitness value and initial reverse fitness value.

[0108] Step 6: Select individuals from the initial population and the initial reverse population that correspond to the fitness values ​​set after the selection process, and use them as the target initial population. The number of individuals in the target initial population, the initial population, and the initial reverse population are the same.

[0109] Step 7: Based on the target initial population, perform the following input steps:

[0110] Sub-step 1 involves performing cosine position update processing on each initial individual of the target population to generate updated initial individuals, thus obtaining the updated target population. The updated target individuals in the updated target population can be obtained by updating the initial individuals using an update function. The updated initial individuals are obtained through the following steps: First, in response to determining that the target being tracked has been detected as being tracked, it is first determined whether the random value is greater than a preset cosine-sine conversion threshold. This preset cosine-sine conversion threshold can be a pre-defined critical value for the conversion between the cosine and sine functions. For example, the preset cosine-sine conversion threshold can be 0.5. The target being tracked can be any initial individual other than the target initial individual. Second, the position difference between the target being tracked and the target initial individual is determined as the individual position difference. Third, in response to determining that the random value is less than the preset cosine-sine conversion threshold, the product of the search range conversion parameter, the cosine function, the flight step size, and the individual position difference is determined as the target individual position. The search range conversion parameter described above represents the conversion between global and local searches. Then, in response to determining that the random value is greater than or equal to a preset cosine-sine conversion threshold, the product of the search range conversion parameter, the sine function, the flight step size, and the individual position difference is determined as the target individual position. Finally, the sum of the target individual position and the initial target individual is determined as the updated individual. The search range conversion parameter described above represents the conversion parameter between global and local searches. In the second step, in response to determining that the target being tracked by the initial individual has not been detected as being tracked by the initial target individual, a random initial target individual is determined as the updated target individual.

[0111] Sub-step 2 involves performing mutation processing on the updated target population to obtain a mutated target population. The mutated target individuals in the mutated target population can be individuals obtained by multiplying the updated target individuals by the mutation probability value.

[0112] Sub-step 3 involves crossover processing the mutated target population to obtain a crossover target population. The crossover individuals in the crossover target population can be individuals obtained by crossovering the mutated target individual and its corresponding updated target individual using a preset number of dimensions, in response to the determination that the crossover probability value corresponding to the mutated target individual is greater than or equal to a preset crossover probability threshold. The crossover probability value can be a random value. The preset crossover probability threshold can be a pre-set maximum value of the crossover probability value. The preset number of dimensions can be 3.

[0113] Sub-step 4: Determine the number of times the above input steps have been executed.

[0114] Sub-step 5: In response to determining that the number of executions is greater than or equal to a preset execution threshold, the updated target population is input into the aforementioned joint control fitness function to obtain an updated fitness value set. The aforementioned preset execution threshold can be a pre-defined maximum number of executions.

[0115] Step 8: In response to determining that the number of executions is less than the aforementioned preset execution threshold, the target population after crossover is determined as the initial target population, and the sum of the number of executions and the preset value is determined as the number of executions, so that the above input steps are executed again. The aforementioned preset value can be a pre-defined value. For example, the aforementioned preset value can be 1.

[0116] Step 9: Select the highest updated fitness value from the set of updated fitness values ​​and use it as the target updated fitness value.

[0117] Step 10: Determine the set of parameter values ​​included in the crossover target individuals corresponding to the fitness values ​​of the updated targets as the target joint control proportional parameter value, the target joint control differential parameter value, and the target joint control differential order parameter value.

[0118] Step 11: Substitute the target joint control proportional parameter value, target joint control differential parameter value, and target joint control differential order parameter value into the joint response controller to obtain the target joint response controller.

[0119] Step 12: Input the target joint position set into the target joint response controller to obtain the joint torque set of the humanoid robot.

[0120] The above technical solution and its related content, combined with step "107" as an inventive point of this disclosure, solve the third technical problem mentioned in the background: "Because traditional parameter tuning algorithms rely on certain expert experience and have simple calculation steps, their accuracy in adjusting joint response controllers in complex scenarios is low and limited, and they have a certain degree of subjectivity, resulting in low accuracy of the output torque of the joint response controller, which cannot effectively control the humanoid robot's leg-lifting movement training, increasing the damage rate and instability of the humanoid robot." The factors that lead to low accuracy of the output torque of the joint response controller, inability to effectively control the humanoid robot's leg-lifting movement training, and increased damage rate and instability of the humanoid robot are often as follows: Because traditional parameter tuning algorithms rely on certain expert experience and have simple calculation steps, their accuracy in adjusting joint response controllers in complex scenarios is low and limited, and they have a certain degree of subjectivity. If these factors are solved, the accuracy of the output torque of the joint response controller can be improved, the leg-lifting movement training of the humanoid robot can be accurately controlled, the damage rate of the humanoid robot can be reduced, and the stability of the humanoid robot can be improved. To achieve this effect, this disclosure first generates a joint control fitness function to obtain the optimal parameter set under the constraints of the joint control fitness function. Second, it adjusts the initial population using chaotic and back-learning algorithms to improve the uniform distribution and diversity of the initial population in the search space, generating a feasible solution space with a good distribution. Then, iteratively updating the initial population using cosine and sine functions increases the search range in the early global search process and fully searches the neighborhood of the current solution in the later local search process. Next, iterative mutation and crossover processing are performed on the updated target population to increase its diversity, enhance the search capability of the parameter set, reduce the probability of getting trapped in local optima, and accelerate convergence. Finally, a target joint response controller is generated, and the target joint position set is input into the controller to obtain joint torques. This controller is then used to train the humanoid robot to perform foot lift-off movements, improving the accuracy of the joint torques, precisely controlling the humanoid robot to reach the expected joint positions, and enhancing the robot's stability.

[0121] Step 107: Based on the joint torque set, control the humanoid robot to perform foot lift-off movement training on the simulated elevation terrain corresponding to the simulated elevation terrain information.

[0122] In some embodiments, the aforementioned execution entity can control the humanoid robot to perform foot lift-off movement training on simulated elevation terrain corresponding to the aforementioned simulated elevation terrain information, based on the aforementioned joint torque set. This foot lift-off control training avoids the humanoid robot moving by friction with its feet firmly on the ground.

[0123] Step 108: In response to the determination that the numerical set of motion probability distribution and the set of motion evaluation value satisfy the preset loss condition, the training of the motor position control network ends.

[0124] In some embodiments, the execution entity may terminate the training of the motor position control network in response to determining that the numerical set of the action probability distribution and the numerical set of the action evaluation value satisfy a preset loss condition. The preset loss condition may be that the sum of the numerical value of the action loss and the value of the action value loss function is less than or equal to a preset loss threshold. The preset loss threshold may be a pre-set minimum loss value for training the motor position control network. The preset loss threshold can be determined according to specific circumstances and is not limited thereto.

[0125] As an example, the aforementioned execution entity can first determine the difference between the aforementioned action probability distribution value set and the sample action probability distribution value set, as the action loss value. The sample action probability distribution value set can be the action probability value set corresponding to the desired action selected by the actual humanoid robot. Then, it determines the difference between the aforementioned action evaluation value set and the sample action evaluation value set, as the action value loss function value. The sample action evaluation value set can be the real reward value set corresponding to the action selected by the real humanoid robot. Finally, in response to determining that the sum of the aforementioned action loss value and the aforementioned action value loss function value is less than or equal to a preset loss threshold, the training of the motor position control network ends.

[0126] In some optional implementations of some embodiments, in response to determining that the above-mentioned action loss function value and the above-mentioned action reward value loss function value do not meet the preset loss conditions, the observation sample information of the humanoid robot is reacquired as the current observation sample information, the relevant parameters of the above-mentioned motor position control network are readjusted to obtain the adjusted motor position control network, and the adjusted motor position control network is determined as the motor position control network, so as to retrain the above-mentioned motor position control network.

[0127] The above embodiments of this disclosure have the following beneficial effects: The humanoid robot foot-lifting movement training method of some embodiments of this disclosure, through a motor position control network and a joint response controller, controls the humanoid robot to perform foot-lifting training on rough ground, which can improve the movement stability of the humanoid robot and the performance and generalization of the motor position control network. Specifically, the reasons for the poor movement stability and high damage rate of the humanoid robot are: the need to collect motion reference trajectories and gait libraries increases additional data collection costs and wastes storage and computing resources; and the strategy model using motion reference trajectories and gait libraries learns specific gaits for specific environments, resulting in poor environmental adaptability for the humanoid robot, leading to poor movement stability and a high damage rate. Based on this, the humanoid robot foot-lifting movement training method of some embodiments of this disclosure can first obtain the current observation sample information of the humanoid robot, wherein the humanoid robot is a bipedal humanoid robot with clog-shaped feet, and the current observation sample information includes at least one of the following: the joint position, joint velocity, joint angle, and ground coupling pressure information of the humanoid robot's feet. Here, the current observation sample information is used for the generation of subsequent model elevation and terrain information and for training the humanoid robot's foot-lifting-off-the-ground movement control. Next, simulated elevation and terrain information is generated for the humanoid robot. This simulated elevation and terrain information can characterize any environmental information used for the humanoid robot's movement, improving its adaptability to various environments. Subsequently, the current observation sample information is input into the motion network included in the motor position control network to obtain a set of motion probability distribution values. Here, the motion network included in the motor position control network performs policy selection for the humanoid robot's actions, removing unnecessary trajectory references and improving the robot's decision-making ability for action selection. Next, the current observation sample information is input into the policy network included in the motor position control network to obtain a set of action evaluation values ​​corresponding to the motion probability distribution value set. Here, the policy network included in the motor position control network evaluates the cumulative reward of the humanoid robot's selected actions, improving the accuracy of the humanoid robot's action selection and enabling alternating support of both feet and foot-lifting-off-the-ground movement control during walking, thus improving the humanoid robot's steady movement. Subsequently, based on the aforementioned action probability distribution numerical set and action evaluation value set, a target joint position set for the humanoid robot is generated. Here, by using the action probability distribution numerical set and action evaluation value set to determine the action with the highest cumulative reward value, the accuracy of the motor position control network and the stability of the humanoid robot's action execution can be improved. Then, the aforementioned target joint position set is input to the joint response controller to generate a joint torque set for the humanoid robot.Here, the joint response controller enables precise tracking, high-precision positioning, and real-time dynamic response of the humanoid robot, improving its control stability. Then, based on the aforementioned joint torque set, the humanoid robot is controlled to perform foot-lifting movement training on simulated elevation terrain corresponding to the simulated elevation terrain information. This avoids the humanoid robot using sliding friction for movement training, ensuring the robot produces a foot-lifting motion. Finally, in response to the determination that the aforementioned action probability distribution value set and action evaluation value set satisfy a preset loss condition, the training of the motor position control network ends. Here, iterative training of the motor position control network using the action probability distribution value set and action evaluation value set improves the accuracy of the motor position control network in the humanoid robot's foot-lifting training scenario. Therefore, this humanoid robot foot-lifting movement training method, through the motor position control network and joint response controller, controls the humanoid robot to perform foot-lifting movement training on rough ground, improving the humanoid robot's movement stability and the performance and generalization of the motor position control network.

[0128] Further reference Figure 8 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a humanoid robot foot-lifting and movement training device, which are similar to... Figure 1 Corresponding to the method embodiments shown, this humanoid robot foot-lifting and movement training device can be specifically applied to various electronic devices.

[0129] like Figure 8As shown, a humanoid robot foot-lifting movement training device 800 includes: an acquisition unit 801, a first generation unit 802, a first input unit 803, a second input unit 804, a second generation unit 805, a third input unit 806, a control unit 807, and an termination unit 808. The acquisition unit 801 is configured to acquire current observation sample information of the humanoid robot, wherein the humanoid robot is a bipedal humanoid robot with clog-shaped feet, and the current observation sample information includes at least one of the following: joint position, joint velocity, joint angle, and ground coupling pressure information of the humanoid robot's feet. The first generation unit 802 is configured to generate simulated elevation terrain information for the humanoid robot. The first input unit 803 is configured to input the current observation sample information into the motion network included in the motor position control network to obtain a motion probability distribution value set. The second input unit 804 is configured to input the current observation sample information into the policy network included in the motor position control network to obtain a motion evaluation value set corresponding to the motion probability distribution value set. The second generation unit 805 is configured to generate a target joint position set for the humanoid robot based on the aforementioned action probability distribution value set and the aforementioned action evaluation value set. The third input unit 806 is configured to input the target joint position set to the joint response controller to generate a joint torque set for the humanoid robot. The control unit 807 is configured to control the humanoid robot to perform foot lift-off movement training on simulated elevation terrain corresponding to the aforementioned simulated elevation terrain information, based on the aforementioned joint torque set. The termination unit 808 is configured to terminate the training of the motor position control network in response to determining that the aforementioned action probability distribution value set and the aforementioned action evaluation value set satisfy a preset loss condition.

[0130] It is understandable that the units described in the humanoid robot foot-lifting and movement training device 800 are related to the reference Figure 1 The steps in the described method correspond accordingly. Therefore, the operations, features, and beneficial effects described above for the method also apply to the humanoid robot leg-lifting and movement training device 800 and the units contained therein, and will not be repeated here.

[0131] The following is for reference. Figure 9 It shows a schematic diagram of the structure of an electronic device (e.g., an electronic device) 900 suitable for implementing some embodiments of the present disclosure. Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0132] like Figure 9As shown, electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 901, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 902 or a program loaded from storage device 908 into random access memory (RAM) 903. RAM 903 also stores various programs and data required for the operation of electronic device 900. Processing device 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.

[0133] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 9 Each box shown can represent a device or multiple devices as needed.

[0134] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by the processing device 901, it performs the functions defined in the methods of some embodiments of this disclosure.

[0135] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0136] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0137] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: acquire current observation sample information of a humanoid robot, wherein the humanoid robot is a bipedal humanoid robot with clog-shaped feet, and the current observation sample information includes at least one of the following: joint positions, joint velocities, joint angles, and ground coupling pressure information of the humanoid robot's feet; generate simulated elevation and terrain information for the humanoid robot; input the current observation sample information into the motion network included in the motor position control network to obtain a set of motion probability distribution values; and input the current observation sample information into... The motor position control network, including a policy network, obtains the action evaluation value set corresponding to the aforementioned action probability distribution value set; based on the aforementioned action probability distribution value set and the aforementioned action evaluation value set, it generates the target joint position set of the humanoid robot; it inputs the aforementioned target joint position set to the joint response controller to generate the joint torque set for the humanoid robot; based on the aforementioned joint torque set, it controls the humanoid robot to perform foot lift-off movement training on the simulated elevation terrain corresponding to the aforementioned simulated elevation terrain information; in response to determining that the aforementioned action probability distribution value set and the aforementioned action evaluation value set satisfy a preset loss condition, it terminates the training of the aforementioned motor position control network.

[0138] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0140] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including an acquisition unit, a generation unit, a first input unit, a second input unit, a third input unit, a control unit, and a termination unit. The names of these units do not necessarily limit the specific unit; for example, the acquisition unit may also be described as "a unit that acquires information about currently observed samples."

[0141] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0142] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for training a humanoid robot to lift its foot and move, comprising: Obtain current observation sample information of a humanoid robot, wherein the humanoid robot is a bipedal humanoid robot with clog-shaped feet, and the current observation sample information includes at least one of the following: joint position, joint velocity, joint angle, and ground coupling pressure information of the humanoid robot's feet; Generate simulated elevation and terrain information for the humanoid robot; The current observation sample information is input into the motion network included in the motor position control network to obtain a numerical set of motion probability distribution; The current observation sample information is input into the strategy network included in the motor position control network to obtain the action evaluation value set corresponding to the action probability distribution value set; Based on the numerical set of the action probability distribution and the set of action evaluation values, the target joint position set of the humanoid robot is generated; The target joint position set is input to the joint response controller to generate a set of joint torques for the humanoid robot; Based on the joint torque set, the humanoid robot is controlled to perform foot lift-off movement training on the simulated elevation terrain corresponding to the simulated elevation terrain information. In response to determining that the numerical set of the action probability distribution and the set of action evaluation values ​​satisfy a preset loss condition, the training of the motor position control network is terminated.

2. The method according to claim 1, wherein, The method further includes: In response to the determination that the numerical set of the action probability distribution and the set of action evaluation values ​​do not meet the preset loss condition, the observation sample information of the humanoid robot is reacquired as the current observation sample information, the relevant parameters of the motor position control network are readjusted to obtain the adjusted motor position control network, and the adjusted motor position control network is determined as the motor position control network for retraining.

3. The method according to claim 1, wherein, The action network includes: an action input layer, an attention mechanism layer, a first action hidden layer, a second action hidden layer, a third action hidden layer, and an action output layer; and The step of inputting the current observation sample information into the motor position control network, including the action network, to obtain a numerical set of action probability distributions includes: The current observation sample information and the simulated elevation and terrain information are input into the action input layer to obtain the first humanoid robot action feature vector; The first humanoid robot motion feature vector is input into the first motion hidden layer to obtain the second humanoid robot motion feature vector; The second humanoid robot motion feature vector is input into the second motion hidden layer to obtain the third humanoid robot motion feature vector; The third humanoid robot motion feature vector is input into the third motion hidden layer to obtain the fourth humanoid robot motion feature vector; The fourth humanoid robot motion feature vector is input into the attention mechanism layer to obtain the humanoid robot motion weight feature vector. The humanoid robot's motion weight feature vector is input into the motion output layer to obtain the motion mean and motion variance. Construct a numerical set of action probability distributions for the action mean and the action variance.

4. The method according to claim 1, wherein, The policy network includes: a value input layer, a long short-term memory neural network, a first value hidden layer, a second value hidden layer, a third value hidden layer, and a value output layer; and The step of inputting the current observed sample information into the strategy network of the motor position control network to obtain the action evaluation value set corresponding to the action probability distribution value set includes: The current observation sample information and the simulated elevation and terrain information are input into the value input layer to obtain the first humanoid robot action value feature vector; The first humanoid robot action value feature vector is input into the first value hidden layer to obtain the second humanoid robot action value feature vector; The second humanoid robot action value feature vector is input into the second value hidden layer to obtain the third humanoid robot action value feature vector; The third humanoid robot action value feature vector is input into the third value hidden layer to obtain the fourth humanoid robot action value feature vector; The action value feature vector of the fourth humanoid robot is input into the long short-term memory neural network to obtain the walking posture temporal feature vector set; The walking posture temporal feature vector set is input into the value output layer to obtain the action evaluation value set corresponding to the action probability distribution.

5. The method according to claim 1, wherein, The humanoid robot includes: thighs, calves, and feet, wherein: The thigh and the lower leg are connected by a locking structure of the knee joint, which is composed of a ratchet and pawl assembly and a four-bar linkage assembly. The locking structure of the knee joint is used to control the humanoid robot to perform knee extension and bending movements. The foot and the lower leg are connected by the ankle joint. The foot is a foot with a target number of circular collision shapes added to the bottom of the foot. The circular collision shapes are collision shapes that couple with the simulated elevation terrain corresponding to the simulated elevation terrain information.

6. The method according to claim 1, wherein, The generation of simulated elevation and terrain information for the humanoid robot includes: Obtain a terrain plane mesh of a preset size, wherein each grid vertex in the terrain plane mesh includes a pseudo-random gradient vector; Determine the adjacent grid vertex group of each input random coordinate point in the set of input random coordinate points in the terrain plane grid to obtain the adjacent grid vertex group set, wherein the input random coordinate point is a randomly generated arbitrary two-dimensional coordinate point located in the terrain plane grid; For each adjacent grid vertex group in the set of adjacent grid vertex groups, perform the following linear interpolation steps: Determine the coordinate distance vector between each adjacent grid vertex and its corresponding input random coordinate point in the adjacent grid vertex group to obtain a coordinate distance vector group; Perform a vector dot product operation on each coordinate distance vector in the coordinate distance vector group and the corresponding pseudo-random gradient vector to obtain a coordinate dot product vector group; Linear interpolation is performed on the coordinate point multiplication vector group to obtain a random noise value group for the input random coordinate point; Based on the preset easing curve, and based on the obtained sets of random noise values, initial noise fractal terrain information is generated; The initial noise fractal terrain information is rendered to obtain the rendered initial noise fractal terrain information, which is used as the simulated elevation terrain information.

7. A training device for humanoid robots to lift their feet and move, comprising: The acquisition unit is configured to acquire current observation sample information of a humanoid robot, wherein the humanoid robot is a bipedal humanoid robot with clog-shaped feet, and the current observation sample information includes at least one of the following: joint position, joint velocity, joint angle, and ground coupling pressure information of the humanoid robot's feet; The first generation unit is configured to generate simulated elevation terrain information for the humanoid robot; The first input unit is configured to input the current observation sample information into the motion network included in the motor position control network to obtain a numerical set of motion probability distribution; The second input unit is configured to input the current observation sample information into the strategy network included in the motor position control network to obtain the action evaluation value set corresponding to the action probability distribution value set; The second generation unit is configured to generate the target joint position set of the humanoid robot based on the numerical set of the action probability distribution and the set of action evaluation values. The third input unit is configured to input the target joint position set to the joint response controller to generate a set of joint torques for the humanoid robot; The control unit is configured to control the humanoid robot to perform foot lift-off movement training on the simulated elevation terrain corresponding to the simulated elevation terrain information, based on the joint torque set. The termination unit is configured to terminate the training of the motor position control network in response to determining that the numerical set of the action probability distribution and the set of action evaluation values ​​satisfy a preset loss condition.

8. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.

9. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Animation processing method and device, computer storage medium and electronic equipment

    CN111292401A

  • Data processing method based on reinforcement learning, electronic equipment and readable medium

    CN116776097A