Motion control model training method, motion control method, device and medium
By combining knowledge distillation and reinforcement learning to optimize student models, the problem of poor performance of student models in existing technology in real scenarios is solved, and better motion control effects are achieved.
Patent Information
- Application Number
- CN202510405362.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-02
AI Technical Summary
In the prior art In the motion control of humanoid robots, the training method of the student model limits its performance in real scenarios and cannot surpass the behavior of the teacher model, resulting in poor performance in actual environments.
By combining knowledge distillation and reinforcement learning, the strategy training method of the student model is optimized, and the distillation loss and strategy optimization loss between the teacher model and the student model are used to adjust the strategies of the student model to generate better action strategies.
The motion control performance of the student model in real scenarios is improved, ensuring that it not only approaches the teacher model, but also generates better strategies and improves practical application effects.
Smart Images

Figure CN119918573B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method for training a motion control model, a motion control method, a device, and a medium. Background Art
[0002] In recent years, humanoid robots have attracted much attention in both academia and industry due to their high adaptability to the environment and human-like characteristics.
[0003] The development of reinforcement learning and deep learning has brought revolutionary changes to the walking control of humanoid robots, enabling them to go beyond the scope of preset actions and tasks and showing more possibilities. Among the many tasks of humanoid robots, motion control is one of the most critical and challenging tasks.
[0004] The existing technology usually first trains a teacher model in a simulated environment, and then uses distillation technology to transfer the learned privileged information to a student model in the form of latent features or actions. This method may limit the potential of the student model, making it only gradually approach the behavior of the teacher model and unable to surpass it, which is contrary to the core idea of reinforcement learning to learn better strategies through continuous trial-and-error interaction with the environment, resulting in poor performance in real scenarios. Summary of the Invention
[0005] The purpose of this application is to provide a method for training a motion control model, a motion control method, a device, and a medium for the above-mentioned deficiencies in the existing technology, so as to optimize the strategy of the student model by combining knowledge distillation and reinforcement learning and improve the performance of the motion control model in real scenarios.
[0006] To achieve the above purpose, the technical solutions adopted in the embodiments of this application are as follows:
[0007] In a first aspect, an embodiment of this application provides a method for training a motion control model for a humanoid robot, the method including:
[0008] Obtain first scan point information, second scan point information in a simulation environment, and first perception information of a target robot in the simulation environment, where the second scan point information is information obtained by adding noise scan points to the first scan point information;
[0009] According to the first scan point information and the first perception information, use a pre-trained teacher model to predict first action information of the target robot at the current moment;
[0010] According to the second scan point information and the first perception information, use a student model to be trained to predict second action information of the target robot at the current moment based on a new strategy, where the new strategy is the strategy adopted by the student model in the current training round;
[0011] Calculate a first training loss according to the first action information and the second action information;
[0012] Calculate a second training loss according to the second action information and the third action information, where the third action information is the action information generated by the student model using the old policy in the previous training round;
[0013] Train the policy of the student model according to the first training loss and the second training loss to obtain a motion control model of the target robot.
[0014] Optionally, the obtaining the first scan point information in the simulation environment includes:
[0015] Obtain the three-dimensional point cloud data of the simulation environment through a sensor;
[0016] Generate an elevation map of the simulation environment according to the three-dimensional point cloud data, where the elevation map is used to represent the terrain of the simulation environment;
[0017] Perform interval sampling on the elevation map to determine the first scan point information, where the first scan point information includes the terrain information of multiple scan points obtained by interval sampling.
[0018] Optionally, the predicting the first action information of the target robot at the current moment by using a pre-trained teacher model according to the first scan point information and the first perception information includes:
[0019] Generate a first terrain feature, a first historical perception feature, and a first estimated perception feature by using a feature extraction module in the teacher model according to the first scan point information and the historical perception information;
[0020] Encode and splice the first terrain feature, the first historical perception feature, and the first estimated perception feature to generate a first spliced feature;
[0021] Predict the first joint position information by using a policy module in the teacher model according to the first spliced feature and the current perception information, where the first action information includes: the first joint position information.
[0022] Optionally, the predicting the second action information of the target robot at the current moment by using a student model to be trained based on a new policy according to the second scan point information and the first perception information includes:
[0023] Generate a second terrain feature, a second historical perception feature, and a second estimated perception feature by using a feature extraction module in the student model according to the second scan point information and the historical perception information;
[0024] Encode and splice the second topographic feature, the second historical perception feature, and the second estimated perception feature to generate a second spliced feature;
[0025] According to the second spliced feature and the current perception information, use the policy module in the student model to predict the second joint position information;
[0026] According to the current perception information, use the evaluation module in the student model to determine the value information, and the second action information includes: the second joint position information and the value information.
[0027] Optionally, the calculating the second training loss according to the second action information and the third action information includes:
[0028] Calculate the policy loss according to the second joint position information and the third joint position information in the third action information;
[0029] Calculate the value function loss according to the second joint position information and the value information;
[0030] Calculate the entropy reward according to the current perception information;
[0031] Calculate the second training loss according to the policy loss, the value function loss, and the entropy reward.
[0032] Optionally, the calculating the value function loss according to the second joint position information and the value information includes:
[0033] Calculate at least one reward value according to at least one predefined reward function and the second joint position information;
[0034] Calculate the return value of the second joint position information at a future time according to the at least one reward value;
[0035] Calculate the value function loss according to the return value and the value information.
[0036] Optionally, the reward function includes at least one of the following: the bipedal motion reward function of the target robot, the joint limit reward function, the speed reward function, the acceleration reward function, the arm degree of freedom penalty function, the torso offset reward function, the command reward function.
[0037] In a second aspect, an embodiment of the present application further provides a motion control method for a humanoid robot, which is applied to the controller of the humanoid robot, and the method includes:
[0038] Obtain the third scan point information and the second perception information of the working environment;
[0039] Generate action information according to the third scanning point information and the second sensing information by using a pre-trained motion control model, where the motion control model is trained by using the motion control model training method for a humanoid robot according to any one of the first aspects;
[0040] Control the humanoid robot to move according to the action information.
[0041] In a third aspect, an embodiment of the present application provides a motion control model training device for a humanoid robot. The device includes:
[0042] A first information acquisition module, configured to acquire first scanning point information, second scanning point information in a simulation environment, and first sensing information of a target robot in the simulation environment, where the second scanning point information is information obtained by adding noise scanning points to the first scanning point information;
[0043] A first prediction module, configured to predict first action information of the target robot at the current moment by using a pre-trained teacher model according to the first scanning point information and the first sensing information;
[0044] A second prediction module, configured to predict second action information of the target robot at the current moment by using a student model to be trained based on a new strategy according to the second scanning point information and the first sensing information, where the new strategy is a strategy adopted by the student model in the current training round;
[0045] A loss calculation module, configured to calculate a first training loss according to the first action information and the second action information;
[0046] The loss calculation module is further configured to calculate a second training loss according to the second action information and third action information, where the third action information is action information generated by the student model by using an old strategy in the previous training round;
[0047] A model training module, configured to train the strategy of the student model according to the first training loss and the second training loss to obtain a motion control model of the target robot.
[0048] Optionally, the first information acquisition module is configured to obtain three-dimensional point cloud data of the simulation environment through a sensor; generate an elevation map of the simulation environment according to the three-dimensional point cloud data, where the elevation map is used to represent the terrain of the simulation environment; perform interval sampling on the elevation map to determine first scanning point information, where the first scanning point information includes terrain information of multiple scanning points sampled at intervals.
[0049] Optionally, the first prediction module is specifically configured to generate a first terrain feature, a first historical perception feature, and a first estimated perception feature by using the feature extraction module in the teacher model according to the first scan point information and the historical perception information; encode and splice the first terrain feature, the first historical perception feature, and the first estimated perception feature to generate a first spliced feature; and predict first joint position information by using the policy module in the teacher model according to the first spliced feature and the current perception information. The first action information includes: first joint position information.
[0050] Optionally, the second prediction module is specifically configured to generate a second terrain feature, a second historical perception feature, and a second estimated perception feature by using the feature extraction module in the student model according to the second scan point information and the historical perception information; encode and splice the second terrain feature, the second historical perception feature, and the second estimated perception feature to generate a second spliced feature; predict second joint position information by using the policy module in the student model according to the second spliced feature and the current perception information; and determine value information by using the evaluation module in the student model according to the current perception information. The second action information includes: the second joint position information and the value information.
[0051] Optionally, the loss calculation module is specifically configured to calculate a policy loss according to the second joint position information and the third joint position information in the third action information; calculate a value function loss according to the second joint position information and the value information; calculate an entropy reward according to the current perception information; and calculate the second training loss according to the policy loss, the value function loss, and the entropy reward.
[0052] Optionally, the loss calculation module is further configured to calculate at least one reward value according to at least one predefined reward function and the second joint position information; calculate a return value of the second joint position information at a future time according to the at least one reward value; and calculate the value function loss according to the return value and the value information.
[0053] Optionally, the reward function includes at least one of the following: a bipedal motion reward function of the target robot, a joint limit reward function, a speed reward function, an acceleration reward function, an arm degree of freedom penalty function, a torso offset reward function, an instruction reward function.
[0054] Fourthly, an embodiment of the present application further provides a motion control device for a humanoid robot, which is applied to a controller of the humanoid robot. The device includes:
[0055] A second information acquisition module, configured to acquire third scan point information and second perception information of the working environment.
[0056] A third prediction module, configured to generate action information according to the third scan point information and the second perception information by using a pre-trained motion control model, where the motion control model is trained by using the motion control model training method for a humanoid robot according to any one of the first aspects;
[0057] A control module, configured to control the humanoid robot to move according to the action information.
[0058] In a fifth aspect, an embodiment of the present application further provides an electronic device, including: a processor, a storage medium, and a bus. The storage medium stores program instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus. The processor executes the program instructions to execute the steps of the motion control model training method for a humanoid robot according to any one of the first aspects, or the steps of the motion control method for a humanoid robot according to the second aspect.
[0059] In a sixth aspect, an embodiment of the present application further provides a computer-readable storage medium, where a computer program is stored on the storage medium, and when the computer program is run by a processor, it executes the steps of the motion control model training method for a humanoid robot according to any one of the first aspects, or the steps of the motion control method for a humanoid robot according to the second aspect.
[0060] The beneficial effects of the present application are as follows:
[0061] For the motion control model training method for a humanoid robot provided in the foregoing embodiment, the strategy of the student model is optimized by combining the distillation loss between the teacher model and the student model and the strategy optimization loss for the student model to adjust the strategy in adjacent training rounds, so that the student model is not only able to approximate the teacher model. By adding reinforcement learning, it can be ensured that the student model generates better strategies and improves the performance of the motion control model in a real scenario. Description of the Drawings
[0062] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0063] Figure 1 Schematic flowchart of the motion control model training method for a humanoid robot provided in the embodiment of the present application Figure 1 ;
[0064] Figure 2 Flow schematic of the motion control model training method for a humanoid robot provided by an embodiment of the present application Figure 2 ;
[0065] Figure 3 Training architecture diagram of the student model provided by an embodiment of the present application;
[0066] Figure 4 Flow schematic of the motion control model training method for a humanoid robot provided by an embodiment of the present application Figure 3 ;
[0067] Figure 5 Flow schematic of the motion control model training method for a humanoid robot provided by an embodiment of the present application Figure 4 ;
[0068] Figure 6 Flow schematic diagram of the motion control method for a humanoid robot provided by an embodiment of the present application;
[0069] Figure 7 Structure schematic diagram of the motion control model training device for a humanoid robot provided by an embodiment of the present application;
[0070] Figure 8 Structure schematic diagram of the motion control device for a humanoid robot provided by an embodiment of the present application;
[0071] Figure 9 Schematic diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners
[0072] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application.
[0073] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0074] In addition, the terms "first", "second", etc. in the description, claims, and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0075] It should be noted that, without conflict, the features in the embodiments of this application can be combined with each other.
[0076] For the motion control model training method for a humanoid robot provided in the embodiments of this application, the electronic device used can be a computer device. For the motion control method of the humanoid robot provided in the embodiments of this application, the electronic device used can be a controller in the humanoid robot.
[0077] Figure 1 It is a flow schematic of the motion control model training method for a humanoid robot provided in the embodiments of this application Figure 1 , as Figure 1 shown, the method may include:
[0078] S101. Obtain first scan point information, second scan point information in the simulation environment, and first perception information of the target robot in the simulation environment. The first scan point information is information including noisy scan points.
[0079] In this embodiment, a simulation environment is generated by a simulator. The simulation environment can have an uneven terrain, such as slopes, steps, etc. Interval sampling is performed on the uneven terrain of the model simulation environment to obtain the scan information of multiple first scan points. The scan information of the first scan points may include: the height information of multiple first scan points. Based on the height information of multiple first scan points, the geometric shape of the terrain of the simulation environment can be represented.
[0080] In order to ensure that the motion control model trained in the simulation environment can bridge the gap between simulation and display, it is necessary to add noisy scan points on the basis of the first scan points to generate second scan point information including the first scan points and the noisy scan points. Among them, the noisy scan points can be randomly generated or generated by mixing and weighting based on the first scan points. This embodiment does not limit this.
[0081] The first perception information of the target robot in the simulation environment includes the perception state information generated by the target robot based on the perceived environmental information. The perception state information may include: the gait information, posture information, joint information, and action information of the target robot.
[0082] In some embodiments, the first perception information may include collecting the perception information of the target robot in the simulation environment during the training of the teacher model, and this embodiment does not limit this.
[0083] S102. According to the first scan point information and the first perception information, use the pre-trained teacher model to predict the first action information of the target robot at the current moment.
[0084] In this embodiment, the pre-trained teacher model can be a model trained using scan point information and perception information that do not include noisy scan points. During the training of the teacher model, the action reward can be calculated according to the Markov decision process (MDP), and the action generation strategy of the teacher model can be optimized based on the action reward.
[0085] The first perception information at least includes the perception information at the current moment. Input the first scan point information and the first perception information into the teacher model. The teacher model predicts the first action information of the target robot at the current moment according to the first scan point information and the first perception information. The first action information is the action of the target robot at the current moment determined based on the perception state information of the target robot in the simulation environment at the current moment. After the target robot controls its limbs to move based on the first action information, the perception information will be updated to the perception information after the movement.
[0086] S103. According to the second scan point and the first perception information, use the student model to be trained to predict the second action information of the target robot at the current moment based on the new strategy, where the new strategy is the strategy adopted by the student model in the current training round.
[0087] In this embodiment, after the training of the teacher model is completed, the model weights of the teacher model are transferred to the student model to be trained through knowledge distillation, so that the student model to be trained inherits the model weights of the teacher model.
[0088] It should be noted that since the scale of the student model is much smaller than that of the teacher model, the model weights of the teacher model inherited by the student model are not exactly the same weights, but the weights after knowledge distillation. That is, even if the student model inherits the model weights of the teacher model, it is not equivalent to the teacher model.
[0089] Input the second scan point information and the first perception information into the student model to be trained at the same time. The student model to be trained predicts the second action information of the target robot at the current moment based on the new strategy.
[0090] Among them, the new policy is the action generation policy of the current training round of the student model to be trained. The training process of the student model is actually a process of adjusting the action generation policy of the student model, so that the student model can output action information with the maximum cumulative reward based on the scan point information and perception information. The student model to be trained optimizes the old policy adopted in the previous training round in the previous training round to obtain the new policy.
[0091] S104. Calculate the first training loss according to the first action information and the second action information.
[0092] In this embodiment, according to the first action information output by the teacher model and the second action information output by the student model, the distillation loss between the teacher model and the student model is determined. The distillation loss is the loss caused during the process of distilling the model weights of the teacher model to the student model.
[0093] In some embodiments, the first training loss can be determined according to the mean square error of the first action information and the second action information.
[0094] S105. Calculate the second training loss according to the second action information and the third action information. The third action information is the action information generated by the student model using the old policy in the previous training round.
[0095] In this embodiment, reinforcement learning is used to optimize the model policy so that the model outputs actions with the maximum cumulative reward. The third action information is the action information generated by the student model using the old policy in the previous training round. In the previous training round, the distillation loss and the policy optimization loss are calculated based on the third action information to optimize the action generation policy of the student model and obtain the new policy of the current training round.
[0096] According to the second action information of the current training round and the third action information of the previous training round, the proximal policy optimization (PPO) loss during the policy optimization process of the student model is determined.
[0097] The proximal policy optimization loss can limit the policy update of the action generation model during the training process of the student model and the role of the reinforcement learning reward in the policy optimization of the student model.
[0098] S106. Train the policy of the student model according to the first training loss and the second training loss to obtain the motion control model of the target robot.
[0099] In this embodiment, the first training loss and the second training loss are weighted, and the strategy of the student model is trained according to the weighted training loss. When the training round reaches the preset round or the weighted training loss converges, the training of the student model is completed, and a motion control model of the target robot is obtained.
[0100] The motion control model training method for a humanoid robot provided in the above embodiment optimizes the strategy of the student model by combining the distillation loss between the teacher model and the student model and the strategy optimization loss of the student model for strategy adjustment in adjacent training rounds, so that the student model can not only approximate the teacher model. By adding reinforcement learning, it can ensure that the student model generates better strategies and improve the performance of the motion control model in the real scenario.
[0101] In a possible implementation manner, the process of obtaining the first scan point information in the simulation environment in S101 may include:
[0102] Obtain the three-dimensional point cloud data of the simulation environment through a sensor; generate a height map of the simulation environment according to the three-dimensional point cloud data, and the height map is used to represent the terrain of the simulation environment; perform interval sampling on the height map to determine the first scan point information, and the first scan point information includes the terrain information of multiple scan points obtained by interval sampling.
[0103] In this embodiment, the sensor may be, for example, a lidar. The three-dimensional point cloud data of the simulation environment is obtained through the sensor, and the terrain of the simulation environment is rendered based on the three-dimensional point cloud data to generate a height map of the simulation environment. The height map records the ground height of each point in the simulation environment, and the ground height is the ground height centered on the target robot.
[0104] By performing interval sampling on the height map, multiple first scan points can be obtained. The first scan point information is the terrain information of multiple first scan points, that is, the location height information. For example, the height map can be compressed into scan points, and multiple first scan points are obtained at a preset distance as the sampling interval.
[0105] Furthermore, since there is noise in the three-dimensional point cloud data collected by the sensor, the height information of the scan points can be updated using a Kalman filter.
[0106] For example, the update formula can be expressed as:
[0107]
[0108] Among them, is the height estimation variance of each cell in the height map, is the noise model variance of the sensor, and the noise model , d is the distance between the scanning point and the sensor, is a hyperparameter, is the height measurement value.
[0109] The motion control model training method for a humanoid robot provided by the above embodiment generates scanning point information based on an elevation map, so that the trained motion control model can control the motion of the humanoid robot in the face of various complex environments and improve the performance of the motion control model in the real environment.
[0110] In a possible implementation manner, Figure 2 is the flowchart of the motion control model training method for a humanoid robot provided by the embodiment of the present application Figure 2 , as Figure 2 shown, the process of the above S102 predicting the first action information of the target robot at the current moment according to the first scanning point information and the first perception information may include:
[0111] S201. Generate a first terrain feature, a first historical perception feature, and a first estimated perception feature according to the first scanning point information and the historical perception information by using the feature extraction module in the teacher model.
[0112] S202. Encode and splice the first terrain feature, the first historical perception feature, and the first estimated perception feature to generate a first spliced feature.
[0113] S203. Predict the first joint position information according to the first spliced feature and the current perception information by using the policy module in the teacher model. The first action information includes: the first joint position information.
[0114] In this embodiment, for example, Figure 3 is the training architecture diagram of the student model provided by the embodiment of the present application. As Figure 3 shown, the teacher model may include: a feature extraction module, a policy module, and an evaluation module. Among them, the evaluation model only calculates the state value of the robot when training the teacher model to optimize the policy of the policy module. After the model training is completed, the evaluation model is no longer used.
[0115] The feature extraction model of the teacher model extracts features from the first scanning point information to obtain a first terrain feature. For example, a one-dimensional convolution can be used to compress the first scanning point information and compress the first scanning point information into a latent space of the target dimension to obtain the first terrain feature.
[0116] The feature extraction module of the teacher model also extracts features from the historical perception information of a preset number of frames in the first perception information to obtain the first historical perception feature and the first estimated perception feature at the current moment estimated based on the historical perception information. Exemplarily, an external encoder can be used to compress the historical perception information into a latent space of a target dimension to obtain the first historical perception feature and the first estimated perception feature.
[0117] Encode the first terrain feature, the first historical perception feature, and the first estimated perception feature to obtain the first scan point vector, the first historical perception vector, and the first estimated perception vector, and splice the first scan point vector, the first historical perception vector, and the first estimated perception vector to generate the first spliced vector, and the first spliced vector is the historical latent vector.
[0118] The first perception information also includes the current perception information at the current moment. According to the perception latent vector of the current perception information and the first spliced vector, the strategy module in the teacher model is used for prediction to generate the first joint position information of the target robot, and the first joint position information is used to represent the action that the target robot is to perform based on the current perception information.
[0119] In a possible implementation manner, Figure 4 is the process flow diagram of the motion control model training method for a humanoid robot provided by the embodiments of the present application Figure 3 , as Figure 4 shown, the process of S103 using the student model to be trained to predict the second action information of the target robot at the current moment based on the second scan point and the first perception information may include:
[0120] S301: Generate a second terrain feature, a second historical perception feature, and a second estimated perception feature by using the feature extraction module in the student model according to the second scan point information and the historical perception information.
[0121] S302: Encode and splice the second terrain feature, the second historical perception feature, and the second estimated perception feature to generate a second spliced feature.
[0122] S303: Predict the second joint position information by using the strategy module in the student model according to the second spliced feature and the current perception information.
[0123] S304: Determine the value information by using the evaluation module in the student model according to the current perception information, and the second action information includes: the second joint position information and the value information.
[0124] In this embodiment, as Figure 3As shown, the student model adopts a structure corresponding to the teacher model, and also includes a feature extraction module, a policy module, and an evaluation module. The feature extraction module of the student model performs the processes of feature extraction, encoding, and splicing on the second scan point information and historical perception information, which is the same as the process of the feature extraction module of the teacher model performing feature extraction, encoding, and splicing on the first scan point information and historical perception information, and will not be elaborated here.
[0125] The current state information is used to refer to the current state s of the target robot. According to the second spliced feature and the current perception information, the policy module in the student model is used to predict the action of the target robot at the current moment, that is, the second joint position information, and according to the current perception information, the evaluation module in the student model is used to evaluate the current state s of the target robot to determine the value information of the current state.
[0126] In some embodiments, the process of calculating the first training loss according to the first action information and the second action information in S104 may include:
[0127] Calculate the first training loss according to the first joint position information and the second joint position information.
[0128] In other embodiments, Figure 5 is the flowchart of the motion control model training method for a humanoid robot provided by the embodiments of the present application Figure 4 as Figure 5 shown, the process of calculating the second training loss according to the second action information and the third action information in S105 may include:
[0129] S401. Calculate the policy loss according to the second joint position information and the third joint position information in the third action information.
[0130] In this embodiment, the policy module of the student model will adopt a new policy to generate multiple joint position information according to the second spliced feature and the current perception information. The multiple joint position information is the joint position information of multiple possible actions generated by the student module for the target robot. According to the mean and variance of the multiple joint position information, calculate the Gaussian distribution probability of each joint position information, and determine the joint position information with the largest Gaussian distribution probability as the second joint position information.
[0131] The student model determines the third joint position information generated by adopting the old policy in the previous round of training using the same scheme.
[0132] Calculate the policy loss according to the ratio of the Gaussian distribution probability of the second joint position information to the Gaussian distribution probability of the third joint position information .
[0133] In some embodiments, the probability ratio of the new and old policies is calculated according to the ratio of the Gaussian distribution probability of the second joint position information to the Gaussian distribution probability of the third joint position information. The policy loss is calculated according to the probability ratio, combined with the advantage function of performing the action corresponding to the second joint position information in the current state s and a preset clipping parameter. 。
[0134] Exemplarily, the first policy loss is calculated according to the product of the probability ratio and the advantage function. The clipped probability ratio is calculated by using a truncation function according to the probability ratio and the clipping parameter. The second policy loss is calculated according to the product of the clipped probability ratio and the advantage function. The minimum value of the first policy loss and the second policy loss is determined as the policy loss. 。
[0135] Among them, the clipping parameter is used to limit the probability ratio within the range of [1 - , 1 + to avoid excessive policy updates.
[0136] In some embodiments, according to the reward value and discount factor at the current moment determined based on the second joint position information, the return value of performing an action using the second joint position information from the current moment to a preset future moment is determined, that is, the discounted cumulative value of the reward. The advantage function is calculated according to the difference between the discounted cumulative value and the value information.
[0137] S402. Calculate the value function loss according to the second joint position information and the value information.
[0138] In this embodiment, according to the reward value and discount factor at the current moment determined based on the second joint position information, the return value of performing an action using the second joint position information from the current moment to a preset future moment is determined, that is, the discounted cumulative value of the reward. The value function loss is calculated according to the mean square error of the value information minus the discounted cumulative value. 。
[0139] S403. Calculate the entropy reward according to the current perception information.
[0140] In this embodiment, the entropy reward is calculated according to the Gaussian distribution probability and the logarithmic probability of the student model outputting multiple joint position information under the current perception information. 。
[0141] S404. Calculate the second training loss according to the policy loss, the value function loss, and the entropy reward.
[0142] In this embodiment, the policy loss minus the value function loss plus the entropy reward , to obtain the second training loss.
[0143] Exemplarily, the formula for calculating the second training loss can be expressed as:
[0144]
[0145] where c1 and c2 are coefficients for balancing the value function loss and the entropy reward .
[0146] For the motion control model training method for a humanoid robot provided in the above embodiment, the policy loss term ensures that the new policy does not deviate too much from the old policy through a clipping coefficient, the value function loss minimizes the difference between the estimated value and the actual return, and the entropy reward explores more policies by maximizing the entropy of the policy distribution. Training the student model with the second training loss composed of these three items of data can improve the stability and efficiency of the student model training, avoid premature convergence of the model training, and ensure the model training effect.
[0147] In a possible implementation manner, the process of calculating the value function loss according to the second joint position information and the value information in S402 may include:
[0148] Calculating at least one reward value according to at least one predefined reward function and the second joint position information; calculating the return value of the second joint position information at future times according to the at least one reward value; calculating the value function loss according to the return value and the value information.
[0149] In this embodiment, at least one reward function is predefined. The reward function is used to calculate the reward value of the target robot performing an action using the second joint position information in one dimension. According to at least one reward value of at least one reward function, the total reward value at the current moment is determined. According to the total reward value and the discount factors at multiple future times, the total return value of performing an action using the second joint position information from the current moment to multiple future times, that is, the discounted cumulative value of the reward, is determined. The value function loss is calculated according to the mean square error of subtracting the discounted cumulative value from the value information .
[0150] In some embodiments, the reward function includes at least one of the following: the bipedal motion reward function of the target robot, the joint limit reward function, the speed reward function, the acceleration reward function, the arm degree of freedom penalty function, the torso offset reward function, the command reward function.
[0151] In this embodiment, the bipedal motion reward function includes a swing-phase reward and a stance-phase reward during bipedal motion. The swing phase is the stage when the foot moves in the air, and the stance phase is the stage when the foot touches the ground. The reward function for each foot is a weighted function of the product of the phase indicator of the swing phase and the reward function of the swing phase, and the product of the phase indicator of the stance phase and the reward function of the stance phase. The reward function of the swing phase is calculated based on the square of the foot speed, and the reward function of the stance phase is calculated based on the square of the normal force of the foot.
[0152] The joint limit reward function is used to limit the joint degrees of freedom of the target robot, ensure that the motion is within the physical capabilities of the target robot, and prevent actions that may damage the mechanical structure of the robot or exceed the operating range of the robot. The joint limit reward function can be calculated based on the joint angles, preset minimum joint angles, and preset maximum joint angles in the second joint position information.
[0153] The speed reward function is used to ensure that the target robot maintains the optimal motion speed, and the acceleration reward function is used to ensure that the target robot accelerates and decelerates smoothly, ensuring the stability of the target robot's motion.
[0154] The arm degree-of-freedom penalty function is used to suppress the excessive motion of the robot's arm, and the torso offset reward function is used to maintain the balance and correct motion direction of the target robot.
[0155] The command reward function is used to encourage the target robot to move in a preset direction. The command reward function is calculated based on the current motion speed and expected speed of the robot, as well as the reward weights in each direction.
[0156] The motion control model training method for a humanoid robot provided in the above embodiment calculates the reward value based on multiple reward functions, and calculates the value function loss based on the reward value. When training the student model based on the value function loss, the student model can adjust the strategy to the optimal to output the best motion information to control the motion of the target robot.
[0157] Based on the motion control model training method for a humanoid robot provided in the above embodiment, the embodiment of the present application further provides a motion control method for a humanoid robot, which is applied to the controller of the humanoid robot.
[0158] Figure 6 It is a schematic flowchart of the motion control method for a humanoid robot provided in the embodiment of the present application, as Figure 6 shown, the method may include:
[0159] S501. Obtain the third scan point information and the second perception information of the working environment.
[0160] S502. Generate action information using a pre-trained motion control model according to the third scan point information and the second perception information.
[0161] S503. Control the humanoid robot to move according to the action information.
[0162] In this embodiment, the third scan point information of the working environment can be determined by scanning the working environment with a lidar installed on the robot, and the second perception information can be obtained through various sensors installed on the robot. The type of the sensor is determined by the type of information included in the second perception information, which is not limited herein.
[0163] Deploy the motion control model trained by the above-mentioned motion control model training method for humanoid robots into the controller of the humanoid robot. The controller generates action information using the motion control model according to the third scan point information and the second perception information, and controls the limbs of the humanoid robot to move according to the action information.
[0164] The motion control method for humanoid robots provided in the above embodiment can accurately control the movement of the humanoid robot by using the trained motion control model, improve the performance of the humanoid robot in the working environment, and the humanoid robot can work in the working environment with complex terrain.
[0165] Based on the above method embodiment, an embodiment of the present application provides a motion control model training device for a humanoid robot. Figure 7 As shown in the structural schematic diagram of the motion control model training device for a humanoid robot provided by an embodiment of the present application, Figure 7 as shown, the device may include:
[0166] The first information acquisition module 601 is configured to acquire the first scan point information, the second scan point information in the simulation environment, and the first perception information of the target robot in the simulation environment, where the second scan point information is the information obtained by adding noise scan points to the first scan point information;
[0167] The first prediction module 602 is configured to predict the first action information of the target robot at the current moment using a pre-trained teacher model according to the first scan point information and the first perception information;
[0168] The second prediction module 603 is configured to predict the second action information of the target robot at the current moment using the student model to be trained based on a new strategy according to the second scan point information and the first perception information, where the new strategy is the strategy adopted by the student model in the current training round;
[0169] The loss calculation module 604 is configured to calculate the first training loss according to the first action information and the second action information;
[0170] The loss calculation module 604 is further configured to calculate a second training loss according to the second action information and the third action information, where the third action information is the action information generated by the student model using the old policy in the previous training round;
[0171] The model training module 605 is configured to train the policy of the student model according to the first training loss and the second training loss to obtain a motion control model of the target robot.
[0172] Optionally, the first information acquisition module 601 is configured to obtain three-dimensional point cloud data of the simulation environment through a sensor; generate an elevation map of the simulation environment according to the three-dimensional point cloud data, where the elevation map is used to represent the terrain of the simulation environment; perform interval sampling on the elevation map to determine first scan point information, and the first scan point information includes terrain information of multiple scan points sampled at intervals.
[0173] Optionally, the first prediction module 602 is specifically configured to generate a first terrain feature, a first historical perception feature, and a first estimated perception feature by using a feature extraction module in the teacher model according to the first scan point information and the historical perception information; encode and splice the first terrain feature, the first historical perception feature, and the first estimated perception feature to generate a first spliced feature; and predict first joint position information by using a policy module in the teacher model according to the first spliced feature and the current perception information. The first action information includes: the first joint position information.
[0174] Optionally, the second prediction module 603 is specifically configured to generate a second terrain feature, a second historical perception feature, and a second estimated perception feature by using a feature extraction module in the student model according to the second scan point information and the historical perception information; encode and splice the second terrain feature, the second historical perception feature, and the second estimated perception feature to generate a second spliced feature; predict second joint position information by using a policy module in the student model according to the second spliced feature and the current perception information; and determine value information by using an evaluation module in the student model according to the current perception information. The second action information includes: the second joint position information and the value information.
[0175] Optionally, the loss calculation module 604 is specifically configured to calculate a policy loss according to the second joint position information and the third joint position information in the third action information; calculate a value function loss according to the second joint position information and the value information; calculate an entropy reward according to the current perception information; and calculate a second training loss according to the policy loss, the value function loss, and the entropy reward.
[0176] Optionally, the loss calculation module 604 is further configured to calculate at least one reward value according to at least one predefined reward function and the second joint position information; calculate the return value of the second joint position information at a future time according to the at least one reward value; and calculate the value function loss according to the return value and the value information.
[0177] Optionally, the reward function includes at least one of the following: bipedal motion reward function of the target robot, joint limit reward function, speed reward function, acceleration reward function, arm degree of freedom penalty function, torso offset reward function, instruction reward function.
[0178] Based on the foregoing method embodiments, an embodiment of the present application further provides a motion control device for a humanoid robot, which is applied to a controller of the humanoid robot. Figure 8 The structural schematic diagram of the motion control device for a humanoid robot provided by the embodiment of the present application is shown in Figure 8 As shown, the device may include:
[0179] A second information acquisition module 701, configured to acquire third scan point information and second perception information of the working environment;
[0180] A third prediction module 702, configured to generate action information by using a pre-trained motion control model according to the third scan point information and the second perception information, where the motion control model is trained by using the motion control model training method for a humanoid robot according to any item of the first aspect;
[0181] A control module 703, configured to control the humanoid robot to perform motion according to the action information.
[0182] The foregoing device is used to execute the method provided by the foregoing embodiment, and its implementation principle and technical effects are similar, and will not be described in detail here.
[0183] The foregoing modules may be one or more integrated circuits configured to implement the foregoing method, for example: one or more application specific integrated circuits (ASICs), or, one or more microprocessors, or, one or more field programmable gate arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. Again, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0184] Figure 9 It is a schematic diagram of the electronic device provided by the embodiment of the present application. The electronic device 800 may include: a processor 801, a storage medium 802, and a bus. The storage medium 802 stores program instructions executable by the processor 801. When the electronic device 800 runs, the processor 801 communicates with the storage medium 802 through the bus, and the processor 801 executes the program instructions to execute the above method embodiment. The specific implementation manners and technical effects are similar and will not be elaborated here.
[0185] Optionally, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the above method embodiment.
[0186] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0187] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0188] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit exists physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0189] The integrated unit implemented in the form of software functional units can be stored in a computer-readable storage medium. The above-mentioned software functional units are stored in a storage medium and include several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (English: Read-Only Memory, abbreviated as: ROM), random access memories (English: Random Access Memory, abbreviated as: RAM), magnetic disks, or optical discs.
[0190] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for training a motion control model for a humanoid robot, characterized in that, The method includes: Obtaining first scan point information, second scan point information in a simulation environment, and first perception information of a target robot in the simulation environment, where the second scan point information is information obtained by adding noise scan points to the first scan point information; According to the first scan point information and the first perception information, using a pre-trained teacher model to predict first action information of the target robot at the current moment; According to the second scan point information and the first perception information, using a student model to be trained to predict second action information of the target robot at the current moment based on a new policy, where the new policy is the policy adopted by the student model in the current training round; Calculating a first training loss according to the first action information and the second action information, where the first training loss is a distillation loss; Calculating a second training loss according to the second action information and third action information, where the third action information is action information generated by the student model using an old policy in the previous training round, and the second training loss is a proximal policy optimization loss; Training the policy of the student model according to the first training loss and the second training loss to obtain a motion control model of the target robot.
2. The method according to claim 1, wherein The obtaining of the first scan point information in the simulation environment includes: Obtaining three-dimensional point cloud data of the simulation environment through a sensor; Generating a height map of the simulation environment according to the three-dimensional point cloud data, where the height map is used to represent the terrain of the simulation environment; Performing interval sampling on the height map to determine the first scan point information, where the first scan point information includes terrain information of multiple scan points obtained by interval sampling.
3. The method according to claim 1, characterized in that The predicting of the first action information of the target robot at the current moment using the pre-trained teacher model according to the first scan point information and the first perception information includes: According to the first scan point information and historical perception information, using a feature extraction module in the teacher model to generate a first terrain feature, a first historical perception feature, and a first estimated perception feature; Encoding and splicing the first terrain feature, the first historical perception feature, and the first estimated perception feature to generate a first spliced feature; According to the first spliced feature and current perception information, using a policy module in the teacher model to predict first joint position information, where the first action information includes: first joint position information.
4. The method according to claim 1, wherein The predicting of the second action information of the target robot at the current moment using the student model to be trained based on the new policy according to the second scan point information and the first perception information includes: According to the second scan point information and historical perception information, using a feature extraction module in the student model to generate a second terrain feature, a second historical perception feature, and a second estimated perception feature; Encoding and splicing the second terrain feature, the second historical perception feature, and the second estimated perception feature to generate a second spliced feature; According to the second spliced feature and current perception information, using a policy module in the student model to predict second joint position information; Based on the current perception information, the evaluation module in the student model is used to determine value information, and the second action information includes: the second joint position information and the value information.
5. The method according to claim 4, wherein The calculating the second training loss according to the second action information and the third action information includes: Calculating a policy loss according to the second joint position information and the third joint position information in the third action information; Calculating a value function loss according to the second joint position information and the value information; Calculating an entropy reward according to the current perception information; Calculating the second training loss according to the policy loss, the value function loss, and the entropy reward.
6. The method according to claim 5, characterized in that, The calculating the value function loss according to the second joint position information and the value information includes: Calculating at least one reward value according to at least one predefined reward function and the second joint position information; Calculating a return value of the second joint position information at a future time according to the at least one reward value; Calculating the value function loss according to the return value and the value information.
7. The method according to claim 6, characterized in that, The reward function includes at least one of the following: the bipedal motion reward function of the target robot, the joint limit reward function, the speed reward function, the acceleration reward function, the arm degree of freedom penalty function, the torso offset reward function, the command reward function.
8. A motion control method for a humanoid robot, characterized in that, Applied to the controller of the humanoid robot, the method includes: Obtaining third scan point information and second perception information of the working environment; According to the third scan point information and the second perception information, generating action information by using a pre-trained motion control model, where the motion control model is trained by using the motion control model training method for a humanoid robot according to any one of claims 1 to 7; Controlling the humanoid robot to move according to the action information.
9. An electronic device, characterized in that, Including: A processor, a storage medium, and a bus. The storage medium stores program instructions executable by the processor. When the electronic device runs, the processor communicates with the storage medium through the bus. The processor executes the program instructions to perform the steps of the motion control model training method for a humanoid robot according to any one of claims 1 to 7, or the steps of the motion control method for a humanoid robot according to claim 8.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is run by the processor, it performs the steps of the motion control model training method for a humanoid robot according to any one of claims 1 to 7, or the steps of the motion control method for a humanoid robot according to claim 8.
Citation Information
Patent Citations
Control device and machine learning device
CN109814615A
Decision model training method, and strategy control method and device of target object
CN115238891A