Motion control method and device, electronic equipment and storage medium

By distilling and training a diverse motion expert policy network, motion control commands are generated, solving the problem of insufficient sensor configuration in robot motion control from simulation to real environment, and achieving high-performance motion control.

CN122185160APending Publication Date: 2026-06-12BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE
Filing Date
2026-01-30
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

In existing technologies, robot motion control methods face the problem of insufficient sensor configuration when deploying from simulation environments to real environments, which makes it impossible to directly apply expert strategies to real humanoid robot systems.

Method used

By distilling and training a diverse motion expert policy network, the robot's own state data is obtained and input into the motion control model to generate motion control commands. High-performance motion control is then achieved using the robot's available state data.

Benefits of technology

This enables the direct application of robot motion control in real-world environments, reducing reliance on costly or unstable environmental information and improving the practicality and reliability of motion control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122185160A_ABST
    Figure CN122185160A_ABST
Patent Text Reader

Abstract

The application provides a motion control method and device, electronic equipment and storage medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: acquiring self-state data of a robot; inputting the self-state data of the robot into a motion control model to obtain a motion control instruction output by the motion control model; wherein the motion control model is obtained by training an original control model according to sample self-state data and sample control instructions; the original control model is obtained by distillation training of a diversity motion expert policy network; the diversity motion expert policy network is obtained by reinforcement learning training based on sample complete state data related to diversity motion of the robot; and the robot is controlled according to the motion control instruction. The state data available for the robot can be used to realize motion control, and the method can be directly applied to system deployment of the robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a motion control method, device, electronic device, and storage medium. Background Technology

[0002] With the rapid development of artificial intelligence and reinforcement learning technologies, data-driven methods, especially deep learning-based control strategies, have become the mainstream approach for controlling robot motion in the field of robot motion control.

[0003] Currently, mainstream high-performance motion control methods in the industry mainly rely on training expert policy networks using reinforcement learning or imitation learning in a simulation environment with complete state data. However, when these well-trained expert policies are deployed to real robots, they often face severe practical limitations: the sensor configuration of real robot systems usually cannot provide all the complete state data relied upon during simulation training, making it impossible to directly apply expert policies to the system deployment of real humanoid robots. Summary of the Invention

[0004] To address the problems existing in the prior art, the present invention provides a motion control method, device, electronic device, and storage medium.

[0005] This invention provides a motion control method, comprising: Obtain the robot's own state data; The robot's own state data is input into the motion control model to obtain the motion control commands output by the motion control model; wherein, the motion control model is obtained by training an original control model based on the sample's own state data and sample control commands; the original control model is obtained by distillation training of a diverse motion expert policy network; the diverse motion expert policy network is obtained by reinforcement learning training based on the complete state data of the robot's diverse motion-related samples. The robot is motion controlled according to the motion control command.

[0006] According to a motion control method provided by the present invention, the distillation training of a diverse motion expert policy network includes: Obtain a complete motion dataset of diverse motions, and extract complete state data for at least two samples from the complete motion dataset; The complete state data of the sample is input into the diverse motion expert policy network to obtain the sample motion control instructions for diverse motion generated by the diverse motion expert policy network. The robot's own state data is determined from the complete state data of the sample; A first training dataset is obtained based on the sample's own state data corresponding to the complete state data of at least two samples and the sample motion control instructions of the diverse motion, so as to perform distillation training on the diverse motion expert policy network through the first training dataset.

[0007] According to a motion control method provided by the present invention, before performing distillation training on a diverse motion expert policy network using the training dataset, the method further includes: The value function of the diverse sports expert strategy network is inherited to the student strategy network, and the value function of the student strategy network is frozen. The step of distilling and training the diverse motion expert policy network using the first training dataset includes: The student policy network is trained under supervision using the first training dataset.

[0008] According to a motion control method provided by the present invention, the step of obtaining at least two sample complete state data from the complete motion dataset includes: Random sampling is performed from the complete motion dataset to obtain complete state data for at least two samples.

[0009] According to a motion control method provided by the present invention, the step of training the original control model based on the sample's own state data and sample control commands includes: Obtain a complete motion dataset of periodic motion, and extract complete state data for at least two samples from the complete motion dataset; The complete state data of the samples is input into a periodic motion expert policy network to obtain sample motion control commands for periodic motion generated by the periodic motion expert policy network; wherein, the periodic motion expert policy network is obtained by reinforcement learning training based on the complete state data of the samples related to the periodic motion of the robot. The robot's own state data is determined from the complete state data of the sample, so as to train the residual network of the original control model based on the own state data of the sample and the corresponding sample motion control command of the periodic motion.

[0010] According to a motion control method provided by the present invention, obtaining a complete motion dataset of periodic motion includes: Obtain template motion data for periodic motion; The template motion data of the periodic motion is input into the encoding model to obtain the template latent representation of the periodic motion output by the encoding model; wherein, the encoding model is trained by combining Fourier analysis and latent variable dynamic modeling. The template latent representation of the periodic motion is sampled to obtain at least two data latent representations; Decode the at least two latent representations of the data respectively to obtain the complete motion dataset of the periodic motion.

[0011] According to a motion control method provided by the present invention, the complete motion dataset of the diverse motion and the complete motion dataset of the periodic motion are both multi-source datasets.

[0012] The present invention also provides a motion control device, including: a self-state data acquisition module, used to acquire the self-state data of the robot; A motion control command determination module is used to input the robot's own state data into a motion control model to obtain motion control commands output by the motion control model; wherein, the motion control model is obtained by training an original control model based on the sample's own state data and sample control commands; the original control model is obtained by distillation training of a diverse motion expert policy network; the diverse motion expert policy network is obtained by reinforcement learning training based on the complete state data of the robot's diverse motion-related samples. The motion control module is used to control the motion of the robot according to the motion control instructions.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the motion control method described above.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the motion control method as described above.

[0015] This invention provides a motion control method, device, electronic device, and storage medium. By acquiring the robot's own state data and inputting this data into a motion control model, the invention obtains motion control commands output by the model. The robot can then be motion-controlled according to these commands. The motion control model is trained using sample state data and sample control commands. The original control model is obtained by distilling and training a diverse motion expert policy network. This network is trained using reinforcement learning based on complete sample state data related to the robot's diverse motions. This allows for high-performance determination of corresponding control commands based on state data, while simultaneously achieving motion control using the robot's available state data. This technology can be directly applied to robot system deployment. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating the motion control method provided by the present invention.

[0018] Figure 2 This is a schematic diagram of the motion control device provided by the present invention.

[0019] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0021] The following is combined Figures 1 to 3 The present invention describes a motion control method, apparatus, electronic device, and storage medium.

[0022] Figure 1 This is a flowchart illustrating the motion control method provided by the present invention, as shown below. Figure 1 As shown, the method includes the following: Step 101: Obtain the robot's own state data.

[0023] Robot self-state data refers to the data related to the robot's motion state that can be acquired in a real robot system. For example, robot self-state data may include pose data, IMU (Inertial Measurement Unit) data, etc.

[0024] Step 102: Input the robot's own state data into the motion control model to obtain the motion control commands output by the motion control model.

[0025] The motion control model is obtained by training the original control model based on the sample's own state data and the sample's control commands. The original control model was obtained by distilling and training a diverse motion expert policy network. The diverse motion expert policy network is obtained by reinforcement learning training based on the complete state data of the robot's diverse motion-related samples.

[0026] Among them, motion control commands refer to the control quantities that drive the robot to produce physical motion.

[0027] For example, motion control commands may include the target angle of the joint, the angular velocity of the joint, and the joint torque.

[0028] One feasible training scheme for the motion control model may include: The sample's own state data is input into the initial control model to obtain the control commands output by the initial control model. Then, based on the control commands and the sample control commands, the loss function value is calculated. Finally, based on the loss function value, the model parameters of the initial control model are updated. The above input and calculation processes are iteratively executed until the loss function converges or the preset number of iterations is reached, thus obtaining the motion control model. The preset number of iterations can be set as needed and is not specifically limited here.

[0029] One feasible training scheme for a diverse sports expert policy network may include: The complete state data of the sample of diverse motion and the linear velocity of the reference base are input into the initial policy network to obtain the control command output by the initial policy network. In the cyclic training, the reward for the interaction between the control command and the environment and the complete state data of the sample of the next state are obtained according to the first reward function until the preset number of iterations or performance index is reached. The parameters of the initial policy network and its corresponding value network are iteratively updated using a reinforcement learning algorithm with the goal of maximizing the cumulative reward. After training is completed, the policy network with updated parameters is used as the final usable expert policy network for diverse motion.

[0030] The first reward function can be constructed in the world coordinate system based on base linear velocity reference tracking, actual velocity consistency, and whole-body stability constraints.

[0031] There are many ways to distill and train a diverse motion expert policy network to obtain the original control model, and you can choose the appropriate method according to the actual working conditions. This embodiment does not limit this method.

[0032] Understandably, diverse motion expert policy networks rely on complete state data of samples. However, obtaining some environmental information from this complete state data is often costly and unreliable in real-time during deployment. This step uses distillation training to compress the complex motion knowledge and diverse policies of the diverse motion expert policy network into a motion control model that uses only its own state data as input. This inherits the high performance of the diverse motion expert policy network while reducing reliance on costly or unstable environmental information during deployment, thus laying the foundation for the practicality, reliability, and deployability of the robot's motion control.

[0033] Step 103: Perform motion control on the robot according to the motion control command.

[0034] It should be noted that, based on the motion control command, the robot's motion target in the next control cycle can be converted into physical actions of the robot's various actuators through a specific control algorithm.

[0035] The motion control method provided in this invention acquires the robot's own state data, inputs this data into a motion control model, and obtains motion control commands output by the model. This allows for motion control of the robot based on these commands. The motion control model is trained on an original control model using sample state data and sample control commands. The original control model is obtained by distilling a diverse motion expert policy network, which is trained using reinforcement learning based on complete sample state data related to the robot's diverse motions. This approach enables high-performance determination of corresponding control commands based on state data while simultaneously achieving motion control using the robot's available state data, allowing for direct application to robot system deployment.

[0036] Based on the above embodiments, the distillation training of the diverse motion expert policy network includes: Obtain a complete motion dataset of diverse motions, and extract complete state data for at least two samples from the complete motion dataset; The complete state data of the sample is input into the diverse motion expert policy network to obtain the sample motion control instructions for diverse motion generated by the diverse motion expert policy network. The robot's own state data is determined from the complete state data of the sample; A first training dataset is obtained based on the sample's own state data corresponding to the complete state data of at least two samples and the sample motion control instructions of the diverse motion, so as to perform distillation training on the diverse motion expert policy network through the first training dataset.

[0037] Diverse motion refers to a set of non-repetitive, multimodal motion patterns performed by a robot. These motion patterns do not have a definite temporal or spatial periodicity.

[0038] Examples include irregular gait variations, multi-directional random movement, and adaptive environmental interactions.

[0039] A complete motion dataset is a collection of complete motion data. In diverse motion scenarios, the robot's own state data is a subset of the complete motion data. Specifically, the complete motion data for diverse motion scenarios can be the collection of all information needed to describe the robot's diverse motion states.

[0040] The complete motion dataset for diverse motions can be a multi-source dataset, such as online reinforcement learning data, expert example data, and structured periodic motion data.

[0041] For example, the complete motion data portion of the robot's diverse motions, in addition to its own state data, may include friction parameters between the robot and the contact surface, the robot's inertia, etc.

[0042] It should be noted that inputting the complete sample state data into the diverse motion expert policy network yields sample motion control commands for diverse motions generated by the network. Simultaneously, the robot's own sample state data can be determined from this complete sample state data. Thus, based on the sample's own state data and the sample motion control commands generated by the diverse motion expert policy network, behavioral cloning training of the student policy network can be performed, achieving distillation training of the diverse motion expert policy network and obtaining the original control model for policy initialization.

[0043] As mentioned earlier, during the training of the diverse motion expert policy network, the reference base velocity corresponding to the motion is used as an explicit input feature and directly provided to the network. In this embodiment, during the behavior cloning training of the student policy network, it can implicitly learn the ability to estimate data outside its own state, such as the base reference linear velocity, within the complete state. This transforms data outside its own state from external input into part of the internal state inference, thereby maintaining stable and high-performance generation of motion control commands even when the input data does not include data outside its own state, such as the base reference linear velocity.

[0044] Based on any of the above embodiments, before performing distillation training on the diverse motion expert policy network using the training dataset, the method further includes: The value function of the diverse sports expert strategy network is inherited to the student strategy network, and the value function of the student strategy network is frozen. The step of distilling and training the diverse motion expert policy network using the first training dataset includes: The student policy network is trained under supervision using the first training dataset.

[0045] One feasible training scheme for the original control model may include: First, the value function of the diverse sports expert strategy network is inherited to the student strategy network, and then the value function of the student strategy network is frozen. Then, the sample state data from the first training dataset is input into the student policy network to obtain the control commands output by the student policy network. Based on the control commands and the sample control commands, the loss function value is calculated. Finally, the model parameters of the initial control model are updated based on the loss function value. This input and calculation process is iteratively executed until the loss function converges or the preset number of iterations is reached, resulting in the original control model. The preset number of iterations can be set as needed and is not specifically limited here.

[0046] It is understandable that inheriting the value function corresponding to the expert policy network into the student policy network and keeping it fixed during the subsequent training of the student policy network can reduce the risk of bias update caused by simultaneous updates of the policy network and the value network, and improve the stability of the training process.

[0047] Furthermore, by inheriting the value function of the expert policy network to the student policy network, the student policy network can implicitly learn the ultimate goal pursued by the expert policy network without the need for a reward function, thus simplifying the training process.

[0048] Furthermore, by inheriting the value function of the diverse motion expert strategy network, the student strategy network can select motion control instructions that are more consistent with the expert's long-term strategy when faced with multiple candidate motion control instructions, thus making the training process smoother.

[0049] To improve the unbiasedness of the reference motion distribution, based on any of the above embodiments, obtaining at least two complete state data samples from the complete motion dataset includes: randomly sampling from the complete motion dataset to obtain at least two complete state data samples.

[0050] Understandably, compared to directly using the captured raw motion data or simply perturbing the captured raw motion data, this embodiment randomly samples the reference motion from the complete motion dataset, which can control the motion distribution and reduce the risk of introducing abnormal motion that does not conform to dynamic constraints.

[0051] Based on any of the above embodiments, the step of training the original control model according to the sample's own state data and sample control commands includes: Obtain a complete motion dataset of periodic motion, and extract complete state data for at least two samples from the complete motion dataset; The complete state data of the samples is input into the periodic motion expert policy network to obtain the sample motion control commands for periodic motion generated by the periodic motion expert policy network; the periodic motion expert policy network is obtained by reinforcement learning training based on the complete state data of the samples related to the periodic motion of the robot. The robot's own state data is determined from the complete state data of the sample, so as to train the residual network of the original control model based on the own state data of the sample and the corresponding sample motion control command of the periodic motion.

[0052] Periodic motion refers to a set of repetitive motion patterns performed by a robot. These motion patterns have a definite temporal or spatial periodicity.

[0053] For example, exercises with fixed movement patterns, such as walking and running.

[0054] A complete motion dataset is a collection of complete motion data. In periodic motion, the robot's own state data is also a subset of the complete motion data. Specifically, the complete motion data for periodic motion can be the collection of all the information needed to describe the robot's periodic motion state.

[0055] The complete motion dataset for periodic motion can be a multi-source dataset, such as online reinforcement learning data, expert example data, and structured periodic motion data.

[0056] It should be noted that after training the basic network of the student policy network to obtain the original control model based on a complete motion dataset of diverse motions, the residual network of the original control model can be trained on a complete motion dataset of periodic motions to obtain the motion control model. In this way, the basic network can process the robot's own state data to control its diverse actions, and the residual network can be used to fine-tune the basic network, enabling the motion control model to process the robot's own state data to control its periodic actions. This efficiently integrates the broad adaptability of diverse actions with the high precision requirements of periodic actions in a single model, significantly improving the accuracy of robot control in complex mixed motion scenarios.

[0057] Understandably, the aforementioned phased training strategy can reduce the overall optimization complexity. Moreover, the residual network structure helps the motion control model converge quickly during the periodic motion training phase, and on the other hand, it can promote the effective integration and complementarity of knowledge of both diverse motion and periodic motion in the motion control model, reducing the risk of mutual interference between different expert strategies.

[0058] One feasible training scheme for periodic motion expert policy networks may include: The complete state data of the periodic motion samples and the linear velocity of the reference base are input into the initial policy network to obtain the control commands output by the initial policy network. During cyclic training, the rewards for the interaction between the control commands and the environment, as well as the complete state data of the next state samples, are obtained according to the second reward function until the preset number of iterations or performance indicators are reached. The parameters of the initial policy network and its corresponding value network are iteratively updated using a reinforcement learning algorithm with the goal of maximizing the cumulative reward. After training, the policy network with updated parameters is used as the final usable periodic motion expert policy network.

[0059] The second reward function can be constructed in the base coordinate system, with explicit inclusion of walking-related dynamic constraints.

[0060] Based on any of the above embodiments, obtaining the complete motion dataset of periodic motion includes: Obtain template motion data for periodic motion; The template motion data of the periodic motion is input into the encoding model to obtain the template latent representation of the periodic motion output by the encoding model; wherein, the encoding model is trained by combining Fourier analysis and latent variable dynamic modeling. The template latent representation of the periodic motion is sampled to obtain at least two data latent representations; Decode the at least two latent representations of the data respectively to obtain the complete motion dataset of the periodic motion.

[0061] Among them, the template motion data of periodic motion can be the raw motion data of captured periodic motion.

[0062] It should be noted that the template motion data of periodic motion is input into the encoding model, which performs Fourier transform on the template motion data of periodic motion to extract the frequency domain transform vector of the template motion data of periodic motion; the encoding model can simultaneously capture the time domain feature sequence of the template motion data of periodic motion; then, latent variable dynamics processing is performed on the frequency domain transform vector and the time domain feature sequence to obtain the latent representation of the template motion of periodic motion in a low-dimensional, continuous latent variable space.

[0063] Based on this, sampling can be performed on the template latent representation of the periodic motion to generate a large amount of structured complete motion data of the periodic motion from a limited number of template motion data of the periodic motion, providing a foundation for the training of the periodic motion expert policy network.

[0064] The motion control device provided by the present invention is described below. The motion control device described below can be referred to in correspondence with the motion control method described above.

[0065] Figure 2 This is a schematic diagram of the motion control device provided by the present invention, as shown below. Figure 2 As shown, the method includes: The self-state data acquisition module 210 is used to acquire the robot's own state data; The motion control command determination module 220 is used to input the robot's own state data into the motion control model to obtain the motion control commands output by the motion control model; wherein, the motion control model is obtained by training an original control model based on the sample's own state data and sample control commands; the original control model is obtained by distillation training of a diverse motion expert policy network; the diverse motion expert policy network is obtained by reinforcement learning training based on the complete state data of the robot's diverse motion-related samples; The motion control module 230 is used to perform motion control on the robot according to the motion control instructions.

[0066] Based on any of the above embodiments, a distillation training module is further included, for: Obtain a complete motion dataset of diverse motions, and extract complete state data for at least two samples from the complete motion dataset; The complete state data of the sample is input into the diverse motion expert policy network to obtain the sample motion control instructions for diverse motion generated by the diverse motion expert policy network. The robot's own state data is determined from the complete state data of the sample; A first training dataset is obtained based on the sample's own state data corresponding to the complete state data of at least two samples and the sample motion control instructions of the diverse motion, so as to perform distillation training on the diverse motion expert policy network through the first training dataset.

[0067] Based on any of the above embodiments, it further includes inheriting the value function of the diverse motion expert strategy network to the student strategy network and freezing the value function of the student strategy network; The distillation training module is used to: supervise the training of the student policy network using the first training dataset.

[0068] Based on any of the above embodiments, obtaining at least two complete state data samples from the complete motion dataset includes: randomly sampling from the complete motion dataset to obtain at least two complete state data samples.

[0069] Based on any of the above embodiments, a raw control model training module is further included, used for: Obtain a complete motion dataset of periodic motion, and extract complete state data for at least two samples from the complete motion dataset; The complete state data of the samples is input into a periodic motion expert policy network to obtain sample motion control commands for periodic motion generated by the periodic motion expert policy network; wherein, the periodic motion expert policy network is obtained by reinforcement learning training based on the complete state data of the samples related to the periodic motion of the robot. The robot's own state data is determined from the complete state data of the sample, so as to train the residual network of the original control model based on the own state data of the sample and the corresponding sample motion control command of the periodic motion.

[0070] Based on any of the above embodiments, the original control model training module is also used to acquire template motion data of periodic motion; The template motion data of the periodic motion is input into the encoding model to obtain the template latent representation of the periodic motion output by the encoding model; wherein, the encoding model is trained by combining Fourier analysis and latent variable dynamic modeling. The template latent representation of the periodic motion is sampled to obtain at least two data latent representations; Decode the at least two latent representations of the data respectively to obtain the complete motion dataset of the periodic motion.

[0071] Based on any of the above embodiments, the complete motion dataset of the diverse motion and the complete motion dataset of the periodic motion are both multi-source datasets.

[0072] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3As shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340, wherein the processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 can call logical instructions in the memory 330 to execute a motion control method, which includes: acquiring the robot's own state data; inputting the robot's own state data into a motion control model to obtain motion control instructions output by the motion control model; wherein the motion control model is obtained by training an original control model based on sample own state data and sample control instructions; the original control model is obtained by distillation training of a diverse motion expert policy network; the diverse motion expert policy network is obtained by reinforcement learning training based on the complete sample state data related to the robot's diverse motions; and performing motion control on the robot according to the motion control instructions.

[0073] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0074] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the motion control method provided by the above methods. The method includes: acquiring the robot's own state data; inputting the robot's own state data into a motion control model to obtain motion control commands output by the motion control model; wherein the motion control model is obtained by training an original control model based on sample own state data and sample control commands; the original control model is obtained by distillation training of a diverse motion expert policy network; the diverse motion expert policy network is obtained by reinforcement learning training based on the complete state data of the robot's diverse motion-related samples; and performing motion control on the robot according to the motion control commands.

[0075] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the motion control method provided by the above methods. The method includes: acquiring the robot's own state data; inputting the robot's own state data into a motion control model to obtain motion control commands output by the motion control model; wherein the motion control model is obtained by training an original control model based on sample own state data and sample control commands; the original control model is obtained by distillation training of a diverse motion expert policy network; the diverse motion expert policy network is obtained by reinforcement learning training based on complete sample state data related to the robot's diverse motions; and performing motion control on the robot according to the motion control commands.

[0076] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0077] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A motion control method, characterized in that, include: Obtain the robot's own state data; The robot's own state data is input into the motion control model to obtain the motion control commands output by the motion control model; wherein, the motion control model is obtained by training an original control model based on the sample's own state data and sample control commands; the original control model is obtained by distillation training of a diverse motion expert policy network; the diverse motion expert policy network is obtained by reinforcement learning training based on the complete state data of the robot's diverse motion-related samples. The robot is motion controlled according to the motion control command.

2. The motion control method according to claim 1, characterized in that, The distillation training of the diverse motion expert policy network includes: Obtain a complete motion dataset of diverse motions, and extract complete state data for at least two samples from the complete motion dataset; The complete state data of the sample is input into the diverse motion expert policy network to obtain the sample motion control instructions for diverse motion generated by the diverse motion expert policy network. The robot's own state data is determined from the complete state data of the sample; A first training dataset is obtained based on the sample's own state data corresponding to the complete state data of at least two samples and the sample motion control instructions of the diverse motion, so as to perform distillation training on the diverse motion expert policy network through the first training dataset.

3. The motion control method according to claim 2, characterized in that, Before distilling the diverse motion expert policy network using the training dataset, the method further includes: The value function of the diverse sports expert strategy network is inherited to the student strategy network, and the value function of the student strategy network is frozen. The step of distilling and training the diverse motion expert policy network using the first training dataset includes: The student policy network is trained under supervision using the first training dataset.

4. The motion control method according to claim 2, characterized in that, The step of obtaining at least two complete state data samples from the complete motion dataset includes: Random sampling is performed from the complete motion dataset to obtain complete state data for at least two samples.

5. The motion control method according to claim 1, characterized in that, The step of training the original control model based on the sample's own state data and sample control commands includes: Obtain a complete motion dataset of periodic motion, and extract complete state data for at least two samples from the complete motion dataset; The complete state data of the samples is input into a periodic motion expert policy network to obtain sample motion control commands for periodic motion generated by the periodic motion expert policy network; wherein, the periodic motion expert policy network is obtained by reinforcement learning training based on the complete state data of the samples related to the periodic motion of the robot. The robot's own state data is determined from the complete state data of the sample, so as to train the residual network of the original control model based on the own state data of the sample and the corresponding sample motion control command of the periodic motion.

6. The motion control method according to claim 5, characterized in that, The process of obtaining a complete motion dataset for periodic motion includes: Obtain template motion data for periodic motion; The template motion data of the periodic motion is input into the encoding model to obtain the template latent representation of the periodic motion output by the encoding model; wherein, the encoding model is trained by combining Fourier analysis and latent variable dynamic modeling. The template latent representation of the periodic motion is sampled to obtain at least two data latent representations; Decode the at least two latent representations of the data respectively to obtain the complete motion dataset of the periodic motion.

7. The motion control method according to claim 5 or 6, characterized in that, The complete motion datasets for the diverse motions and the complete motion datasets for the periodic motions are both multi-source datasets.

8. A motion control device, characterized in that, include: The self-state data acquisition module is used to acquire the robot's own state data; A motion control command determination module is used to input the robot's own state data into a motion control model to obtain motion control commands output by the motion control model; wherein, the motion control model is obtained by training an original control model based on the sample's own state data and sample control commands; the original control model is obtained by distillation training of a diverse motion expert policy network; the diverse motion expert policy network is obtained by reinforcement learning training based on the complete state data of the robot's diverse motion-related samples. The motion control module is used to control the motion of the robot according to the motion control instructions.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the motion control method as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the motion control method as described in any one of claims 1 to 7.