Robot motion control model training method and device based on deep reinforcement learning

By adopting a teacher-student model framework for deep reinforcement learning in robot motion control, the linear velocity encoder and encoder generate potential vectors, combined with the value estimation of the value output from the value network, and updating the parameters of the policy network, it solves the problem that traditional methods are difficult to achieve stable and precise control in complex dynamic environments, and improves the control performance of the robot in high-speed or high-dynamic environments.

CN120065751AActive Publication Date: 2025-05-30SHENZHEN ZHUJI POWER TECH CO LTD

Patent Information

Application Number
CN202510527132.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

Traditional robot motion control methods are difficult to achieve stable and accurate control in complex dynamic environments, scenes with severe velocity changes or high-dimensional control tasks, especially in the case of unknown disturbances or partial observation information.

Method used

The teacher-student model framework based on deep reinforcement learning is used to train the policy network, and potential vectors are generated through linear speed encoder, student encoder and teacher encoder, and combined with the value estimation of the value network output, the deep reinforcement learning algorithm is used to update the parameters of the policy network.

Benefits of technology

It improves the control accuracy and stability of the robot in speed changes or high dynamic scenarios, reduces the control lag phenomenon, improves the response speed and stability of the control strategy, and enhances the robot's adaptability in high-speed or high-dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120065751A_ABST
    Figure CN120065751A_ABST
Patent Text Reader

Abstract

The invention provides a robot motion control model training method and device based on deep reinforcement learning, and relates to the technical field of sensors and robots. The method comprises the following steps: performing deep reinforcement learning-based training on a strategy network for outputting an action strategy for controlling the robot to move by utilizing a teacher-student model framework; wherein the training process of the strategy network further comprises the steps of encoding historical linear velocity information of the robot through a linear velocity encoder to generate a first potential vector, and inputting the first potential vector into the strategy network for auxiliary training. According to the method, the motion trend, the state change track and the surrounding terrain structure of the robot are comprehensively considered during action decision making, and strategy output is more stably and accurately controlled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical fields of sensors and robotics, and relates to a method and device for training a robot motion control model based on deep reinforcement learning. Background Art

[0002] In the field of robot motion control, traditional control methods usually rely on model predictive control (MPC), proportional-integral-derivative control (PID), or classical optimization algorithms. However, these methods often struggle to achieve stable and precise control when faced with complex dynamic environments, scenarios with drastic speed changes, or high-dimensional control tasks. Especially in the presence of unknown disturbances or partial observation information, the adaptability of traditional methods is poor, and it is difficult to quickly adjust the strategy, thus affecting the motion performance of the robot in complex environments.

[0003] In recent years, control methods based on deep reinforcement learning (DRL) have gradually become an important research direction in the field of robot control. Deep reinforcement learning uses neural networks to extract features from high-dimensional input information and learns optimal control strategies through policy optimization algorithms (such as PPO, DDPG, etc.).

[0004] For example, in the paper "CTS Concurrent Teacher-Student Reinforcement Learning for Legged Locomotion" (arXiv:2405.10830v2 [cs.RO], 20240901), a teacher-student model is adopted to improve training efficiency and policy stability. Similar teacher-student models are also disclosed in the Chinese patent application with publication number CN119458315A, the Chinese patent application with publication number CN116931475A, and the Korean patent application with publication number KR1020250009137A. In such model architectures, the teacher policy is trained in an environment with complete information, while the student policy learns in a restricted information environment and gradually approximates the teacher policy through knowledge distillation.

[0005] However, since the student model usually only has access to limited observation information during deployment, and the privileged information relied on by the teacher model is not available during deployment, the learning effect of the student model is limited and it is difficult to reach the level of the teacher policy. For example, the inventors found that the control policy output by the policy network is prone to robot control lag or instability in environments with large speed changes or high-speed motion.

[0006] Therefore, although the above method improves the training efficiency and stability of the policy network, the accuracy and stability of the policy network in scenarios with speed changes or high dynamics still need to be further optimized. Summary of the Invention

[0007] The present disclosure provides a method and apparatus for training a robot motion control model based on deep reinforcement learning, which are used to improve the control accuracy and stability of a robot in scenarios with speed changes or high dynamics.

[0008] Additional aspects and advantages of the present disclosure will be partly set forth in the description below, and partly will be obvious from the description, or can be learned by practice of the present disclosure.

[0009] According to a first aspect of the present disclosure, there is provided a method for training a robot motion control model based on deep reinforcement learning, including: Using a teacher-student model framework, performing deep reinforcement learning-based training on a policy network for outputting an action policy for controlling the motion of a robot; wherein, the training process of the policy network includes: Encoding the historical linear velocity information of the robot through a linear velocity encoder to generate a first latent vector, and inputting the first latent vector into the policy network; Based on the value estimation of the current state output by the value network in the teacher-student model framework and the action policy output by the policy network in combination with the first latent vector, updating the parameters of the policy network using a deep reinforcement learning algorithm.

[0010] In an exemplary embodiment of the present disclosure, the teacher-student model framework includes a student encoder and a teacher encoder; Using a teacher-student model framework, performing deep reinforcement learning-based training on a policy network for outputting an action policy for controlling the motion of a robot, including: Encoding the historical motion state information of the robot itself through the student encoder to generate a second latent vector, and encoding the privileged state information of the robot through the teacher encoder to generate a third latent vector; Selecting the second latent vector or the third latent vector according to a preset policy and inputting it into the policy network; Based on the received first latent vector, the second latent vector or the third latent vector selected by the preset policy, and the current motion state information, the policy network outputs an action policy for controlling the motion of the robot; The value network uses the privileged state information to output a value estimation of the current state; Based on the action policy and the value estimation, updating the parameters of the policy network using a deep reinforcement learning algorithm.

[0011] In an exemplary embodiment of the present disclosure, the value network utilizes privileged state information to output a value estimate of the current state, including: The value network utilizes privileged state information, and in combination with the first latent vector, the second latent vector or the third latent vector selected by a preset policy, outputs a value estimate of the current state.

[0012] In an exemplary embodiment of the present disclosure, the linear velocity encoder includes: An input layer, configured to receive the linear velocity information of the robot at multiple historical moments and preprocess the linear velocity information; A hidden layer, configured to extract temporal features of the preprocessed linear velocity information by using a neural network model; An output layer, configured to generate a first latent vector according to the temporal features extracted by the hidden layer.

[0013] In an exemplary embodiment of the present disclosure, the neural network model in the hidden layer includes a long short-term memory neural network or a gated recurrent unit neural network.

[0014] In an exemplary embodiment of the present disclosure, based on the received first latent vector, the second latent vector or the third latent vector selected by a preset policy, and the current motion state information, the policy network outputs an action policy for controlling the motion of the robot, including: The policy network is based on: Outputs an action policy for controlling the motion of the robot; Wherein, represents the probability distribution of taking action in the current state , MLP represents a multi-layer perceptron, represents the second latent vector or the third latent vector selected according to a preset policy, represents the first latent vector.

[0015] In an exemplary embodiment of the present disclosure, the training process of the policy network further includes: Encoding the historical elevation map information of the robot by a topographic map encoder to generate a fourth latent vector, and inputting the fourth latent vector into the policy network for auxiliary training.

[0016] In an exemplary embodiment of the present disclosure, the teacher-student model framework includes a student encoder and a teacher encoder; Utilizing the teacher-student model framework to perform training based on deep reinforcement learning on the policy network for outputting an action policy for controlling the motion of the robot, including: Encode the historical motion state information of the robot itself through the student encoder and generate a second latent vector, and encode the privileged state information of the robot through the teacher encoder and generate a third latent vector; Select the second latent vector or the third latent vector according to a preset policy and input it into the policy network; Based on the received first latent vector, fourth latent vector, the second latent vector or the third latent vector selected by the preset policy, and the current motion state information, the policy network outputs an action policy for controlling the movement of the robot; The value network uses the privileged state information to output a value estimate of the current state; Based on the action policy and the value estimate, use the deep reinforcement learning algorithm to update the parameters of the policy network.

[0017] In an exemplary embodiment of the present disclosure, the value network uses the privileged state information to output a value estimate of the current state, including: The value network uses the privileged state information and combines the first latent vector, the fourth latent vector, and the second latent vector or the third latent vector selected by the preset policy to output a value estimate of the current state.

[0018] In an exemplary embodiment of the present disclosure, based on the received first latent vector, fourth latent vector, the second latent vector or the third latent vector selected by the preset policy, and the current motion state information, the policy network outputs an action policy for controlling the movement of the robot, including: The policy network is based on: Output an action policy for controlling the movement of the robot; Wherein, represents the probability distribution of taking action in the current state MLP represents a multi-layer perceptron, represents the second latent vector or the third latent vector selected according to the preset policy, represents the first latent vector, represents the fourth latent vector.

[0019] In an exemplary embodiment of the present disclosure, selecting the second latent vector or the third latent vector according to the preset policy includes: Configure a first selection probability for the second latent vector and a second selection probability for the third latent vector; Randomly select from the second latent vector and the third latent vector according to the first selection probability and the second selection probability.

[0020] In an exemplary embodiment of the present disclosure, the method further includes: Dynamically increase the ratio of the first selection probability to the second selection probability as the training time increases; wherein, at the initial stage of training, the ratio of the first selection probability to the second selection probability is less than 1.

[0021] In an exemplary embodiment of the present disclosure, when training a policy network for outputting an action policy for controlling the movement of a robot by using a teacher-student model framework based on deep reinforcement learning, it further includes: Optimize the parameters of the student encoder based on the difference between the second latent vector and the third latent vector.

[0022] In an exemplary embodiment of the present disclosure, optimizing the parameters of the student encoder based on the difference between the second latent vector and the third latent vector includes: Calculate the difference between the second latent vector and the third latent vector through a supervised learning module; Construct a loss function according to the difference; Update the parameters of the student encoder through backpropagation according to the loss function, so that the second latent vector output by the student encoder gradually approaches the third latent vector output by the teacher encoder.

[0023] In an exemplary embodiment of the present disclosure, updating the parameters of the policy network by using a deep reinforcement learning algorithm based on the action policy and value estimation includes: Obtain a reward signal generated based on environmental interaction and the current state value output by the value network, and construct an advantage function based on the reward signal and the current value estimation ; According to the ratio of the current action policy output by the policy network to the old policy and the advantage function , construct an objective function for proximal policy optimization : wherein, is a preset hyperparameter for restricting the amplitude of policy update; represents restricting within the interval , represents the expected value; Perform gradient update on the parameters of the policy network through the objective function for proximal policy optimization.

[0024] In an exemplary embodiment of the present disclosure, updating the parameters of the policy network by using a deep reinforcement learning algorithm based on the action policy and value estimation further includes: Set an entropy coefficient for the objective function for proximal policy optimization; Decay the entropy coefficient every preset time interval.

[0025] In an exemplary embodiment of the present disclosure, setting an entropy coefficient for the objective function optimized by proximal policy includes: According to: Calculate the total loss; Wherein, is the total loss, is the entropy coefficient that decays over time, represents the entropy of the action distribution of the current policy at the current state of the current policy

[0026] According to the second aspect of the present disclosure, a robot motion control method based on deep reinforcement learning is provided, including: Encoding the historical linear velocity information of the robot through a linear velocity encoder to generate a first latent vector, and inputting the first latent vector into the policy network; Encoding the historical motion state information of the robot itself through a student encoder to generate a second latent vector, and inputting it into the policy network; Based on the received latent vectors and the current motion state information, the policy network outputs an action policy for controlling the motion of the robot; Wherein, the student encoder, the linear velocity encoder, and the policy network are obtained by the method for training a robot motion control model based on deep reinforcement learning in the above embodiment.

[0027] In an exemplary embodiment of the present disclosure, the method further includes: Encoding the historical elevation map information of the robot through a topographic map encoder to generate a fourth latent vector, and inputting the fourth latent vector into the policy network; Wherein, the student encoder, the linear velocity encoder, the topographic map encoder, and the policy network are obtained by the method for training a robot motion control model based on deep reinforcement learning in the above embodiment.

[0028] According to the third aspect of the present disclosure, a device for training a robot motion control model based on deep reinforcement learning is provided, including: A policy network training module for training a policy network for outputting an action policy for controlling the motion of the robot by using a teacher-student model framework; wherein, the training process of the policy network includes: Encoding the historical linear velocity information of the robot through a linear velocity encoder to generate a first latent vector, and inputting the first latent vector into the policy network for auxiliary training.

[0029] According to the fourth aspect of the present disclosure, a robot motion control device based on deep reinforcement learning is provided, including: A linear velocity feature encoding module, configured to encode historical linear velocity information of the robot to generate a first latent vector, and input the first latent vector into a policy network; A student encoder encoding module, configured to encode historical motion state information of the robot itself to generate a second latent vector, and input it into the policy network; A policy network inference module, configured to output an action policy for controlling the motion of the robot based on the received latent vectors and the current motion state information; Wherein, the student encoder, the linear velocity encoder, and the policy network are obtained according to the robot motion control model training method based on deep reinforcement learning in the above-mentioned embodiments.

[0030] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: A processor; and A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the methods in the above-mentioned embodiments are implemented.

[0031] According to a sixth aspect of the present disclosure, there is provided a robot, including: A processor; and A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the methods in the above-mentioned embodiments are implemented.

[0032] In an exemplary embodiment of the present disclosure, the robot includes any one of a legged robot, a quadruped robot, a biped robot, a wheeled robot, a wheel-legged robot, a four-wheel-legged robot, a humanoid robot, a cleaning robot, a transportation robot, a mobile robot, and a robotic arm.

[0033] According to a seventh aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program code instructions are stored, and when the computer program code instructions are called by a processor of the robot, the robot is caused to execute the methods in the above-mentioned embodiments.

[0034] From the above technical solutions, it can be seen that the present disclosure has at least one of the following advantages and positive effects: The robot motion control model training method based on deep reinforcement learning provided in the exemplary embodiment of the present disclosure encodes the robot's own historical motion state information through the student encoder to generate a first potential vector, and encodes the robot's privileged state information through the teacher encoder to generate a second potential vector, and selects the first potential vector or the second potential vector according to the preset strategy to input into the policy network, and encodes the robot's historical linear velocity information through the linear velocity encoder to generate a third potential vector, and inputs it into the policy network, so that the policy network can optimize the control decision based on the historical velocity information. By introducing the linear velocity encoder to encode the robot's historical linear velocity information and through the synergy of the student encoder, the teacher encoder and the linear velocity encoder, and adding the motion trend information in the time dimension to the policy network, the control strategy can more accurately predict the robot's future motion state and make a control decision that is more in line with the law of motion. Therefore, compared with the traditional reinforcement learning method that only relies on the current state to make decisions, the present disclosure can improve the sensitivity of the policy network to speed changes by supplementing the state representation with linear velocity information, thereby reducing the control lag phenomenon in an environment with high-speed motion or drastic speed changes, improving the response speed and stability of the control strategy, and improving the adaptability of the robot in a high-speed or high-dynamic environment. In addition, since the linear velocity information directly reflects the changing trend of the robot's motion state, the present disclosure provides additional information support through the linear velocity encoder during the training phase, so that the policy network can learn the control strategy that adapts to the speed change more quickly, thereby reducing the randomness in the exploration process and improving the training efficiency. Compared with the traditional method that requires a large amount of data for training to adapt to different speed scenarios, the present disclosure can make the control strategy adapt to different speed states more quickly, improve the convergence speed of learning, and reduce the demand for training data. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0036] Figure 1 A system architecture diagram is shown to which the robot motion control model training method based on deep reinforcement learning and the robot motion control method based on deep reinforcement learning in the embodiments of the present disclosure can be applied.

[0037] Figure 2 A flow chart of a robot motion control model training method based on deep reinforcement learning in an embodiment of the present disclosure is shown.

[0038] Figure 3 Shows a schematic diagram of a two-stage training-inference framework for a policy network in an embodiment of the present disclosure.

[0039] Figure 4 Shows a schematic diagram of the process of a pre-trained policy network in an embodiment of the present disclosure.

[0040] Figure 5 Shows a schematic diagram of the process of updating the parameters of a student encoder in an embodiment of the present disclosure.

[0041] Figure 6 Shows a schematic diagram of another two-stage training-inference framework for a policy network in an embodiment of the present disclosure.

[0042] Figure 7 Shows a schematic diagram of yet another two-stage training-inference framework for a policy network in an embodiment of the present disclosure.

[0043] Figure 8 Shows a schematic diagram of the process of another pre-trained policy network in an embodiment of the present disclosure.

[0044] Figure 9 Shows a schematic diagram of yet another two-stage training-inference framework for a policy network in an embodiment of the present disclosure.

[0045] Figure 10 Shows a schematic diagram of the process of a robot motion control method based on deep reinforcement learning in an embodiment of the present disclosure.

[0046] Figure 11 Shows a block diagram of a robot motion control model training device based on deep reinforcement learning in an embodiment of the present disclosure.

[0047] Figure 12 Shows a block diagram of a robot motion control device based on deep reinforcement learning in an embodiment of the present disclosure.

[0048] Figure 13 Shows a schematic diagram of a robot in an embodiment of the present disclosure.

[0049] Figure 14 Shows another schematic diagram of a robot in an embodiment of the present disclosure.

[0050] Figure 15 Shows yet another schematic diagram of a robot in an embodiment of the present disclosure.

[0051] Figure 16 Shows a schematic diagram of the structure of an electronic device in an embodiment of the present disclosure.

[0052] Figure 17 Shows a schematic diagram of the structure of a program product in an embodiment of the present disclosure. Detailed implementation manners

[0053] In the description of the present disclosure, the terms "first" and "second" are only used for description and do not indicate relative importance or imply the number of technical features. Therefore, the features of "first" and "second" may explicitly or implicitly include at least one of such features. The meaning of "a plurality" is at least two, unless otherwise clearly defined.

[0054] Figure 1 The system architecture diagram to which the method for training a robot motion control model based on deep reinforcement learning in the embodiments of the present disclosure can be applied is shown.

[0055] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a robot 102, a network 103, and a server 104. Among them, the terminal device 101 includes, but is not limited to, a desktop computer, a portable computer, a smart phone, a tablet computer, and the like. The terminal device 101 may serve as an interaction interface, providing a visualization function to display the operating state, motion trajectory, etc. of the robot 102, and at the same time supporting sending motion control instructions to the robot 102.

[0056] A variety of sensors are mounted on the robot 102 to collect its own motion state information under different tasks, and it can execute an action policy under the control of the server 104 to complete corresponding training actions or tasks. The server 104 deploys a training framework based on deep reinforcement learning, including modules such as a teacher-student model structure, a policy network, a value network, and a linear velocity encoder. During the training process, the server 104 receives historical linear velocity information from the robot 102, generates a first latent vector through the linear velocity encoder and inputs it into the policy network, combines the policy output with the current state value estimated by the value network, and uses a deep reinforcement learning algorithm to optimize the parameters of the policy network. The server 104 can also undertake the supervised learning task in the teacher-student model, and perform latent vector alignment through the teacher encoder and the student encoder to guide the student encoder to learn an effective state representation.

[0057] After the training is completed, the server 104 can deploy the trained policy network to the robot 102 to realize its autonomous inference and execution of the action policy during actual operation. The terminal device 101 can also be used to call the deployed policy, monitor the motion state of the robot, or remotely schedule and control.

[0058] The network 103 is used to provide a medium for communication links between the terminal device 101, the robot 102, and the server 104. The network 103 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. It should be understood, Figure 1The number and type of terminal devices, robots, networks, and servers therein are merely illustrative. According to the implementation requirements, there can be any number and any type of terminal devices, robots, networks, and servers.

[0059] The exemplary embodiments of the present disclosure provide a method for training a robot motion control model based on deep reinforcement learning. In this method, a teacher-student model framework is used to perform deep reinforcement learning training on a policy network for outputting an action policy for controlling the motion of a robot.

[0060] Among them, the teacher-student model framework includes a student encoder and a teacher encoder. Through the two collaborative structures of the teacher encoder and the student encoder, the policy network is guided to obtain higher-quality state representations and action outputs during the deep reinforcement learning process. At the same time, this model framework provides a structured knowledge transfer path, enabling the policy network to not only learn the optimal policy from the environment but also accelerate convergence and improve sample efficiency with the guidance of the teacher encoder. In the exemplary embodiments of the present disclosure, deep reinforcement learning includes, but is not limited to, algorithms such as PPO (Proximal Policy Optimization), DDPG (Deep Deterministic Policy Gradient), and SAC (Soft Actor-Critic). The present disclosure does not limit the specific algorithm types of deep reinforcement learning.

[0061] In some exemplary embodiments, referring to Figure 2 as shown, the training process of the policy network includes step S201 and step S202: Step S201, encoding the historical linear velocity information of the robot through a linear velocity encoder to generate a first latent vector, and inputting the first latent vector into the policy network.

[0062] Among them, linear velocity is the distance that a point moves along a straight or curved trajectory per unit time. In the exemplary embodiments of the present disclosure, the linear velocity information is the displacement velocity of the robot changing with time in three-dimensional space. For example, the linear velocity information of the robot is represented by the spatial velocity of the center of mass or the center of the torso of the robot, which is used to characterize the overall motion trend and dynamic state of the robot. The linear velocity encoder is used to extract motion trend features with dynamic continuity and time dependence from the linear velocity information of the robot at multiple historical moments and encode them into a latent representation suitable for input to the policy network. For example, the linear velocity encoder can be a one-dimensional convolutional neural network, a recurrent neural network, a multi-layer perceptron, etc. The present disclosure does not limit the specific type of the linear velocity encoder.

[0063] Exemplarily, the linear velocity encoder may include an input layer, a hidden layer, and an output layer. First, the input layer receives the linear velocity information of the robot at multiple historical moments, and preprocesses the linear velocity information, such as performing operations like normalization and dimension alignment, to ensure the stability and effectiveness of subsequent network calculations. Then, the hidden layer uses a neural network model to extract temporal features from the preprocessed linear velocity information, and can effectively capture the dynamic patterns and acceleration / deceleration trends during the speed change process. For example, the neural network model can be a long short-term memory neural network or a gated recurrent unit neural network, which can retain important information from past moments and model time correlations. Finally, the output layer maps the temporal features extracted by the hidden layer into a latent vector with a fixed dimension, that is, the first latent vector.

[0064] The first latent vector is fed into the policy network, which can provide auxiliary information in the time dimension for the policy network, enhancing the policy's perception ability of speed changes, enabling the policy network to not only rely on the current state but also perceive past motion trajectories and speed changes, and helping to improve the stability and responsiveness of the control policy in high-speed motion or dynamic change scenarios.

[0065] Step S202, based on the value estimation of the current state output by the value network in the teacher-student model framework and the action policy output by the policy network combined with the first latent vector, use the deep reinforcement learning algorithm to update the parameters of the policy network.

[0066] Among them, the value network is used to estimate the long-term return of the current state, that is, starting from this state, if the current policy is continuously executed, how much cumulative reward can be obtained in the future. Specifically, after receiving the first latent vector, the policy network outputs the current action policy, and at the same time combines the value network in the teacher-student model framework to estimate the value of the current state. A loss function is constructed through the deep reinforcement learning algorithm, and the policy output and value estimation are used to jointly optimize the policy network. For example, through the policy gradient and backpropagation mechanism, the parameters of the policy network are updated, thereby continuously improving the control performance of the policy in the current environment. Taking the PPO algorithm as an example, the "clip ratio" strategy is adopted to limit the update amplitude each time to improve the training stability and sample utilization rate.

[0067] In this step, the teacher encoder provides stable feature supervision for the student encoder, guiding the student encoder to learn higher-quality state representations, enabling the policy network to have better convergence speed and generalization ability during training. Importantly, by introducing the linear velocity encoder, the policy network gains the ability to model the motion trend in the time dimension, thereby improving the prediction accuracy of future states, making action decisions more in line with the dynamic laws of the robot body, further enhancing the response speed and stability of the control policy in high-speed or dynamic environments. At the same time, combined with the global evaluation information provided by the value network and the guidance of the teacher model, the overall training efficiency, policy quality, and adaptability in complex task scenarios are improved.

[0068] Reference Figure 3 As shown, a schematic diagram of a two-stage training-inference framework for a policy network is presented. Among them, the training stage combines a teacher-student model framework including a teacher encoder 301 and a student encoder 302, a linear velocity encoder 303, and a deep reinforcement learning framework. The deep reinforcement learning framework is jointly composed of a policy network 304, a value network 305, and the PPO algorithm. During the training stage, by introducing privileged state information and historical linear velocity information to guide policy learning, the student encoder 302 can stably output high-quality control actions depending only on observable data during the inference stage. The observable data includes the current motion state information of the robot.

[0069] It should be noted that the teacher encoder 301 is only used to extract the high-dimensional semantic representation in the privileged state information during the training stage for subsequent use as input to the policy to guide learning. The student encoder 302 encodes the historical motion state information of the robot itself during both the training and inference stages, and the linear velocity encoder 303 encodes the historical linear velocity information of the robot during both the training and inference stages to obtain dynamic features related to action decisions. Then, together with the current motion state information, it is input into the policy network 304 to generate control actions. In addition, the student encoder 302 ultimately needs to imitate the representation output by the teacher encoder 301, so during the training process, distillation training is performed by minimizing the MSE (mean squared error).

[0070] The policy network 304 is the core module for generating action policies. During the training stage, the policy network 304 receives the latent vectors from the teacher encoder 301 or the student encoder 302, the latent vector from the linear velocity encoder 303, and the current motion state information, and combines with the value network 305 to update the policy using the PPO algorithm. During the inference stage, the policy network 304 can receive the latent vectors from the student encoder 302 and the linear velocity encoder 303 and the current motion state information, and output the action policy for controlling the robot's motion.

[0071] Based on Figure 3 the framework schematic diagram shown, reference Figure 4As shown, the training process of the policy network for outputting an action policy for controlling the movement of a robot based on deep reinforcement learning may include the following steps S401 to S405: Step S401: Encode the historical motion state information of the robot itself through a student encoder to generate a second latent vector, and encode the privileged state information of the robot through a teacher encoder to generate a third latent vector.

[0072] Among them, the historical motion state information of the robot itself may include joint angles, joint velocities, sole contact states, center-of-gravity trajectories, etc. of several past frames. For example, the historical motion state sequence of the past 10 frames is input into the student encoder 302 for encoding to obtain a second latent vector. The second latent vector can capture the continuity and dynamic characteristics of the robot's actions.

[0073] The privileged state information may include the actual contact force of the robot with the ground, the environmental height map, disturbance information, etc. The privileged state information is input into the teacher encoder 301 for encoding to obtain a third latent vector. The third latent vector can highly condense key semantics such as terrain structure, obstacle distribution, and the interaction between the robot and the environment, which helps to construct a more complete high-dimensional feature representation of the environmental state.

[0074] It should be noted that the privileged state information can only be obtained during the training phase but is not observable during the testing or deployment phase. Therefore, the teacher encoder 301 can use the complete information to learn the best latent representation, thereby guiding the student encoder 302 to learn. Both the teacher encoder 301 and the student encoder 302 can compress high-dimensional, time-series state information into low-dimensional latent representations, providing behavioral semantic representations from different sources for the policy network 304. For example, the teacher encoder 301 and the student encoder 302 can be multi-layer perceptrons, temporal convolutional networks, or recurrent neural networks, which have the ability to extract temporal features and compress them into fixed-length semantic vectors. In addition, the network architectures of the teacher encoder 301 and the student encoder 302 can be the same or different, and the present disclosure does not limit this.

[0075] Step S402: Select the second latent vector or the third latent vector according to a preset policy and input it into the policy network.

[0076] According to a preset policy, select the second latent vector generated by the student encoder 302 or the third latent vector generated by the teacher encoder 301 and send it to the policy network 304 for decision-making. Among them, the preset policy can be set according to the real-time environmental state, the training phase, or specific performance metrics, and the specific performance metrics include action execution error thresholds, simulation and real environment difference degrees, etc.

[0077] Exemplarily, a first selection probability can be configured for the second latent vector and a second selection probability can be configured for the third latent vector, and a random selection can be made from the second latent vector and the third latent vector according to the first selection probability and the second selection probability. It should be noted that, as the training time increases, the ratio of the first selection probability to the second selection probability can also be dynamically increased; wherein, at the initial stage of training, the ratio of the first selection probability to the second selection probability is less than 1.

[0078] For example, at the initial stage of training, the high-quality third latent vector generated by the teacher encoder 301 is preferentially used to guide the policy network 304 to quickly converge to an approximate optimal solution. When the student encoder 302 is optimized through knowledge distillation, it gradually transitions to only using the second latent vector to reduce the dependence on privileged information. Accordingly, the third latent vector can be selected according to the second selection probability p, the second latent vector can be selected according to the first selection probability 1 - p, and the value of p can be gradually reduced to dynamically adjust the ratio between the first selection probability 1 - p and the second selection probability p. It can be understood that at the initial stage of training, the third latent vector needs to be preferentially selected. Therefore, the second selection probability p is greater than the first selection probability 1 - p, and at this time, the ratio of the first selection probability 1 - p to the second selection probability p is less than 1.

[0079] Of course, the second latent vector and the third latent vector can also be concatenated and then sent into the policy network 304, and dynamically weighted through an attention mechanism or a gating module, so that the policy network 304 can flexibly combine historical experience and privileged knowledge in complex scenarios.

[0080] The selective input mechanism can not only utilize the prior knowledge of the ideal state provided by the teacher encoder 301 to accelerate the training process, but also, during actual operation, cope with problems such as sensor limitations or environmental disturbances through the generalization ability of the student encoder 302. At the same time, through the conditional processing of the latent vector by the policy network 304, smooth switching and robust decision-making of motion control are achieved. For example, when the robot encounters unknown terrain, it preferentially adjusts its gait based on the second latent vector, and relies on the first latent vector to maintain efficiency during the stable walking stage. Finally, the policy selection rule is optimized through closed-loop feedback, enabling the robot to balance motion performance and adaptability under different stages and environmental conditions.

[0081] In addition, in addition to the selected latent vector, the first latent vector and the current motion state information of the robot can also be input into the policy network 304 for decision-making.

[0082] Step S403, based on the received first latent vector, the second latent vector or the third latent vector selected by the preset policy, and the current motion state information, the policy network outputs an action policy for controlling the movement of the robot.

[0083] Among them, the policy network 304 can be a multi-layer perceptron, a Transformer structure, etc. The policy network 304 performs multi-modal feature fusion on the second latent vector or the third latent vector, the first latent vector, and the current motion state information, and outputs an action policy, denoted as .

[0084] For example, the policy network can be based on: (1) Output an action policy for controlling the movement of the robot; Among them, represents the probability distribution of taking action in the current state , MLP represents a multi-layer perceptron, represents the second latent vector or the third latent vector selected according to a preset policy, represents the first latent vector.

[0085] Again, for example, the second latent vector or the third latent vector, the first latent vector, and the current motion state information are concatenated or weighted and interacted in the embedding space, and action policies such as joint angle targets, torque commands, or gait phase parameters are generated through non-linear transformation.

[0086] Step S404, using the privileged state information through the value network, output the value estimate of the current state; During the process of training the policy network 304, the long-term return of the current state is also estimated through the value network 305 to optimize the policy network 304.

[0087] Among them, the value network 305 can be a multi-layer fully connected perceptron. For example, when using the privileged state information through the value network 305 to estimate the value of the current state, the privileged state information is first encoded into a high-dimensional feature vector, for example, dynamic features are extracted through convolutional or fully connected layers, and then fused with the current motion state information in the latent space, and then the value estimate representing the current state is output through multi-layer non-linear transformation, denoted as . This value estimate is used for policy optimization in reinforcement learning to guide the policy network 304 to learn better behaviors.

[0088] Step S405, based on the action policy and the value estimate, use the deep reinforcement learning algorithm to update the parameters of the policy network.

[0089] Taking the PPO algorithm as an example for illustration. The reward signal generated based on environmental interaction and the current state value output by the value network can be obtained, and the advantage function is constructed based on the reward signal and the current value estimate, and the advantage function Indicates the superiority or inferiority of the current action relative to the average policy.

[0090] Exemplarily, during the training process of the policy network 304, the current policy network 304 is used to interact with the environment. Based on the current state Select an action policy , and after executing the action, return the reward and the next state , from which an interaction trajectory sequence ( ) is sampled for subsequent policy optimization. Then, the value network 305 is introduced to evaluate the current value of each state .

[0091] Based on this, the constructed advantage function can be: (2) where is the discount factor, which is used to measure the importance of future rewards.

[0092] Furthermore, according to the ratio of the current action policy output by the policy network 304 to the old policy and the advantage function , the objective function of proximal policy optimization is constructed : (3) where is a preset hyperparameter used to limit the amplitude of policy update; means to limit within the interval , represents the expected value. Formula (3) ensures that the policy does not grow too fast when the advantage function is positive and does not decrease too much when the advantage function is negative, thus making a "proximal" constraint on policy update.

[0093] Finally, the parameters of the policy network 304 are updated by gradient using the objective function of proximal policy optimization. For example, the loss value is calculated using the objective function shown in formula (3), and then the gradient is calculated through the backpropagation algorithm, and the parameters of the policy network 304 are iteratively updated in combination with the optimizer. This update process can guide the policy network 304 to gradually learn the action policy that maximizes the long-term cumulative reward while maintaining a stable output, thereby improving the robustness and execution efficiency of the control policy in the real environment.

[0094] In some example embodiments, by introducing an entropy term into the objective function, the policy network 304 can be encouraged to output an action distribution with high uncertainty at the initial stage of training, thereby enhancing the comprehensive exploration of the action space and preventing the policy from prematurely falling into a local optimum. However, if a large entropy weight is maintained throughout, the policy will not be able to fully converge to a stable action. Therefore, during the training process of the policy network 304, an entropy coefficient can also be set for the objective function of proximal policy optimization, and the entropy coefficient can be decayed after every preset time interval.

[0095] For example, it can be calculated according to: (4) to obtain the total loss; where is the total loss, is the entropy coefficient that decays over time, represents the entropy of the action distribution of the current policy at the current state , which is used to measure the uncertainty of the policy's selection of each action.

[0096] In this example, on the one hand, the policy diversity and sample efficiency at the initial stage of training are improved, enabling the model to more fully explore the environmental state space; on the other hand, by reducing the entropy weight in the later stage of training, the policy tends to be deterministic, enhancing control stability and execution consistency. Generally speaking, it can help the policy network achieve a dynamic balance between exploration and convergence during the training process, improving the final training effect and the reliability of policy deployment.

[0097] In addition, based on Figure 4 the training process of the policy network for outputting an action policy for controlling the movement of a robot based on deep reinforcement learning, the training process further includes optimizing the parameters of the student encoder based on the difference between the second latent vector and the third latent vector. This step quantifies the distribution difference between the two in the latent space and backpropagates the difference gradient to update the network weights of the student encoder 302, thereby guiding the student encoder 302 to learn to generate a latent representation close to that of the teacher encoder 301 and achieving teacher knowledge distillation.

[0098] Exemplarily, during the training phase, the parameters of the teacher encoder 301 are fixed. After using the historical motion state information and the corresponding privileged state information as parallel inputs to generate latent vectors through the student encoder 302 and the teacher encoder 301 respectively, the distance between the two is minimized using a contrastive learning framework, or the output distribution of the student encoder 302 is made to approximate the latent space characteristics of the teacher encoder 301 through adversarial training. At the same time, noise injection or data augmentation is introduced to simulate sensor errors in actual deployment, forcing the student encoder 302 to still extract feature expressions compatible with the privileged information encoding under limited input information conditions.

[0099] In some example embodiments, with reference to Figure 5 as shown, based on the difference between the second latent vector and the third latent vector, the process of optimizing the parameters of the student encoder may include the following steps S501 to S503: Step S501, calculate the difference between the second latent vector and the third latent vector through a supervised learning module.

[0100] Among them, the supervised learning module receives the second latent vector and the third latent vector and performs a calculation representing the difference. It should be noted that the supervised learning module in step S501 is not a specific neural network structure, but a training mechanism for measuring the proximity between the output of the student encoder 302 and the output of the teacher encoder 301 in the feature space. For example, vector distance metrics such as mean squared error, cosine distance, or L1 norm are used to quantify the difference between the two latent vectors.

[0101] The difference result will be used as the basis for the supervised loss function in step S502 to train the student encoder 302, enabling the student encoder 302 to generate a latent state representation similar to that of the teacher encoder 301 even without privileged information, thereby achieving high-quality representation knowledge transfer and providing effective support for the decision-making of the subsequent policy network 304.

[0102] Step S502, construct a loss function based on the difference.

[0103] Correspondingly, the loss function can adopt mean squared error loss, L1 loss, cosine similarity loss, or contrast loss, etc., to reflect the similarity between the representation generated by the student encoder 302 and the high-quality semantic features output by the teacher encoder 301. The smaller the difference, the more successfully the student encoder 302 learns the feature extraction ability of the teacher encoder 301.

[0104] The constructed loss function will act on the parameter update process of the student encoder 302 through the error backpropagation mechanism in step S503, guiding the student encoder 302 to learn a more task-relevant and representative latent state vector under the condition of no privileged information, thereby improving the input quality and overall performance of the policy network 304.

[0105] Step S503, update the parameters of the student encoder through backpropagation according to the loss function, so that the second latent vector output by the student encoder gradually approaches the third latent vector output by the teacher encoder.

[0106] Using the error backpropagation algorithm, the parameters of the student encoder are updated according to the differential loss function, so that the second latent vector output by the student encoder gradually approaches the second latent vector output by the teacher encoder. Through continuous optimization, the student encoder learns to extract feature expressions close to the privileged information from the historical motion states, thereby enhancing the generalization ability of the policy network 304, improving the decision-making quality of the policy network 304 in the real environment, and enabling it to approach the teacher level without relying on privileged information during deployment.

[0107] Reference Figure 6 As shown, a schematic diagram of a two-stage training-inference framework of another policy network is presented. During the training stage, the policy network 304 receives the latent vector from the teacher encoder 301 or the student encoder 302, the first latent vector from the linear velocity encoder 303, and the current motion state information, and combines with the value network 305 to perform policy updates using the PPO algorithm. During the inference stage, the policy network 304 can receive the second latent vector from the student encoder 302, the first latent vector from the linear velocity encoder 303, and the current motion state information, and output an action policy for controlling the movement of the robot.

[0108] Based on Figure 6 the schematic diagram of the framework shown, the training process of the policy network is Figure 4 similar to the training process shown, except that the step of using the privileged state information to output the value estimate of the current state through the value network 305 in this training process is different from the implementation method of step S404. Specifically, in Figure 6 the training stage shown, the value network 305 uses the privileged state information, and combines with the first latent vector, the second latent vector or the third latent vector selected by the preset policy, to output the value estimate of the current state.

[0109] In this example, fusing the latent features and the privileged information and inputting them into the value network 305 not only retains the high credibility of the privileged information for environmental understanding, but also enhances the adaptability of the value estimate to the policy behavior trajectory and the motion trend, and can better reflect the long-term reward expectation that the robot can obtain after taking actions in the current state. In addition, fusing the latent vector input also improves the representational richness and generalization ability of the value network, making its output more consistent with the value distribution under the real policy behavior, thereby improving the stability and performance of the overall policy training.

[0110] In some exemplary embodiments, Figure 2 the training process of the policy network shown further includes encoding the historical elevation map information of the robot by a topographic map encoder to generate a fourth latent vector, and inputting the fourth latent vector into the policy network for auxiliary training.

[0111] Among them, the historical elevation map information refers to the elevation map information collected by the robot at several historical moments. The elevation map information is sourced from lidar, depth cameras, or other 3D sensors, and can reflect spatial features such as the undulation, slope, and obstacle positions of the terrain where the robot is located. The topographic map encoder is used to extract the temporal evolution features of the terrain from the historical elevation map information for assisting in the training of the policy network. For example, the topographic map encoder can be a convolutional neural network, a vision Transformer architecture, or other models suitable for image and point cloud processing, which can encode the original spatial information in the elevation map into a latent vector with a fixed dimension.

[0112] Exemplarily, the topographic map encoder can include an input layer, a hidden layer, and an output layer. First, the input layer receives elevation maps at multiple historical moments, such as two-dimensional or three-dimensional raster maps, and performs preprocessing operations such as size unification, pixel normalization, and occlusion area filling to ensure the standardization and stability of the network input. Next, the hidden layer extracts features from the preprocessed elevation map, capturing information such as local structural changes, landform features, spatial gradients, and their temporal evolution in the terrain. For example, convolutional structures can extract local spatial features, and methods such as temporal convolution or stacking input sequences can retain the dynamic evolution patterns of terrain changes, thus obtaining the ability to understand the terrain with temporal correlation. Finally, the output layer maps the spatio-temporal features extracted in the hidden layer into a latent vector with a fixed length, that is, the fourth latent vector.

[0113] The fourth latent vector, as an abstract representation of the terrain cognitive information, is fed into the policy network and participates in the generation and training of the policy together with other latent vectors, enabling the policy network to refer to the change trends and structural features of the historical terrain when making motion decisions, thereby possessing stronger environmental adaptability and path planning robustness. When facing complex terrains, non-flat ground, or local obstacles, the terrain latent vector provided by the topographic map encoder can help the policy network better understand the constraints and guiding effects of the terrain on motion, improving the generalization ability and response efficiency of the policy in the real environment.

[0114] Reference Figure 7 As shown, a schematic diagram of a two-stage training-inference framework for another policy network is shown. Figure 7In it, a topographic map encoder 306 is introduced. The topographic map encoder 306 encodes the historical elevation map information of the robot to obtain terrain features related to action decisions. In the training phase, the policy network 304 receives the latent vector from the teacher encoder 301 or the student encoder 302, the first latent vector of the linear velocity encoder 303, the fourth latent vector of the topographic map encoder 306, and the current motion state information, and combines with the value network 305 to update the policy using the PPO algorithm. In the inference phase, the policy network 304 can receive the latent vector from the student encoder 302, the latent vector of the linear velocity encoder 303, the latent vector of the topographic map encoder 306, and the current motion state information, and output an action policy for controlling the movement of the robot.

[0115] Based on Figure 7 the schematic diagram of the framework shown, refer to Figure 8 as shown, the training process of the policy network based on deep reinforcement learning may include the following steps S801 to step S805: Step S801, encode the historical motion state information of the robot itself through the student encoder to generate a second latent vector, and encode the privileged state information of the robot through the teacher encoder to generate a third latent vector.

[0116] Step S802, select the second latent vector or the third latent vector according to a preset policy and input it into the policy network.

[0117] Step S803, based on the received first latent vector, fourth latent vector, the preset policy to select the second latent vector or the third latent vector, and the current motion state information, the policy network outputs an action policy for controlling the movement of the robot.

[0118] Exemplarily, the policy network can be based on: (5) output an action policy for controlling the movement of the robot; where represents the probability distribution of taking action in the current state , MLP represents a multi-layer perceptron, represents the second latent vector or the third latent vector selected according to the preset policy, represents the first latent vector, represents the fourth latent vector.

[0119] Step S804, use the privileged state information through the value network to output the value estimate of the current state; Step S805, based on the action policy and the value estimate, use the deep reinforcement learning algorithm to update the parameters of the policy network.

[0120] It should be noted that, except for step S803, the specific details of other steps can be referred to Figure 4 the shown training process of the policy network and supplementary steps. For example, the specific description of the step of optimizing the student encoder parameters based on the difference between the second latent vector and the third latent vector will not be elaborated here.

[0121] Refer to Figure 9 shown, which shows a schematic diagram of a two-stage training-inference framework of another policy network. In Figure 9 the shown training stage, the value network 305 utilizes privileged state information, and combines the latent vector from the teacher encoder 301 or the student encoder 302, the first latent vector of the linear velocity encoder 303, and the fourth latent vector of the topographic map encoder 306 to output the value estimate of the current state.

[0122] In this example, the fourth latent vector generated by the topographic map encoder 306 further introduces semantic information such as the terrain structure and obstacle layout in the spatial environment, enhancing the sensitivity of the value estimate to the spatial context and terrain complexity. The fusion of multi-source latent features enables the value network 305 to not only estimate based on the global privileged state, but also combine the temporal features and environment adaptation features strongly related to the policy behavior, thereby improving the accuracy, stability, and generalization ability of the current state value evaluation, and can more precisely guide the policy update direction in reinforcement learning optimization, improving the sample efficiency and training convergence speed.

[0123] The exemplary embodiment of the present disclosure also provides a robot motion control method based on deep reinforcement learning. Refer to Figure 10 shown, this method may include the following steps S1001 to step S1003: Step S1001, encoding the historical linear velocity information of the robot through a linear velocity encoder to generate a first latent vector, and inputting the first latent vector into the policy network.

[0124] The system encodes the linear velocity data of the robot at multiple historical moments through the linear velocity encoder, extracts the temporal features of the speed change during the movement process, and generates a first latent vector. This latent vector contains dynamic trend information such as acceleration, deceleration, and direction change of the robot within a short time window, which helps the policy network understand the motion context in which the current action is located.

[0125] Step S1002, encoding the historical motion state information of the robot itself through a student encoder to generate a second latent vector, and inputting it into the policy network.

[0126] The student encoder receives the observable state information of the robot at multiple historical time steps, such as joint angles, angular velocities, IMUs, foot contact states, etc., extracts structural motion semantics and generates a second latent vector, which reflects the state evolution trajectory of the robot itself and is an abstract expression of the internal motion process.

[0127] Step S1003, based on the received latent vectors and the current motion state information, the policy network outputs an action policy for controlling the motion of the robot.

[0128] The policy network simultaneously receives the first latent vector, the second latent vector, and the current motion state information, performs joint encoding and fusion reasoning, and outputs the current action policy for controlling the motion of the robot. Through the joint input of multi-source latent features and the current state, the policy network not only has the ability to perceive the immediate state, but also can understand the historical continuity and dynamic background of the action, so as to output a smoother, more responsive, and more environmentally adaptable control policy, showing higher stability and intelligence in complex terrains or high-dynamic scenarios.

[0129] In some exemplary embodiments, the historical elevation map information of the robot can also be encoded by a topographic map encoder to generate a fourth latent vector, and the fourth latent vector is input into the policy network.

[0130] The topographic map encoder receives the elevation map data observed by the robot in the past several time steps. The topographic map encoder extracts the spatial distribution pattern of the elevation map data and generates a fourth latent vector, which encodes structural information such as terrain undulations, surface continuity, and environmental change trends on the historical path.

[0131] The fourth latent vector, together with the first latent vector from the linear velocity encoder, the second latent vector from the student encoder, and the current motion state information, is input into the policy network. The policy network can comprehensively consider the motion trend, state change trajectory of the robot itself, and the surrounding terrain structure during action decision-making, and realize the output of a more robust and terrain-adaptive control policy.

[0132] It can be understood that the student encoder, the linear velocity encoder, the topographic map encoder, and the policy network are trained according to the robot motion control model training method based on deep reinforcement learning described in detail in other embodiments of the present disclosure, which will not be elaborated here.

[0133] In the exemplary embodiments of the present disclosure, a robot motion control model training device based on deep reinforcement learning is also provided. Refer to Figure 11As shown in the figure, the training device 1100 for a robot motion control model based on deep reinforcement learning includes a policy network training module 1101, which is used to train a policy network for outputting an action policy for controlling the motion of the robot by using a teacher-student model framework; wherein, the training process of the policy network includes: Encoding the historical linear velocity information of the robot through a linear velocity encoder to generate a first latent vector, and inputting the first latent vector into the policy network for auxiliary training.

[0134] The specific details of the above-mentioned policy network training module 1101 have been described in detail in the corresponding method for training a robot motion control model based on deep reinforcement learning, so they will not be elaborated here.

[0135] In the exemplary embodiment of the present disclosure, a robot motion control device based on deep reinforcement learning is also provided. Refer to Figure 12 As shown in the figure, the robot motion control device 1200 based on deep reinforcement learning includes a linear velocity feature encoding module 1201, a student encoder encoding module 1202, and a policy network inference module 1203, wherein: The linear velocity feature encoding module 1201 is used to encode the historical linear velocity information of the robot to generate a first latent vector, and input the first latent vector into the policy network; The student encoder encoding module 1202 is used to encode the historical motion state information of the robot itself to generate a second latent vector, and input it into the policy network; The policy network inference module 1203 is used to output an action policy for controlling the motion of the robot based on the received latent vectors and the current motion state information; Wherein, the student encoder, the linear velocity encoder, and the policy network are obtained according to the method for training a robot motion control model based on deep reinforcement learning in the embodiments of the present disclosure.

[0136] The specific details of each module in the above-mentioned robot motion control device based on deep reinforcement learning have been described in detail in the corresponding robot motion control method based on deep reinforcement learning, so they will not be elaborated here.

[0137] In the exemplary embodiment of the present disclosure, a robot is also provided. The robot includes a processor and a memory, and a computer-readable instruction is stored on the memory. When the computer-readable instruction is executed by the processor, the above method is implemented. Wherein, the robot includes any one of a legged robot, a quadruped robot, a biped robot, a wheeled robot, a wheel-legged robot, a four-wheel-legged robot, a humanoid robot, a cleaning robot, a transportation robot, a mobile robot, and a robotic arm. Refer to Figures 13 to 15 As shown in the figure, schematic diagrams of three different robots are respectively shown.

[0138] Reference Figure 16 As shown, an electronic device capable of implementing the above method is also provided. Among them, the electronic device 1600 includes a processor 1601 and a memory 1602. A computer-readable instruction is stored on the memory 1602. When the computer-readable instruction is executed by the processor 1601, the method for training a robot motion control model based on deep reinforcement learning in the embodiments of the present disclosure is implemented.

[0139] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, on which computer program code instructions are stored. When the computer program code instructions are called by a processor of a robot, the robot is enabled to execute the method in the embodiment.

[0140] Reference Figure 17 As shown, a program product 1700 for implementing the above method for training a robot motion control model based on deep reinforcement learning according to an embodiment of the present disclosure is described. It can adopt a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0141] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described here can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0142] Finally, the above preferred embodiments are only used to illustrate the technical solution of the present application and are not restrictive. Although the present application has been described in detail, those skilled in the art should understand that changes in form and details can be made to it without departing from the scope defined by the claims of the present application. The dimensions of the drawings have nothing to do with the specific physical object, and the physical dimensions can be arbitrarily changed.

Claims

1. A robot motion control model training method based on deep reinforcement learning, characterized in that: include: Using the teacher-student model framework, a policy network for outputting an action strategy for controlling robot motion is trained based on deep reinforcement learning; wherein the training process of the policy network includes: Encoding the historical linear velocity information of the robot through a linear velocity encoder to generate a first latent vector, and inputting the first latent vector into the policy network; Based on the value estimate of the current state output by the value network in the teacher-student model framework and the action strategy output by the policy network in combination with the first latent vector, the parameters of the policy network are updated using a deep reinforcement learning algorithm.

2. The robot motion control model training method based on deep reinforcement learning according to claim 1 is characterized in that: The teacher-student model framework includes a student encoder and a teacher encoder; The strategy network for outputting the action strategy for controlling the robot motion is trained based on deep reinforcement learning by using the teacher-student model framework, including: The student encoder is used to encode the historical motion state information of the robot itself and generate a second latent vector, and the teacher encoder is used to encode the privileged state information of the robot and generate a third latent vector; Selecting a second latent vector or a third latent vector according to a preset strategy and inputting the second latent vector into the strategy network; Outputting, by the strategy network, a motion strategy for controlling the motion of the robot based on the received first potential vector, the second potential vector or the third potential vector selected by the preset strategy, and current motion state information; Utilizing the privileged state information through a value network, outputting a value estimate of the current state; Based on the action strategy and the value estimate, the parameters of the policy network are updated using a deep reinforcement learning algorithm.

3. The robot motion control model training method based on deep reinforcement learning according to claim 2 is characterized in that: The step of utilizing the privileged state information through the value network to output a value estimate of the current state includes: The value network utilizes privileged state information and combines the first latent vector and the second latent vector or the third latent vector selected by the preset strategy to output a value estimate of the current state.

4. The robot motion control model training method based on deep reinforcement learning according to claim 1 is characterized in that: The linear velocity encoder comprises: An input layer, used for receiving linear velocity information of the robot at multiple historical moments and preprocessing the linear velocity information; The hidden layer is used to extract time series features from the preprocessed linear velocity information using a neural network model; The output layer is used to generate the first latent vector according to the temporal features extracted by the hidden layer.

5. The robot motion control model training method based on deep reinforcement learning according to claim 4 is characterized in that: The neural network model in the hidden layer includes a long short-term memory neural network or a gated recurrent unit neural network.

6. The robot motion control model training method based on deep reinforcement learning according to claim 2 is characterized in that: The outputting, through the strategy network, an action strategy for controlling the movement of the robot based on the received first potential vector, the second potential vector or the third potential vector selected by the preset strategy, and the current movement state information, comprises: The policy network is based on: Outputting an action strategy for controlling the movement of the robot; in, Indicates the current state Take action The probability distribution of MLP stands for multi-layer perceptron. represents the second latent vector or the third latent vector selected according to the preset strategy, Denote the first latent vector.

7. The robot motion control model training method based on deep reinforcement learning according to claim 1 is characterized in that: The training process of the policy network also includes: The historical elevation map information of the robot is encoded by a terrain map encoder to generate a fourth latent vector, and the fourth latent vector is input into the policy network for auxiliary training.

8. The robot motion control model training method based on deep reinforcement learning according to claim 7 is characterized in that: The teacher-student model framework includes a student encoder and a teacher encoder; The strategy network for outputting the action strategy for controlling the robot motion is trained based on deep reinforcement learning by using the teacher-student model framework, including: The student encoder is used to encode the historical motion state information of the robot itself and generate a second latent vector, and the teacher encoder is used to encode the privileged state information of the robot and generate a third latent vector; Selecting a second latent vector or a third latent vector according to a preset strategy and inputting the second latent vector into the strategy network; Selecting the second latent vector or the third latent vector and the current motion state information through the strategy network based on the received first latent vector, the fourth latent vector, the preset strategy, and outputting an action strategy for controlling the motion of the robot; Utilizing the privileged state information through a value network, outputting a value estimate of the current state; Based on the action strategy and the value estimate, the parameters of the policy network are updated using a deep reinforcement learning algorithm.

9. The robot motion control model training method based on deep reinforcement learning according to claim 8 is characterized in that: The step of utilizing the privileged state information through the value network to output a value estimate of the current state includes: The value network utilizes privileged state information and combines the first latent vector, the fourth latent vector, and the second latent vector or the third latent vector selected by the preset strategy to output a value estimate of the current state.

10. The robot motion control model training method based on deep reinforcement learning according to claim 8, characterized in that: The step of selecting the second potential vector or the third potential vector and the current motion state information based on the received first potential vector, the fourth potential vector, and the preset strategy through the strategy network, and outputting the motion strategy for controlling the motion of the robot includes: The policy network is based on: Outputting an action strategy for controlling the movement of the robot; in, Indicates the current state Take action The probability distribution of MLP stands for multi-layer perceptron. represents the second latent vector or the third latent vector selected according to the preset strategy, represents the first latent vector, Denotes the fourth latent vector.

11. The robot motion control model training method based on deep reinforcement learning according to claim 2 or 8, characterized in that: The selecting the second latent vector or the third latent vector according to a preset strategy includes: configuring a first probability of selection for the second latent vector and configuring a second probability of selection for the third latent vector; The second latent vector and the third latent vector are randomly selected according to the first selection probability and the second selection probability.

12. The robot motion control model training method based on deep reinforcement learning according to claim 10, characterized in that: The method further comprises: According to the increase of training time, the ratio of the first selection probability to the second selection probability is dynamically increased; wherein, in the initial stage of training, the ratio of the first selection probability to the second selection probability is less than 1.

13. The robot motion control model training method based on deep reinforcement learning according to claim 2 or 8, characterized in that: The method of using the teacher-student model framework to train a policy network for outputting an action policy for controlling the motion of the robot based on deep reinforcement learning also includes: Optimizing parameters of the student encoder based on a difference between the second latent vector and the third latent vector.

14. The robot motion control model training method based on deep reinforcement learning according to claim 13 is characterized in that: The optimizing the parameters of the student encoder based on the difference between the second latent vector and the third latent vector comprises: Calculate the difference between the second latent vector and the third latent vector through a supervised learning module; constructing a loss function based on the difference; The parameters of the student encoder are updated by back propagation according to the loss function, so that the second latent vector output by the student encoder gradually approaches the third latent vector output by the teacher encoder.

15. The robot motion control model training method based on deep reinforcement learning according to claim 2 or 8, characterized in that: The updating of the parameters of the strategy network using a deep reinforcement learning algorithm based on the action strategy and the value estimation includes: Obtain a reward signal generated based on the interaction with the environment and the current state value output by the value network, and construct an advantage function based on the reward signal and the current value estimate ; The ratio of the current action strategy output by the policy network to the old strategy And the advantage function , construct the objective function of proximal strategy optimization : in, It is a preset hyperparameter used to limit the magnitude of strategy updates; Indicates that Restricted to the interval Inside, Indicates expected value; The parameters of the policy network are gradient updated through the objective function of the proximal policy optimization.

16. The robot motion control model training method based on deep reinforcement learning according to claim 15, characterized in that: The updating of the parameters of the strategy network using a deep reinforcement learning algorithm based on the action strategy and the value estimation also includes: Setting an entropy coefficient for the objective function of the proximal strategy optimization; After each preset time interval, the entropy coefficient is decayed.

17. The robot motion control model training method based on deep reinforcement learning according to claim 16, characterized in that: The step of setting an entropy coefficient for the objective function of the proximal strategy optimization includes: according to: Calculate the total loss; in, is the total loss, is the entropy coefficient that decays over time, Indicates the current state Next, the current strategy The entropy of the action distribution.

18. A robot motion control method based on deep reinforcement learning, characterized in that: include: Encoding the historical linear velocity information of the robot through a linear velocity encoder to generate a first latent vector, and inputting the first latent vector into a policy network; Encoding the robot's own historical motion state information through a student encoder to generate a second latent vector, and inputting the second latent vector into the policy network; Outputting, through the strategy network, an action strategy for controlling the movement of the robot based on the received potential vectors and the current motion state information; Among them, the student encoder, linear velocity encoder and strategy network are obtained according to the robot motion control model training method based on deep reinforcement learning according to any one of claims 1 to 17.

19. The robot motion control method based on deep reinforcement learning according to claim 18, characterized in that: The method further comprises: encoding the historical elevation map information of the robot through a terrain map encoder to generate a fourth latent vector, and inputting the fourth latent vector into the policy network; Among them, the student encoder, linear velocity encoder, topographic map encoder and strategy network are obtained according to the robot motion control model training method based on deep reinforcement learning according to any one of claims 1 to 17.

20. A robot motion control model training device based on deep reinforcement learning, characterized in that: include: A policy network training module is used to train a policy network for outputting an action strategy for controlling the motion of the robot using a teacher-student model framework; wherein the training process of the policy network includes: The historical linear velocity information of the robot is encoded by a linear velocity encoder to generate a first latent vector, and the first latent vector is input into the policy network for auxiliary training.

21. A robot motion control device based on deep reinforcement learning, characterized in that: include: A linear velocity feature encoding module, used for encoding the historical linear velocity information of the robot to generate a first latent vector, and inputting the first latent vector into the policy network; A student encoder encoding module, used for encoding the robot's own historical motion state information to generate a second potential vector, and inputting the second potential vector into the policy network; A strategy network reasoning module, used for outputting an action strategy for controlling the movement of the robot based on the received potential vectors and the current movement state information; Among them, the student encoder, linear velocity encoder and strategy network are obtained according to the robot motion control model training method based on deep reinforcement learning according to any one of claims 1 to 17.

22. An electronic device, characterized in that: include: processor; as well as A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by the processor, implement the method according to any one of claims 1 to 19.

23. A robot, characterized in that: include: processor; as well as A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by the processor, implement the method according to any one of claims 1 to 19.

24. The robot according to claim 23, characterized in that The robot includes any one of a legged robot, a quadruped robot, a bipedal robot, a wheeled robot, a wheel-legged robot, a quadrupedal robot, a humanoid robot, a cleaning robot, a transport robot, a mobile robot and a robotic arm.

25. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program code instructions, and when the computer program code instructions are called by a processor of the robot, the robot executes the method as described in any one of claims 1 to 19.

Citation Information

Patent Citations

  • Foot type robot control method and system, computer equipment and storage medium

    CN116931475A

  • Method and device for determining robot control strategy model, readable medium and program product

    CN119458315A

  • Learning robust legged robot locomotion with implicit terrain imagination via deep reinforcement learning

    KR1020250009137A

  • Mechanical arm path planning method based on velocity smoothing deterministic policy gradient

    CN110328668A

  • Motion control method and system for quadruped robot under terrain subareas

    CN118192254A

Cited By

  • Hexapod robot leg and arm multiplexing control method and device based on reinforcement learning

    CN120993711A

  • Quadruped robot control method and system based on Transform reinforcement learning architecture

    CN121115515A

  • Motor spiral push mechanism calibration control method and system based on transfer learning

    CN121143054A

  • Motor spiral pushing mechanism calibration control method and system based on transfer learning

    CN121143054B

  • Robot motion control strategy network training method and device based on deep reinforcement learning, robot motion control method and device, equipment, robot and storage medium

    CN121361098A