Robot motion control method and system
Through the BMPC model, the robot multi-dimensional sensor data is mapped into low-dimensional feature vectors, combined with environmental dynamics and reward models, the strategy network is guided to generate efficient action sequences, solving the problem of low robot learning efficiency in high-dimensional tasks, and achieving rapid convergence and low computing cost control effects.
Patent Information
- Application Number
- CN202510561710.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-01
AI Technical Summary
The efficiency and performance of robot reinforcement learning in the existing technology is insufficient in high-dimensional continuous control tasks, expert iterative methods and guided strategy search are difficult to effectively carry out, and model-based reinforcement learning does not perform well in complex tasks.
Using the BMPC model, the robot multi-dimensional sensor data is mapped into low-dimensional hidden space feature vectors through a state encoder, and the environmental dynamics model is used to predict the future state of the candidate action sequence. The reward model evaluates immediate rewards and long-term returns. The strategy network guides the generation of candidate action sequences, and accelerates convergence through expert strategies.
It realizes efficient strategy learning for robots in high-dimensional tasks, significantly reduces training steps and computing resource consumption, and improves the intelligence of control actions and task completion efficiency.
Smart Images

Figure CN120395832A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot motion control and relates to a robot motion control method and system. Background Art
[0002] With the rapid development of artificial intelligence technology, reinforcement learning, as an effective machine learning method, has made remarkable progress in solving robot motion control tasks. Reinforcement learning enables an agent to learn the optimal policy through interactions with the environment, adapt to various complex environments, and demonstrate strong adaptive capabilities.
[0003] In the research of reinforcement learning, the method of combining planning or models has become a popular direction. These methods aim to improve the efficiency and performance of reinforcement learning by introducing planning algorithms or environment models. Among them, expert iteration, guided policy search, and model-based reinforcement learning are several relatively common technical solutions.
[0004] The expert iteration method combines planning and policy learning, using a planning algorithm to generate high-quality expert data to guide policy learning. This method performs well in discrete tasks such as chess, but in tasks with high-dimensional continuous action spaces, its efficiency often drops significantly. This is because the tree search algorithm is difficult to effectively search in continuous action spaces, resulting in difficulties in generating expert data, and the computational cost of large-scale planning is also relatively high.
[0005] The guided policy search method uses the trajectories generated by an external planner to guide the neural network policy learning. Although this method can use the trajectories generated by the planner to help the network policy better explore the state space in complex tasks, the observation spaces of the planner and the neural network are usually different, resulting in the planner being unable to fully utilize the feedback of the network policy. Therefore, it is difficult for the guided policy search to reverse-guide the planner through the network policy to achieve the progressive improvement of the network policy.
[0006] The model-based reinforcement learning method introduces an environment model to simulate future states and rewards, thereby improving data efficiency. This method can perform efficient policy learning in the case of insufficient data, but its applicable scope is limited. For example, DreamerV3 is more suitable for vision-based, discrete tasks and does not perform well in complex high-dimensional continuous control tasks. Although TD-MPC2 combines MPC (Model Predictive Control) and temporal difference learning and performs well in continuous control tasks, its policy learning still has deficiencies. When the network policy learning is insufficient, the value estimation of the Q function will deviate, further affecting the effect of MPC planning, resulting in the agent being unable to effectively plan and execute in high-dimensional tasks. Summary of the Invention
[0007] The object of the present invention is to solve the technical problem that in the prior art, when a robot performs reinforcement learning, the prior art is insufficient in efficiency and performance in high-dimensional continuous control tasks, and to provide a robot motion control method and system.
[0008] To achieve the above object, the present invention adopts the following technical solutions: The first aspect of the present invention provides a robot motion control method, including the following steps: Collect the current state information of the robot; Input the current state information of the robot into the BMPC model to obtain the current control action of the robot; The BMPC model includes a state encoder, an environmental dynamics model, a reward model, a policy network, and a value network; the current state information of the robot is encoded into a latent space feature vector by the state encoder; Based on the latent space feature vector, a part of the candidate action sequences are selected through the policy network, and another part of the candidate action sequences are sampled from a multi-dimensional Gaussian distribution; The environmental dynamics model predicts the future states of all candidate action sequences, and the reward model predicts the reward values of the future states; Calculate the cumulative rewards of all action sequences based on the predicted reward values of the future states, recalculate the parameters of the Gaussian distribution through the action sequence with the highest cumulative reward; return to the step of selecting a part of the candidate action sequences through the policy network until the iteration step is completed; Select the first action of the action sequence with the highest cumulative reward as the current control action, and record it as the expert policy at the same time; The robot executes the current control action; Repeat the above steps until the preset task is completed.
[0009] Further, the current state information of the robot includes the positions, speeds, angular velocities of the respective joints of the robot, as well as the body posture information and the contact information with the ground.
[0010] Further, the reward model includes sparse rewards and dense rewards.
[0011] Further, the expert policy is described as:
[0012] Wherein, is the action distribution in the expert behavior dataset, represents the state encoder, represents the policy network.
[0013] Further, the loss function of the policy network is:
[0014] Among them, is the action distribution in the expert behavior dataset, is the network policy, is the Kullback-Leibler divergence between the expert policy and the network policy; represents the entropy term of the policy, is the time weight factor, is the adjustment coefficient of the entropy term.
[0015] Furthermore, the cumulative reward is specifically:
[0016] Among them, is the discount factor; represents the reward model, is the current state information of the robot, is the expert policy corresponding to the current state information of the robot.
[0017] The second aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned robot motion control method is implemented.
[0018] The third aspect of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned robot motion control method is implemented.
[0019] The fourth aspect of the present invention provides a computer program product, which includes computer instructions. It is characterized in that the computer instructions instruct a computer to execute the above-mentioned robot motion control method.
[0020] The fifth aspect of the present invention provides a robot motion control system, including: A data acquisition module that acquires the current state information of the robot and multiple action sequences; A robot action prediction module that inputs the current state information of the robot into the BMPC model to obtain the control action of the robot; The BMPC model includes a state encoder, an environmental dynamics model, a reward model, a policy network, and a value network; Based on the latent space feature vector, a part of the candidate action sequences are selected through the policy network, and another part of the candidate action sequences are sampled from a multi-dimensional Gaussian distribution; The environmental dynamics model predicts the future states of all candidate action sequences, and the reward model predicts the reward values of the future states; Calculate the cumulative reward for all action sequences based on the predicted future state rewards, and recalculate the parameters of the Gaussian distribution through the action sequence with the highest cumulative reward; return the step of selecting a part of the candidate action sequences through the policy network until the iteration step is completed; Select the first action of the action sequence with the highest cumulative reward as the current control action, and record it as the expert policy at the same time; Execution module, the robot executes the current control action; Iteration module, repeat the above steps until the preset task is completed.
[0021] Compared with the prior art, the present invention has the following beneficial effects: The present invention discloses a robot motion control method, which maps multi-dimensional sensor data of the robot (such as position, speed, environmental features, etc.) into a low-dimensional hidden space feature vector through a state encoder to achieve efficient fusion and dimensionality reduction processing of multi-modal information; uses an environmental dynamics model to predict the future state of each candidate action sequence frame by frame to generate a multi-step state trajectory; the reward model evaluates the immediate reward and long-term return of each action based on the future state trajectory, and the policy network guides the generation of candidate action sequences; through reward-guided policy learning, the robot can dynamically adapt to environmental changes, and at the same time uses the expert policy to accelerate convergence and improve the intelligence and task completion efficiency of the control action; calculates the cumulative reward of the action sequence based on the multi-step reward value, and selects the first action of the optimal sequence as the current control instruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a schematic diagram of the overall framework of the proposed BMPC model; Figure 2 It is the working flow chart of the BMPC model proposed by the present invention; Figure 3 It is the embodiment diagram of the robot controlled in 28 tasks for testing the BMPC model; Figure 4 It is the test result of the BMPC model and the comparison with the results of other algorithms; Figure 5 It is the test result of the training time of the BMPC model and the comparison with the results of other algorithms. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and marked in the accompanying drawings here can be arranged and designed in various different configurations.
[0025] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0026] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0027] The present invention will be further described in detail below with reference to the accompanying drawings: See Figure 1 , the present invention discloses a robot motion control method, including the following steps: Step 1: Construct a world model, a policy network, and a value network. The world model includes a state encoder, an environmental dynamics model, and a reward model; and train the above models and networks. First, let the robot interact with the environment, collect motion data, and use a reinforcement learning algorithm for training to update the parameters of the neural network (including the state encoder, the environmental dynamics model, the reward model, the policy network, and the value network) to maximize the reward. After training, a trained neural network is obtained. The specific process is as Figure 2 shown.
[0028] Specifically: S1, first perform task setting and reward design: The task is set for the robot to run and maintain balance, and it is necessary to precisely control the leg movements to reach the target speed while avoiding falling or losing balance.
[0029] The reward design can select the following reward methods: Sparse reward: When the robot successfully maintains balance and reaches the target speed, a positive reward is given, such as +10. If the robot falls or deviates too much from the target speed, a negative reward is given, such as -5.
[0030] Dense Reward: Rewards are given based on the proximity of the robot to the target speed. For each action step, the reward value is calculated according to the error from the target speed. For example, a +0.1 reward is given when the error is within a certain range, and a -0.1 penalty is given when the error is too large. At the same time, according to the balance state of the robot, such as a +0.05 reward is given when the body tilt angle is within a reasonable range, and a -0.05 penalty is given when it exceeds the range.
[0031] S2, at each time step of the robot , according to the currently observed state , through the state encoder encode the state into a latent space feature vector.
[0032] Based on the current network policy , the robot selects an action and executes it. The environment returns the next state and the immediate reward . At the same time, record the expert policy at this moment .
[0033] Store the five-tuple into the experience replay pool .
[0034] S3, update the parameters of each neural network. Specifically: For the world model, the parameter update method is as follows: Randomly sample a batch of samples from the experience replay pool . .
[0035] Use the state encoder to encode the sampled state to obtain a latent state representation. Then, using the environmental dynamics model and the reward model , predict the next latent state and reward according to the sampled action, and update by minimizing the error between the predicted value and the actual observed value , and . For example, for the environmental dynamics model , the optimization goal is to minimize the difference between the predicted next latent state and the actually observed next latent state (obtained by encoding ).
[0036] For the value network, the parameter update method is as follows: Update the value function based on the model-based online objective .
[0037] According to the currently sampled state-action sequence, use the world model to predict the future The state, action, and reward of the step. Calculate the target value according to the following formula . Then update the parameters of the value function by minimizing the mean squared error between the predicted value of the value function and the target value .
[0038] The calculation of the target value is as follows:
[0039] where is the discount factor, used to measure the attenuation of future rewards; is the preset step size of the TD target; is the target value estimate after a certain number of steps, obtained from the target value network; is the reward estimate at the current moment.
[0040] For the policy network , the parameter update method is as follows:
[0041] where is the action distribution in the expert behavior dataset, is the network policy, is the Kullback-Leibler divergence between the expert policy and the network policy. is the policy loss. Among them, the divergence measures the difference between the expert policy and the network policy, and the entropy term is used to encourage the exploration of the policy. Update the policy network parameters by minimizing the policy loss through the gradient descent algorithm .
[0042] For the target value network, use exponential moving average to update the target value network
[0043] Use the exponential moving average method to update the parameters of the target value network , so that it slowly tracks the parameter changes of the value function to provide a more stable target value estimate.
[0044] S4, every time after the preset reanalysis interval steps, draw a batch of samples from the experience replay pool . For the drawn samples, recalculate the expert policy , and update the corresponding policy distribution in the experience replay pool . This step reduces the computational cost of frequent replanning by using existing high-quality samples, while maintaining the effectiveness of policy learning.
[0045] Step 2: When actually controlling the robot, the following steps are adopted: S1. Obtain the current state information of the robot.
[0046] The sensors of the robot collect the current state information of the robot in real time. These state information include the positions, speeds, angular speeds of each joint, as well as the body posture information (such as the inclination angle of the body) and the contact information with the ground, etc. These data are used as the current state .
[0047] S2. Input the current state information of the robot into the BMPC model to obtain the control actions of the robot; the BMPC model includes a state encoder, an environmental dynamics model, a reward model, a policy network, and a value network.
[0048] S201. First, through the state encoder encode the current state into a latent space feature vector .
[0049] S202. Based on the latent space feature vector, select a part of the candidate action sequences through the policy network, and at the same time sample from a multi-dimensional Gaussian distribution to obtain another part of the candidate action sequences; The environmental dynamics model predicts the future states of all candidate action sequences, and the reward model predicts the reward values of the future states; Calculate the cumulative rewards of all action sequences based on the predicted reward values of the future states, recalculate the parameters of the Gaussian distribution through the action sequence with the highest cumulative reward; return to the step of selecting a part of the candidate action sequences through the policy network until the iteration step is completed; Select the first action of the action sequence with the highest cumulative reward as the current control action, and record it as the expert policy at the same time.
[0050] S3. Output a control signal Convert the control action at the selected current moment to the corresponding control signal at a frequency of about 50Hz and output it to the actuator of the robot to control the joint movement of the robot and achieve running and balance control.
[0051] S4. Continuously loop.
[0052] In the next time step, repeat the above process, continuously update the control actions according to the sensor data, and continuously perform motion control on the robot until the task is completed or the preset termination condition is reached.
[0053] To illustrate the superiority of the present invention, the proposed algorithm BMPC was experimentally verified in this embodiment. Multiple complex continuous control tasks were mainly conducted in the DMControl environment, and the experimental results were analyzed and interpreted. The experimental tasks covered a wide range of high-dimensional motion control scenarios, including various sparse reward tasks and high-complexity motion control tasks. The effectiveness and data efficiency improvement of the proposed BMPC algorithm in these tasks were verified through experiments.
[0054] DMControl is a standard simulation environment for continuous control tasks and is widely used in the research of reinforcement learning and control algorithms. This environment provides multiple challenging simulation tasks, covering from simple basic motion tasks to complex high-dimensional motion control tasks, and can effectively verify the performance of reinforcement learning algorithms. In the experiment, 28 different continuous control tasks were mainly selected, covering a variety of high-dimensional motion control scenarios. In these tasks, all embodied entities controlled by BMPC are as Figure 3 shown.
[0055] These tasks include: Walker Run: The agent controls a bipedal walking robot to run and maintain balance. The agent needs to learn to precisely control the leg movements to reach the target speed while avoiding falling or losing balance.
[0056] Humanoid Run: This is a high-dimensional motion control task where the agent needs to control a humanoid robot to run. Due to the high degrees of freedom and complex movements of the humanoid robot, this task poses extremely high requirements on the agent's control ability.
[0057] Dog Run: The agent controls a quadruped dog-like robot to run. The high-dimensional state and action space of this task require the agent to be able to precisely control each joint to maintain stability and achieve fast movement.
[0058] In the DMControl environment, the action space design for different tasks varies according to the complexity of the tasks and the degrees of freedom of the agents. Generally speaking, the action space defines the set of operations that an agent can take at each time step. The actions in DMControl tasks are usually a continuous multi-dimensional vector, and the dimensions correspond to the controllable joints on the agent. For example, the Walker (biped robot) has multiple joints (such as hip joints, knee joints, etc.), and the action of each joint can be the force or torque applied to that joint. The agent controls the actions of these joints to achieve the goal of walking or running. Humanoid robots have more degrees of freedom and a higher-dimensional action space. Each joint (such as shoulder joints, elbow joints, hip joints, etc.) can receive a continuous action input, usually force or torque. The agent runs by simultaneously controlling multiple joints and coordinating the movement of the whole body.
[0059] The observation space defines the information that a robot can perceive at each time step. In DMControl, the observation space usually includes the state information of the agent itself and the key information in the environment. For example, the positions, velocities, angular velocities of each joint, etc. In addition, the agent can also observe its own posture (such as the tilt angle of the body) and the contact information with the ground. This information helps the agent maintain balance during walking or running. The observation space of humanoid robots is more complex and usually includes the angles, velocities, angular velocities of multiple joints, and the posture information of the body (such as the rotation, tilt, acceleration of the body, etc.). The agent needs to use this information to judge its own balance state and make corresponding adjustments.
[0060] In DMControl, the reward design is usually customized according to the task goals. Most tasks use sparse rewards or a combination of sparse rewards and dense rewards. In the sparse reward design, the agent only receives a reward when it completes the task (such as reaching the target position or achieving a specific target state). For example, in the Walker Run task, the agent only receives a positive reward when it successfully maintains balance and reaches the target speed. To avoid the low learning efficiency of the agent caused by a completely sparse reward signal, dense rewards are sometimes designed to provide more feedback. For example, the reward can be designed according to the proximity of the agent to the target position, and each step of the agent's action will receive a reward or punishment according to its distance from the target.
[0061] The verification results are as follows: In all experimental tasks, BMPC demonstrated a significant improvement in sample efficiency compared to existing baseline algorithms such as TD-MPC2, SAC, and DreamerV3. Especially in high-dimensional motion control tasks (such as Humanoid Run and Dog Run), the data efficiency improvement of BMPC was particularly evident. In these tasks, the guided policy learning mechanism of BMPC helped the agent quickly converge to an efficient policy, significantly reducing the number of environment interaction steps required for training. The complete training curves are shown as Figure 4 shown.
[0062] For example, in the Humanoid Run task, the sample utilization efficiency of BMPC was approximately 300% higher than that of TD-MPC2. BMPC learns network policies by mimicking MPC policies, enabling the agent to achieve a higher task success rate with fewer environment steps. This improvement in data efficiency has been verified in multiple high-dimensional tasks, indicating that BMPC has high applicability in complex control tasks.
[0063] BMPC greatly enhances the stability of policy learning by combining MPC policies with model-based temporal difference learning. In multiple high-dimensional continuous control tasks, the policy learning process of BMPC exhibits the characteristic of stable convergence and can achieve a high task success rate within a short training time.
[0064] In tasks such as Walker Run, the network policy and MPC policy of BMPC ultimately showed similar performance, indicating that through guided policy learning, BMPC can significantly reduce the dependence on MPC online planning during the inference process. In some tasks, the network policy of BMPC can even replace the MPC policy for task execution, effectively reducing the computational resource consumption during the inference process.
[0065] For example, in the Dog Walk task, BMPC obtained a task reward of 900 within the first 400,000 training steps, while in contrast, other baseline algorithms required 2 million or more training steps to achieve similar performance. This result indicates that BMPC can quickly achieve a high task success rate through a stable policy learning process in complex tasks.
[0066] BMPC significantly reduces the computational resource consumption during inference and training by introducing a lazy reanalysis mechanism. Traditional reinforcement learning methods, especially model-based planning methods, usually require frequent replanning, which significantly increases the computational cost. BMPC reduces the frequency of replanning by maintaining an expert behavior dataset and performing reanalysis regularly, thereby reducing the computational resource consumption.
[0067] In the experiment, the inert reanalysis mechanism of BMPC reduced the training time for solving tasks by approximately 50% compared to baseline algorithms such as TD-MPC2. The specific statistical results are as Figure 5 shown. Especially in the later stages of the task, when the network policy can partially replace the MPC policy, BMPC reduces its dependence on online planning, further reducing the computational burden. This enables BMPC to achieve efficient policy learning at a lower computational cost in multiple complex tasks.
[0068] An embodiment of the present invention provides a robot motion control system, comprising: A data acquisition module that acquires the current state information of the robot and multiple action sequences; A robot action prediction module that inputs the current state information of the robot into the BMPC model to obtain the control actions of the robot; The BMPC model includes a state encoder, an environmental dynamics model, a reward model, a policy network, and a value network; Based on the latent space feature vector, a part of the candidate action sequences are selected through the policy network, and another part of the candidate action sequences are sampled from a multi-dimensional Gaussian distribution; The environmental dynamics model predicts the future states of all candidate action sequences, and the reward model predicts the reward values of the future states; Calculate the cumulative rewards of all action sequences based on the predicted reward values of the future states, recalculate the parameters of the Gaussian distribution through the action sequence with the highest cumulative reward; return to the step of selecting a part of the candidate action sequences through the policy network until the iteration step is completed; Select the first action of the action sequence with the highest cumulative reward as the current control action, and record it as the expert policy at the same time; An execution module, where the robot executes the current control action; An iteration module that repeats the above steps until the preset task is completed.
[0069] In another embodiment of the present invention, a terminal device is provided. The terminal device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function. The processor described in the embodiments of the present invention can be used for the operation of the robot motion control method.
[0070] In another embodiment of the present invention, a storage medium is also provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is the memory device in the terminal device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and, of course, the extended storage medium supported by the terminal device. It can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. The computer-readable storage medium provides storage space, and this storage space stores the operating system of the terminal. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that more specific examples (non-exhaustive list) of the computer-readable storage medium here include: electrical connections with one or more wires, portable disks, hard disks, Random Access Memories (RAMs), Read Only Memories (ROMs), Erasable Programmable Read Only Memories (EPROMs or flash memories), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0071] The computer-readable storage medium also includes a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0072] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0073] One or more instructions stored in the computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the robot motion control method in the above embodiments.
[0074] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A robot motion control method, characterized in that, It includes the following steps: Collect the current state information of the robot; Input the current state information of the robot into the BMPC model to obtain the current control action of the robot; The BMPC model includes a state encoder, an environment dynamics model, a reward model, a policy network, and a value network; the current state information of the robot is encoded into a latent space feature vector by the state encoder; Based on the latent space feature vector, select a part of the candidate action sequences through the policy network, and at the same time sample another part of the candidate action sequences from the multi-dimensional Gaussian distribution; The environment dynamics model predicts the future states of all candidate action sequences, and the reward model predicts the reward values of the future states; Calculate the cumulative rewards of all action sequences based on the predicted reward values of the future states, and recalculate the parameters of the Gaussian distribution through the action sequence with the highest cumulative reward; Return to the step of selecting a part of the candidate action sequences through the policy network until the iteration step is completed; Select the first action of the action sequence with the highest cumulative reward as the current control action, and record it as the expert policy at the same time; The robot executes the current control action; Repeat the above steps until the preset task is completed.
2. The robot motion control method according to claim 1, wherein The current state information of the robot includes the positions, speeds, angular speeds of each joint of the robot, as well as the body posture information and the contact information with the ground.
3. The robot motion control method according to claim 1, characterized in that, The reward model includes sparse rewards and dense rewards.
4. The robot motion control method according to claim 1, wherein The expert policy is described as: Among them, is the action distribution in the expert behavior dataset, represents the state encoder, represents the policy network.
5. The robot motion control method according to claim 1, characterized in that, The loss function of the policy network is: Among them, is the action distribution in the expert behavior dataset, is the network policy, is the Kullback-Leibler divergence between the expert policy and the network policy; represents the entropy term of the policy, is the time weight factor, is the adjustment coefficient of the entropy term.
6. The robot motion control method according to claim 1, characterized in that, The cumulative reward is specifically: Among them, is the discount factor; represents the reward model, is the current state information of the robot, is the expert policy corresponding to the current state information of the robot.
7. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the robot motion control method according to any one of claims 1-6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the robot motion control method according to any one of claims 1-6.
9. A computer program product, the computer program product comprising computer instructions, characterized in that, The computer instructions instruct the computer to execute the robot motion control method according to any one of claims 1-6.
10. A robot motion control system, based on the robot motion control method according to claim 1, characterized in that, It includes: A data acquisition module that acquires the current state information of the robot and multiple action sequences; A robot action prediction module that inputs the current state information of the robot into the BMPC model to obtain the control action of the robot; The BMPC model includes a state encoder, an environment dynamics model, a reward model, a policy network, and a value network; Based on the latent space feature vector, select a part of the candidate action sequences through the policy network, and at the same time sample another part of the candidate action sequences from the multi-dimensional Gaussian distribution; The environment dynamics model predicts the future states of all candidate action sequences, and the reward model predicts the reward values of the future states; Calculate the cumulative rewards of all action sequences based on the predicted reward values of the future states, and recalculate the parameters of the Gaussian distribution through the action sequence with the highest cumulative reward; Return to the step of selecting a part of the candidate action sequences through the policy network until the iteration step is completed; Select the first action of the action sequence with the highest cumulative reward as the current control action, and record it as the expert policy at the same time; An execution module where the robot executes the current control action; An iteration module that repeats the above steps until the preset task is completed.
Citation Information
Patent Citations
Robot operation method based on visual reasoning
CN109159113A
Control strategy offline training method based on model uncertainty and behavior prior
CN115972211A
Off-line learning for robot control using a reward prediction model
US20230256593A1