Determination Method, Device, Readable Medium and Program Product of Robot Control Strategy Model

By training robot control strategies in simulation and adapting them using stage-specific transfer algorithms, the method addresses the scarcity of real-world training samples, enhancing model efficiency and accuracy in real-world applications.

CN119458315BActive Publication Date: 2025-07-15BEIJING RUIZHEN INNOVATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411479680.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-22
Publication Date
2025-07-15
Estimated Expiration
2044-10-22

AI Technical Summary

Technical Problem

In the prior art, due to the lack of training samples, the robot control strategy model has poor intelligence performance and insufficient adaptability in real application scenarios.

Method used

Training reinforcement learning strategy models in the simulation environment, adapting them to the real environment through migration algorithms of multiple links, including environment perception, strategy control and task execution links. Various migration algorithms such as domain adaptation, domain randomization, inverse dynamics model, etc. are used to improve the adaptability of the model in the real environment.

Benefits of technology

The training efficiency, accuracy and generalization of the robot control strategy model are improved, ensuring accuracy and adaptability in real environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119458315B_ABST
    Figure CN119458315B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for determining a robot control strategy model, a method and device for determining a robot control strategy, a computer-readable storage medium, and a program product. A specific implementation manner of this application includes: training a reinforcement learning strategy model for determining the control strategy of a simulated robot in a simulation environment; for multiple links involved in the execution process of the control strategy of the simulated robot by the reinforcement learning strategy model, adapting the reinforcement learning strategy model to the real environment by using the respective transfer algorithms corresponding to the multiple links to obtain a robot control strategy model. Based on the simulation environment, this application can efficiently generate a large number of training samples, improving the training efficiency, accuracy, and generalization ability of the reinforcement learning strategy model; based on the respective strategy transfer algorithms corresponding to the multiple links, the adaptability between the robot control strategy model and the real environment is improved, ensuring the accuracy of the robot control strategy model in the real environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and particularly to a method and device for determining a robot control strategy model, a method and device for determining a robot control strategy, a computer-readable medium, and a program product. Background Art

[0002] Embodied intelligence refers to the physical entity robot learning how to solve specific tasks during the continuous interaction with the environment. Based on data-driven robot control strategy models such as reinforcement learning, a large number of training samples are generally required for training to ensure the generalization and accuracy of its control strategy. However, the training samples in the real world are often scarce, and it is difficult to collect training samples, resulting in limited application scope of the trained robot control strategy model and poor intelligent performance in real scenarios.

[0003] This section aims to provide background or context for the embodiments of the present application stated in the claims. The description herein is not considered prior art merely because it is included in this section. Summary of the Invention

[0004] Multiple aspects of the present application provide a method and device for determining a robot control strategy model, a method and device for determining a robot control strategy, a computer-readable storage medium, and a program product, to solve the problem of poor intelligent performance caused by insufficient adaptability of the trained robot control strategy model to real application scenarios due to scarce training samples.

[0005] In a first aspect of the present application, there is provided a method for determining a robot control strategy model, including: training a reinforcement learning strategy model for determining the control strategy of a simulated robot in a simulation environment; for multiple links involved in the execution process of the control strategy of the simulated robot by the reinforcement learning strategy model, adapting the reinforcement learning strategy model to the real environment by using the respective transfer algorithms corresponding to the multiple links, to obtain a robot control strategy model.

[0006] In a second aspect of the present application, there is provided a method for determining a robot control strategy, including: determining the control strategy of the robot according to the perception data of the robot for the real environment and the state data of the robot through the robot control strategy model, where the robot control strategy model is obtained by the method described in the first aspect above; controlling the robot to run according to the control strategy.

[0007] In a third aspect of the present application, there is provided a device for determining a robot control strategy model, including: a reinforcement learning module configured to train a reinforcement learning model for determining a control strategy of a simulation robot in a simulation environment; a policy transfer module configured to, for multiple links involved in the execution process of the control strategy of the simulation robot by the reinforcement learning model, adapt the reinforcement learning policy model to a real environment by using transfer algorithms corresponding to the multiple links respectively, so as to obtain a robot control strategy model.

[0008] In a fourth aspect of the present application, there is provided a device for determining a robot control strategy, including: a policy determination module configured to determine a control strategy of a robot according to real environment perception data of the robot for a real environment and state data of the robot through a robot control strategy model, where the robot control strategy model is obtained by the device described in the above third aspect; a control strategy module configured to control the robot to operate according to the control strategy.

[0009] In a fifth aspect of the present application, there is provided an electronic device, the device including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the methods described in the first aspect and the second aspect as above.

[0010] In a sixth aspect of the present application, there is provided a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions can be executed by a processor to implement the methods described in the first aspect and the second aspect as above.

[0011] In a seventh aspect of the present application, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the methods described in the first aspect and the second aspect as above.

[0012] In the solution provided by the embodiments of the present application, first, a reinforcement learning policy model for controlling the task execution of a simulation robot is trained in a simulation environment. Based on the simulation environment, a large number of training samples can be efficiently generated, improving the training efficiency, accuracy, and generalization ability of the reinforcement learning policy model; then, for multiple links involved in the execution process of the control strategy of the simulation robot by the reinforcement learning policy model, the reinforcement learning policy model is adapted to the real environment by using transfer algorithms corresponding to the multiple links respectively, so as to obtain a robot control strategy model. Based on the transfer algorithms corresponding to the multiple links respectively, the adaptability between the robot control strategy model and the real environment is improved, ensuring the accuracy of the robot control strategy model in the real environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0014] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present application will become more apparent:

[0015] Figure 1 Schematic flowchart of a method for determining a robot control strategy model provided by an embodiment of the present application;

[0016] Figure 2 Schematic diagram of the knowledge distillation process from the teacher strategy model to the student strategy model of the present application;

[0017] Figure 3 Schematic diagram of the migration algorithm corresponding to the motion control link of the present application;

[0018] Figure 4 Schematic flowchart of a robot control strategy provided by an embodiment of the present application;

[0019] Figure 5 Schematic structural diagram of a device for determining a robot control strategy model provided by an embodiment of the present application;

[0020] Figure 6 Schematic structural diagram of a device for determining a robot control strategy provided by an embodiment of the present application;

[0021] Figure 7 Schematic structural diagram of a device suitable for implementing the solution in the embodiments of the present application. The same or similar reference numerals in the drawings represent the same or similar components. Detailed implementation manners

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0023] In a typical configuration of the present application, the devices of the terminal and the service network both include one or more processors (CPUs), input / output interfaces, network interfaces, and memories.

[0024] The memory may include non - permanent memory in the form of computer - readable media, such as random access memory (RAM) and / or non - volatile memory, such as read - only memory (ROM) or flash RAM. The memory is an example of computer - readable media.

[0025] Computer - readable media includes permanent and non - permanent, removable and non - removable media, and information storage can be implemented by any method or technology. The information can be computer program instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase - change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read - only memory (ROM), electrically erasable programmable read - only memory (EEPROM), flash memory or other memory technologies, compact disc read - only memory (CD - ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non - transitory medium that can be used to store information that can be accessed by a computing device.

[0026] An embodiment of this application provides a method for determining a robot control strategy model. The method first trains a reinforcement learning strategy model for controlling the task execution of a simulated robot in a simulation environment. Based on the simulation environment, a large number of training samples can be efficiently generated, improving the training efficiency, accuracy, and generalization of the reinforcement learning strategy model. Then, for multiple links involved in the control strategy execution process of the reinforcement learning strategy model for the simulated robot, a migration algorithm corresponding to each link is used to adapt the reinforcement learning strategy model to the real environment, obtaining a robot control strategy model. Based on the migration algorithms corresponding to multiple links, the fitness between the robot control strategy model and the real environment is improved, ensuring the accuracy of the robot control strategy model in the real environment.

[0027] In an actual scenario, the execution entity of this method can be a user device, or a device formed by integrating a user device and a network device through a network, or it can also be an application program running on the above - mentioned device. The user device includes, but is not limited to, various terminal devices such as computers, mobile phones, tablets, smart watches, bracelets, etc. The network device includes, but is not limited to, implementations such as network hosts, single network servers, multiple network server sets, or computer clusters based on cloud computing, and can be used to implement some processing functions when setting an alarm. Here, the cloud consists of a large number of hosts or network servers based on cloud computing. Among them, cloud computing is a type of distributed computing, consisting of a virtual computer formed by a group of loosely coupled computer sets.

[0028] Figure 1The processing flow of a method provided by an embodiment of the present application is shown. The method at least includes the following processing steps:

[0029] Step S101: Train a reinforcement learning policy model in a simulation environment for determining the control policy of a simulation robot.

[0030] A simulation environment is a virtual environment created through computer technology and software tools, used to simulate physical processes, system behaviors, or environmental conditions in the real world, so as to train and test models therein. The simulation environment allows relevant personnel to iterate and optimize the model without actually building or operating a real system.

[0031] Compared with the model training process in a real environment, the training process in a simulation environment has the following multiple advantages:

[0032] Data generation: The simulation environment can efficiently generate a large amount of training data available for model training. The training data covers various situations and scenarios that the model may encounter, which can not only enrich the training set of the model but also improve the generalization ability of the model.

[0033] Condition control: In the simulation environment, relevant personnel can precisely control various parameters and conditions, such as light, temperature, pressure, etc., to simulate different working environments and test scenarios. This control ability enables relevant personnel to systematically evaluate the performance of the model under different conditions.

[0034] Safety: Compared with model training in a real environment, the simulation environment has higher safety. In the simulation environment, even if the model exhibits errors or abnormal behaviors, it will not cause damage to the actual system or personnel.

[0035] Repeatability: The simulation environment provides repeatable experimental conditions, enabling relevant personnel to run the same experiment multiple times and compare the results. This repeatability helps researchers more accurately evaluate the performance and stability of the model.

[0036] Cost - effectiveness: Compared with model training in a real environment, the training process in the simulation environment does not require purchasing and maintaining expensive hardware devices, nor does it need to bear the losses caused by experimental failures, enabling relevant personnel to conduct more experiments and iterations at a lower cost.

[0037] In this embodiment, a simulation environment corresponding to the task scenario can be specifically created according to the task scenario of the tasks that the robot needs to execute in the real environment. By modeling the real robot, a simulated robot can be obtained. In the modeling process, the STL (Stereolithography) file containing the three-dimensional model of the robot and the URDF (Unified Robot Description Format) or XML (Extensible Markup Language) file used to define the dynamic and kinematic relationships of the robot are mainly imported into the engine, and the XML model is modified into an interactive simulation environment with interactive attributes according to the rules of the simulation engine. The URDF file usually contains the link, joint, sensor, kinematic chain, and collision and visual attributes of the embodied agent.

[0038] Reinforcement Learning (RL) is a machine learning method, the purpose of which is to enable the simulated agent to learn how to obtain the maximum cumulative reward in a specific task through its interaction with the simulation environment. The simulated agent in this embodiment is specifically manifested as a simulated robot and a reinforcement learning policy model for controlling the simulated robot. The simulated robot directly interacts with the simulation environment, collects the simulation environment perception data sequence and the robot state data, and the reinforcement learning policy model determines the control strategy of the robot according to the simulation environment perception data sequence and the current state data of the robot to control the interaction between the robot and the simulation environment, and so on. The simulated agent gradually learns how to achieve a specific goal by selecting the best action sequence by trying different actions and observing the feedback (i.e., reward or punishment) of the simulation environment to its actions, and obtains the trained reinforcement learning policy model.

[0039] In order to further improve the accuracy of the trained reinforcement learning policy model, the deep reinforcement learning method can be adopted in this embodiment. Deep Reinforcement Learning (DRL) combines the technologies of deep learning and reinforcement learning, introduces high-dimensional perception data, and uses a deep neural network as the initial reinforcement learning policy model to learn the mapping relationship from the simulation environment perception data sequence and the state data of the simulated robot to the control strategy, so as to solve the decision-making problem in a complex environment.

[0040] The reward in the reinforcement learning process is a feedback signal used in deep reinforcement learning to guide the agent to learn and optimize its behavior, indicating the goodness or badness of a certain action or policy. The reward function defines a function for each state-action pair. It reflects the goal of performing the task. For example, in this application, a correct moving and obstacle-crossing direction and a safe movement will be given a positive reward, while a collision with an obstacle will be given a negative reward.

[0041] In this embodiment, the rewards in the reinforcement learning process mainly include but are not limited to the following categories:

[0042] Instruction-following reward R inst_follow : Reward the robot for following the predetermined speed and direction instructions to ensure that its behavior conforms to the control goal;

[0043] Motion efficiency reward R motion_effciency : Reward the energy efficiency and motion smoothness of the robot when performing tasks, including reducing unnecessary joint movements and torque consumption;

[0044] Stability reward R stability : Reward the robot for maintaining its own balance and gait stability and avoiding falling or unstable movements;

[0045] Collision avoidance reward -R collision_avoidance : Penalize the robot for unnecessary contact with the environment, such as collisions of the arms, feet, legs or knees with obstacles, to protect the robot from damage;

[0046] Gait smoothness reward R gait_smoothnes : Reward the robot for generating a smooth and coherent gait, and improve the naturalness and comfort of the movement by reducing mutations in the gait;

[0047] Environmental adaptability reward R environment_suitablity : Reward the robot for adjusting its behavior according to environmental changes, such as adjusting the gait and speed on different terrains;

[0048] Behavior consistency reward R consistency : Reward the robot for maintaining consistent behavior in different environments and conditions to ensure the generalization ability of the policy.

[0049] Combining all the above factors, a comprehensive reward function can be constructed, for example:

[0050] R = R inst_follow ×W inst_follow +R motion_effciency ×W motion_effciency

[0051] +R stability ×W stability -R collision_avoidance ×W collision_avoidance

[0052] +R gait_smoothness ×W gait_smppthness +R environment_suitablity

[0053] ×W environment_suitablity +R consistency ×W consistency

[0054] Among them, W inst_follow , W motion_effciency , W stability , W collision_avoidance , W gait_smoothness , W environment_suitablity , W consistency is the weight coefficient corresponding to each category, which is used to adjust the relative importance of each part.

[0055] Through these rewards, the robot is guided to achieve efficient, stable and adaptable motion control in complex environments, while reducing energy consumption and avoiding unnecessary damage. This design enables the robot to exhibit better performance and robustness in diverse tasks and environments.

[0056] In some alternative implementation manners of this embodiment, the above-mentioned execution subject may execute the above-mentioned step 101 in the following manner:

[0057] First step, disassemble the simulation task corresponding to the simulation robot to obtain multiple subtasks.

[0058] According to the set goals, specific actions and constraint conditions of the simulation task, disassemble the task with higher complexity into multiple basic subtasks with lower complexity. For example, disassemble the simulation task into basic subtasks such as the movement of the simulation robot body, object operation, object grasping, etc.

[0059] Second step, according to the multiple subtasks, parallelly train the initial reinforcement learning policy model used to control the operation of the simulation robot in the simulation environment to obtain the reinforcement learning policy model.

[0060] GPU (graphics processing unit) has the ability to process tens of thousands of parallel instructions. In this implementation manner, based on the simulation engine with parallel GPU computing, by simultaneously generating and managing multiple (e.g., thousands of) robots in the same physically accurate simulation environment, efficient parallel collection of policy learning data is achieved. According to different basic subtasks, their corresponding parallel training environments will be generated.

[0061] For the parallel training environment of the simulation robot's movement and obstacle crossing tasks, multiple robots can be placed together in the same terrain grid environment. This grid does not change as a whole during each training reset. Instead, by arranging different terrain types and difficulty levels in parallel, a comprehensive terrain grid is formed. Each simulation robot is assigned a specific terrain type and the corresponding difficulty level. The simulation robot is placed at the center of a simple grid and moves along a terrain trajectory with gradually increasing challenges. When the simulation robot successfully crosses the current terrain, its difficulty level will be automatically increased, and it will start from a higher-difficulty terrain during the next training reset. On the contrary, if the distance the robot moves in a round of training does not reach half of its target distance, its difficulty level will be correspondingly decreased. To increase training diversity and prevent forgetting the obstacle crossing skills for low-level terrains, the simulation robots that reach the highest difficulty level will randomly select the training difficulty during reset. The advantage of this adaptive training method is that it can dynamically adjust the training difficulty according to the actual performance of the robot without external intervention. In addition, it can independently adapt to the difficulty of each terrain type and provide intuitive and quantitative feedback on the training progress.

[0062] Finally, when the simulation robots can be evenly distributed on all terrains and successfully complete the obstacle crossing tasks, it indicates that they have mastered the tasks.

[0063] For tasks such as object manipulation and object grasping, multiple sub-environments will be generated simultaneously. In each sub-environment, the robotic arm interacts with the precisely modeled target object, and the reinforcement learning policy model for controlling the robotic arm is trained. The multiple interacting objects can be interacting objects with the same parameters, or interacting objects of the same parameter type but with different parameter distributions.

[0064] During the training process of the reinforcement learning policy model, the data collected by each simulation robot is crucial for the update of the policy gradient. Assume that the amount of data collected before each batch iteration is determined by the formula N episode = N robot × N steps where N robot represents the number of simulation robots in parallel training, and N steps represents the maximum number of iteration steps that each robot can execute in a round of training.

[0065] Therefore, finding the optimal balance in the number of simulation robots is the key to achieving efficient training. When the number of robots is excessive, the amount of data contributed by each robot to the policy update decreases, which may lead to unstable gradient estimation, increased noise, and thus affect the stability and efficiency of the learning process. Conversely, if the number of simulation robots is small, each simulation robot needs to perform more iteration steps in each policy update. Although this maintains the consistency of the samples in terms of quantity, due to the proximity of the samples in time, it reduces the diversity of the sample data, violates the assumption of independent and identically distributed, and has a negative impact on the training effect. In this embodiment, through experiments and adjustments, the optimal number of robots is found, and according to the optimal data parallel training process, to ensure that while maintaining data diversity, accurate gradient estimation can also be obtained, thereby improving the learning efficiency and model performance.

[0066] In this implementation manner, combining task decomposition and parallel training helps to improve the training efficiency and accuracy of the reinforcement learning model.

[0067] Step S102, for multiple links involved in the process of the reinforcement learning policy model controlling the simulation robots, use the transfer algorithms corresponding to each link to adapt the reinforcement learning policy model to the real environment, and obtain the robot control policy model.

[0068] In this embodiment, the corresponding relationship between the links and the transfer algorithms is determined in advance. Furthermore, for each link involved in the process of the reinforcement learning policy model executing the control policy for the simulation robots, use the transfer algorithm corresponding to this link to adapt this link corresponding to the reinforcement learning policy model to the real environment, so as to obtain the robot control policy model.

[0069] Among them, the Sim2Real (Simulation to Reality) transfer algorithm aims to solve the problem of migrating the model trained in the simulation environment to the real world. Such algorithms achieve cross-domain migration of the model by reducing the differences between the simulation environment and the real environment, including but not limited to domain adaptation algorithms, domain randomization algorithms, inverse dynamics model algorithms, cross-model transfer algorithms. Based on the rich experience of technicians, the corresponding relationship between the links and the transfer algorithms can be determined.

[0070] In this embodiment, for the timing relationship between different nodes in the entire control policy execution process, the whole is divided into multiple links with sequential timing relationships, for example, the environmental perception link, the policy control link, and the task execution link.

[0071] In some alternative implementation manners of this embodiment, the above-mentioned execution entity may execute the above step 102 in the following manner: During the process of iteratively training the student policy model with the reinforcement learning policy model as the teacher policy model, the student policy model during the training process is gradually adapted to the real environment by using the transfer algorithms corresponding to multiple links, so as to obtain the robot control policy model.

[0072] In this implementation manner, the knowledge distillation process from the teacher policy model to the student policy model and the transfer process from the student policy model to the real environment are carried out simultaneously.

[0073] Continue to refer to Figure 2 , which shows a schematic diagram of the knowledge distillation process from the teacher policy model to the student policy model.

[0074] First, input the current state data of the simulation robot and the sequence of simulation environment perception data collected by the simulation robot into the trained reinforcement learning policy model to obtain a control policy; then, the simulation robot executes the control policy to interact with the simulation environment; finally, based on the changed simulation environment and the reward function, determine the reward data.

[0075] The student policy model also inputs the current state data of the simulation robot and the sequence of simulation environment perception data collected by the simulation robot, outputs a prediction instruction, and calculates the loss between the prediction instruction and the control policy output by the reinforcement learning policy model, so as to update the parameters of the student policy model through the loss.

[0076] By iteratively executing the above process, in response to reaching a preset end condition, a trained student policy model is obtained. Among them, the preset end condition is, for example, that the training time exceeds a preset time threshold, the number of training times exceeds a preset number threshold, and the training loss region converges.

[0077] In each training process of the student policy model, for each link in the training process, the transfer algorithm corresponding to this link is used for data processing, so that the student policy model is gradually adapted to the real environment during the training process. In this way, along with the process of distilling the learning policy model from the teacher policy model, the obtained student policy model realizes the transfer from the simulation environment to the real environment, and the finally obtained student policy model is used as the robot control policy model.

[0078] In this implementation manner, through the knowledge distillation process, a student policy model that is easy to deploy, has a high inference speed and generalization ability can be obtained, and during the knowledge distillation process, the model transfer process for the real environment is carried out simultaneously, which helps to improve the overall processing speed while improving the adaptability between the robot control policy model and the real environment.

[0079] In some alternative implementation manners of this embodiment, the above-mentioned execution entity may execute the above-mentioned migration operation in the following manner:

[0080] First, for the motion control link of the student policy model for the simulation robot, the state credibility recursive encoder is used to determine the credibility of the state data of the simulation robot and the simulation environment perception data sequence, and credibility data is obtained. Then, through the student policy model, the control strategy of the simulation robot is determined according to the credibility data, so as to gradually adapt the motion control algorithm of the student policy model to the real environment.

[0081] Continue to refer to Figure 3 , which shows a schematic diagram of the migration algorithm corresponding to the policy control link. Among them, the hidden state represents the hidden state of the student policy model.

[0082] The motion control link refers to the link where the student policy model controls the simulation robot to execute tasks based on the control strategy. The key component of the migration method corresponding to the motion control link is a state credibility recursive encoder that fuses the self-state data of the simulation robot and the simulation environment perception data sequence. The state credibility recursive encoder needs to be trained in the simulation environment to determine the feasibility of the state data and the simulation environment perception data sequence. After end-to-end training, the state credibility encoder can integrate the state data of the simulation robot's own body and the simulation environment perception data sequence, and without relying on heuristic methods, it can learn to use the model prediction value to prospectively plan the support points and accelerating motion of the simulation robot when the simulation environment perception data sequence is reliable, and seamlessly fallback to the motion planning that depends on the body's own state data when needed. Among them, the simulation environment perception data sequence of the robot for the simulation environment includes but is not limited to data of types such as images, point clouds, and sounds collected by cameras, radars, and sound collectors on the robot; the state data of the robot is the perception data of the robot's own state, such as control instructions acting on the body; the body speed and direction (linear speed and angular speed); joint position, speed, and acceleration; phase information for gait generation.

[0083] Therefore, combined with the learning model of the state credibility recursive encoder, the robot can rely on the environmental perception data to bring higher speed and efficiency, and also rely on its own body state data to run stably.

[0084] In this implementation manner, for the motion control link, based on the credibility judgment of the state data of the simulation robot and the simulation environment perception data sequence by the state credibility recursive encoder, the student policy model can judge the credibility of the input data in the real environment, and realize the adaptation to the real environment.

[0085] In some alternative implementation manners of this embodiment, the above-mentioned execution entity may execute the above-mentioned migration operation in the following manner: For multiple simulation environment perception data sequences obtained by the simulation robot in the environment perception link, use the preprocessing methods corresponding to the multiple environment perception data to process the multiple simulation environment perception data sequences, so as to reduce the difference between the corresponding simulation environment perception data sequences and the real environment perception data, and gradually adapt the simulation environment perception data sequences required by the student policy model to the real environment.

[0086] In the environment perception link, it aims to reduce the appearance difference between the simulation environment perception data sequence of the student policy model in the simulation environment and the real environment perception data in the real environment.

[0087] The multiple simulation environment perception data sequences include but are not limited to perception data of types such as images, point clouds, and sounds. Taking the simulation environment perception data sequence in the form of an image as an example, for an RGB (Red, Green, Blue) image, preprocessing such as motion blur and Gaussian noise is performed; for a depth image, preprocessing such as depth cropping, pixel-level Gaussian noise, and random artifacts is performed.

[0088] Taking the simulation environment perception data sequence obtained by lidar as an example, preprocessing such as simulating the characteristics of real sensors and data augmentation is performed.

[0089] In this embodiment, in order to further reduce the appearance difference between the simulation environment perception data sequence of the student policy model in the simulation environment and the real environment perception data in the real environment, the real environment perception data corresponding to the real environment can also be preprocessed. Taking the depth image in the real environment as an example, preprocessing such as depth cropping, hole filling, spatial smoothing, and temporal smoothing can be performed on it.

[0090] In this implementation manner, for the simulation environment perception data sequence obtained in the environment perception link, it is processed through the corresponding preprocessing method to reduce the appearance difference between the simulation environment perception data sequence of the student policy model in the simulation environment and the real environment perception data in the real environment, so that the finally obtained robot control policy model is adapted to the real environment in the environment perception dimension.

[0091] In some alternative implementation manners of this embodiment, the above-mentioned execution entity may execute the above-mentioned migration operation in the following manner: For the operation link of the simulation robot based on the control instructions of the student policy model, perform at least one of the following migration operations to gradually adapt the control instructions of the student policy model during the training process to the real environment:

[0092] 1. Randomly set the initial physical state of the simulation robot in the simulation environment.

[0093] In each training, randomize the mass, initial joint positions, velocities of the body, arms, legs, etc. of the simulated robot, as well as the initial self-pose and velocity of the simulated robot.

[0094] 2. Randomly set the interaction parameters between the simulated robot and the simulated objects that interact with the simulated robot.

[0095] The simulated objects that interact with the simulated robot are the simulated objects that the simulated robot has come into contact with during the process of running according to the control instructions, such as the ground, stairs, and objects to be grasped by the robotic arm. Interaction parameters are, for example, parameters such as the friction coefficient between the simulated robot and the interacting simulated objects.

[0096] 3. Apply external force action parameters that affect the running state of the simulated robot.

[0097] As an example, external forces and torques can be applied to the robot body, or occasionally the friction coefficient of the feet can be set to a low value to introduce a slipping phenomenon.

[0098] 4. Alternately train the student policy model in the simulation environment and the real environment.

[0099] For tasks such as operating objects and object grasping, in order to reduce the problem of inconsistent parameter distributions between the real environment and the simulation environment, the student policy model is alternately trained in the simulation environment and the real environment, so as to adjust the distribution of the simulation parameters. In this way, the parameter distribution of the simulation environment can be changed, and by matching the simulation environment and the real environment, the transfer of the student policy model to the real environment can be realized.

[0100] In this implementation manner, in the task execution link of the simulated robot based on the control instructions of the student policy model, by performing the corresponding transfer operations, the student policy model is adapted to the real environment, improving the adaptability of the student policy model to the real environment in the running link.

[0101] In some optional implementation manners of this embodiment, the above execution entity can execute the process of iteratively training the student policy model with the reinforcement learning policy model as the teacher policy model in the following manner:

[0102] The first step is to establish a simulation model of the real robot based on the kinematic characteristics of the real robot to be controlled by the robot control policy model in the real environment.

[0103] During the training process of the reinforcement learning policy model, there may be differences between the simulated robot and the real robot on which it is based, resulting in the reinforcement learning policy model not being fully adapted to the real robot. In this implementation, a simulation model of the real robot is established based on the kinematic characteristics of the real robot, aiming to train the student policy model completely based on the kinematic characteristics of the real robot, so that the student policy model is adapted to the real robot.

[0104] For example, there are slight differences in the dynamic characteristics (such as mass distribution, friction, etc.) between robot forms such as robotic arms and quadruped robots of the same type but different models. Physical-level modeling needs to be carried out for the real robot of the target model. Among them, the target model is, for example, the model corresponding to the robot to be controlled in the real environment.

[0105] In the second step, using the reinforcement learning policy model as the teacher policy model and the simulation model as the simulated robot controlled by the student policy model, iteratively train the student policy model.

[0106] In this implementation, the simulation model executes the instructions output by the student policy model to interact with the simulation environment. By modeling the real robot in the real environment, the student policy model is gradually adapted to the real robot during the training process.

[0107] In some optional implementation manners of this embodiment, the above execution subject may execute the second step in the following manner: First, through the reinforcement learning policy model, determine the control strategy of the simulated robot according to the state data of the simulated robot, the sequence of simulation environment perception data of the simulated robot for the simulation environment, and the environment interaction data between the simulated robot and the simulation environment; then, use the state data and the sequence of simulation environment perception data as inputs and the control strategy as the expected output to train the student policy model.

[0108] The environment interaction data between the simulated robot and the simulation environment, as a kind of auxiliary training data for training the teacher policy, is data that is easy to obtain in the simulation environment but may not be obtainable or difficult to obtain in the real environment. For example, it is precise terrain height information and friction coefficient; detailed information on the contact state between the robot's feet and the ground, such as contact force and the normal of the contact point; external forces and torques acting on the robot body; contact states of the thighs and calves of the legged robot, etc.

[0109] The teacher policy model can use three types of information: the ontology state data of the simulation robot, the simulation environment perception data sequence, and the auxiliary training data, with the aim of learning how to perform tasks under complex and changeable conditions in the simulation environment. The student policy model can only use the ontology state data and the simulation environment perception data sequence, aiming to imitate the behavior of the teacher policy model in the case of only the ontology state data and the simulation environment perception data sequence, so as to achieve the effective deployment of the student policy model in the real environment.

[0110] In some alternative implementation manners of this embodiment, the above-mentioned execution entity can also perform the following operations: adjust the parameters of the robot control policy model according to the differences between the real tasks in the real environment and the simulation tasks in the simulation environment.

[0111] As an example, if the real task in the real environment is to carry an item of 10 kilograms, while the simulation task in the simulation environment is to carry an item of 5 kilograms, then it is necessary to adjust the parameter of the robot control policy model for the carrying force.

[0112] In this implementation manner, adjusting the parameters of the robot control policy model based on the differences between the real task and the simulation task can make the robot control policy model further adapt to the real tasks in the real environment.

[0113] In some alternative implementation manners of this embodiment, the above-mentioned execution entity can also perform the following operations: trim the robot control policy model according to the computing power of the deployment device corresponding to the robot control policy model in the real environment.

[0114] For example, when the computing power of the deployment device used to deploy the robot control policy model can meet the computing requirements of the robot control policy model, directly deploy the complete robot policy control model on the deployment device; for another example, when the computing power of the deployment device used to deploy the robot control policy model cannot meet the computing requirements of the robot control policy model, it is necessary to trim the robot control policy model to reduce the parameter scale of the robot control policy model.

[0115] In this implementation manner, flexibly trimming the robot control policy model based on the computing power of the deployment device makes the robot control policy model further adapt to the real environment.

[0116] Continue to refer to Figure 4 , which shows the processing flow of a method provided by an embodiment of the present application. The method at least includes the following processing steps:

[0117] Step 401, through the robot control policy model, determine the control policy of the robot according to the real environment perception data of the robot for the real environment and the state data of the robot.

[0118] Among them, the robot control strategy model is obtained through any implementation manner in the method for determining the robot control strategy model.

[0119] As an example, after receiving the target task, through the robot control strategy model, according to the real environment perception data of the robot for the real environment and the state data of the robot, the control instructions required for the robot to complete the target task are determined.

[0120] The real environment perception data of the robot for the real environment includes but is not limited to data of types such as images, point clouds, and sounds collected by cameras, radars, and sound collectors on the robot; the state data of the robot is the perception data of the robot's own state, such as control instructions acting on the body; body speed and direction (linear speed and angular speed); joint position, speed, and acceleration; phase information for gait generation.

[0121] Step 402, control the robot to run according to the control strategy.

[0122] After the robot determines the control instructions, it runs according to the control instructions to interact with the real environment. By repeatedly executing the above steps 401 - 402 until the target task to be executed by the robot is completed.

[0123] In addition, the embodiment of the present application also provides a device for determining a robot control strategy model, and the structure of the device is as Figure 5 shown.

[0124] A device for determining a robot control strategy model includes: a reinforcement learning module 501 configured to train a reinforcement learning strategy model for determining the control strategy of a simulation robot in a simulation environment; a strategy migration module 502 configured to, for multiple links involved in the execution process of the control strategy of the reinforcement learning strategy model for the simulation robot, adapt the reinforcement learning strategy model to the real environment by using the migration algorithms corresponding to the multiple links respectively to obtain a robot control strategy model.

[0125] In some optional implementation manners of this embodiment, the strategy migration module 502 is further configured to: in the process of iteratively training a student strategy model with the reinforcement learning strategy model as the teacher strategy model, gradually adapt the student strategy model during the training process to the real environment by using the migration algorithms corresponding to the multiple links respectively to obtain a robot control strategy model.

[0126] In some alternative implementation manners of this embodiment, the policy migration module 502 is further configured to: for the motion control link of the student policy model for the simulation robot, determine the credibility of the state data of the simulation robot and the sequence of simulation environment perception data through the state credibility recursive encoder to obtain credibility data; according to the credibility data, determine the control policy of the simulation robot through the student policy model, so as to gradually adapt the motion control algorithm of the student policy model to the real environment.

[0127] In some alternative implementation manners of this embodiment, the policy migration module 502 is further configured to: for multiple sequences of simulation environment perception data obtained by the simulation robot in the environment perception link, process the multiple sequences of simulation environment perception data by using the respective preprocessing methods corresponding to the multiple types of environment perception data, so as to reduce the difference between the corresponding sequences of simulation environment perception data and the real environment perception data, and gradually adapt the sequence of simulation environment perception data required by the student policy model to the real environment.

[0128] In some alternative implementation manners of this embodiment, the policy migration module 502 is further configured to: for the operation link of the simulation robot based on the control instruction of the student policy model, perform at least one of the following migration operations to gradually adapt the control instruction of the student policy model during the training process to the real environment: randomly set the initial physical state of the simulation robot in the simulation environment; randomly set the interaction parameters between the simulation robot and the simulation objects that interact with the simulation robot; apply an external force action parameter that affects the running state of the simulation robot; alternately train the student policy model in the simulation environment and the real environment.

[0129] In some alternative implementation manners of this embodiment, the policy migration module 502 is further configured to: establish a simulation model of the real robot based on the kinematic characteristics of the real robot to be controlled by the robot control policy model in the real environment; use the reinforcement learning policy model as the teacher policy model and the simulation model as the simulation robot controlled by the student policy model, and iteratively train the student policy model.

[0130] In some alternative implementation manners of this embodiment, the policy migration module 502 is further configured to: determine the control policy of the simulation robot through the reinforcement learning policy model according to the state data of the simulation robot, the sequence of simulation environment perception data of the simulation robot for the simulation environment, and the environment interaction data between the simulation robot and the simulation environment; use the state data and the sequence of simulation environment perception data as the input and the control policy as the expected output to train the student policy model.

[0131] In some alternative implementation manners of this embodiment, the above device further includes: an adjustment module (not shown in the figure), configured to adjust parameters of the robot control policy model according to the difference between the real task in the real environment and the simulated task in the simulated environment.

[0132] In some alternative implementation manners of this embodiment, the above device further includes: a pruning module (not shown in the figure), configured to prune the robot control policy model according to the computing power of the deployment device corresponding to the robot control policy model in the real environment.

[0133] In some alternative implementation manners of this embodiment, the reinforcement learning module 501 is further configured to: disassemble the simulated task corresponding to the simulated robot to obtain a plurality of subtasks; and based on the plurality of subtasks, parallelly train an initial reinforcement learning policy model for controlling the operation of the simulated robot in the simulated environment to obtain a reinforcement learning policy model.

[0134] In this embodiment, first, a reinforcement learning policy model for determining the control policy of the simulated robot is trained in the simulated environment. Based on the simulated environment, a large number of training samples can be efficiently generated, improving the training efficiency, accuracy, and generalization ability of the reinforcement learning policy model. Then, for multiple links involved in the execution process of the control policy of the simulated robot by the reinforcement learning policy model, the reinforcement learning policy model is adapted to the real environment by using the migration algorithms corresponding to the multiple links respectively to obtain a robot control policy model. Based on the migration algorithms corresponding to the multiple links respectively, the adaptability between the robot control policy model and the real environment is improved, ensuring the accuracy of the robot control policy model in the real environment.

[0135] In addition, an embodiment of the present application further provides a robot control device, and the structure of the device is as Figure 6 shown.

[0136] A device for determining a robot control policy includes: a policy determination module 601, configured to determine the control policy of the robot through the robot control policy model according to the real environment perception data of the robot for the real environment and the state data of the robot, where the robot control policy model is obtained through the implementation manners in the embodiment of the device for determining the robot control policy model; and a control policy module 602, configured to control the robot to operate according to the control policy.

[0137] Based on the same inventive concept, an electronic device is further provided in an embodiment of the present application. The method corresponding to the electronic device may be the method for determining a robot control strategy model and the robot control method in the foregoing embodiments, and the principle of solving problems is similar to that of the method. The electronic device provided in the embodiment of the present application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the methods and / or technical solutions of multiple foregoing embodiments of the present application.

[0138] The electronic device may be a user device, or a device formed by integrating a user device and a network device through a network, or may also be an application program running on the above devices. The user device includes, but is not limited to, various terminal devices such as a computer, a mobile phone, a tablet computer, a smart watch, and a bracelet. The network device includes, but is not limited to, a network host, a single network server, a set of multiple network servers, or a computer set based on cloud computing, etc., and can be used to implement some processing functions when setting an alarm. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing. Among them, cloud computing is a type of distributed computing, which consists of a virtual computer formed by a group of loosely coupled computer sets.

[0139] Figure 7 The structure of a device suitable for implementing the methods and / or technical solutions in the embodiments of the present application is shown. The device 700 includes a central processing unit (CPU, Central Processing Unit) 701, which can execute various appropriate actions and processes according to the program stored in the read-only memory (ROM, Read Only Memory) 702 or the program loaded from the storage part 708 into the random access memory (RAM, Random Access Memory) 703. In the RAM 703, various programs and data required for system operation are also stored. The CPU 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O, Input / Output) interface 705 is also connected to the bus 704.

[0140] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, a touch screen, a microphone, an infrared sensor, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), an LED display, an OLED display, etc., and a speaker; a storage section 708 including one or more computer-readable media such as a hard disk, an optical disk, a magnetic disk, a semiconductor memory, etc.; and a communication section 709 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet.

[0141] Specifically, the method and / or embodiments in the embodiments of the present application can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. When the computer program is executed by a central processing unit (CPU) 701, the above functions defined in the method of the present application are executed.

[0142] Another embodiment of the present application also provides a computer-readable storage medium, on which computer program instructions are stored, and the computer program instructions can be executed by a processor to implement the method and / or technical solution of any one or more of the foregoing embodiments of the present application.

[0143] Specifically, this embodiment can adopt any combination of one or more computer-readable media. The computer-readable media can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device.

[0144] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including - but not limited to - electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0145] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including - but not limited to - wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0146] The computer program code for performing the operations of this application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0147] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions denoted in the blocks may occur in an order different from that denoted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0148] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0149] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or page components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can be in electrical, mechanical, or other forms.

[0150] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0151] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.

[0152] The above integrated units implemented in the form of software functional units can be stored in a computer-readable storage medium. The above software functional units stored in a storage medium include several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

[0154] In addition, obviously the word "including" does not exclude other units or steps, and the singular does not exclude the plural. A plurality of units or devices stated in the apparatus claims can also be implemented by one unit or device through software or hardware. The terms such as first and second are used to denote names, rather than indicating any particular order.

Claims

1. A method for determining a robot control model, wherein, The method includes: Training a reinforcement learning policy model in a simulation environment for determining a control strategy of a simulation robot; During the process of iteratively training a student policy model with the reinforcement learning policy model as the teacher policy model, for multiple links involved in the execution process of the control strategy of the simulation robot by the reinforcement learning policy model, using the respective corresponding transfer algorithms for the multiple links to gradually adapt the student model during the training process to the real environment, obtaining a robot control strategy model; Wherein, the using the respective corresponding transfer algorithms for the multiple links to gradually adapt the student policy model during the training process to the real environment includes: For the motion control link of the simulation robot by the student policy model, determining the credibility of the state data of the simulation robot and the sequence of simulation environment perception data through a state credibility recursive encoder, obtaining credibility data; Based on the student policy model, determining the control strategy of the simulation robot according to the credibility data, so as to gradually adapt the motion control algorithm of the student policy model to the real environment.

2. The method according to claim 1, wherein, The using the respective corresponding transfer algorithms for the multiple links to gradually adapt the student policy model during the training process to the real environment includes: For multiple sequences of simulation environment perception data obtained by the simulation robot in the environment perception link, processing the multiple sequences of simulation environment perception data using the respective corresponding preprocessing methods for the multiple sequences of simulation environment perception data, so as to reduce the difference between the corresponding sequence of simulation environment perception data and the real environment perception data, and gradually adapt the sequence of simulation environment perception data required by the student policy model to the real environment.

3. The method according to claim 1, wherein The using the respective corresponding transfer algorithms for the multiple links to gradually adapt the student policy model during the training process to the real environment includes: For the operation link of the simulation robot based on the control strategy of the student policy model, performing at least one of the following transfer operations to gradually adapt the control strategy of the student policy model during the training process to the real environment: Randomly setting the initial physical state of the simulation robot in the simulation environment; Randomly setting the interaction parameters between the simulation robot and the simulation objects that interact with the simulation robot; Applying an external force action parameter that affects the running state of the simulation robot to the simulation robot; Alternately training the student policy model in the simulation environment and the real environment.

4. The method according to claim 1, wherein, The iteratively training the student policy model with the reinforcement learning policy model as the teacher policy model includes: Based on the kinematic characteristics of the real robot to be controlled by the robot control strategy model in the real environment, establishing a simulation model of the real robot; Using the reinforcement learning policy model as the teacher policy model and using the simulation model as the simulation robot controlled by the student policy model, iteratively training the student policy model.

5. The method according to claim 4, wherein The using the reinforcement learning policy model as the teacher policy model and using the simulation model as the simulation robot controlled by the student policy model to iteratively train the student policy model includes: Based on the state data of the simulation robot, the sequence of simulation environment perception data of the simulation robot for the simulation environment, and the environmental interaction data between the simulation robot and the simulation environment, determine the control strategy of the simulation robot through the reinforcement learning policy model; Using the state data and the sequence of simulation environment perception data as inputs and the control strategy as the desired output, train the student policy model.

6. The method according to any one of claims 1-5, wherein, It further includes: Adjust the parameters of the robot control strategy model according to the difference between the real task in the real environment and the simulation task in the simulation environment.

7. The method according to any one of claims 1-5, wherein, It further includes: Prune the robot control strategy model according to the computing power of the deployment device corresponding to the robot control strategy model in the real environment.

8. The method according to any one of claims 1-5, wherein, The reinforcement learning policy model trained in the simulation environment for determining the control strategy of the simulation robot includes: Decompose the simulation task corresponding to the simulation robot to obtain multiple subtasks; According to the multiple subtasks, parallelly train the initial reinforcement learning policy model for controlling the operation of the simulation robot in the simulation environment to obtain the reinforcement learning policy model.

9. A method for determining a robot control strategy, wherein, The method includes: Through the robot control strategy model, determine the control strategy of the robot according to the real environment perception data of the robot for the real environment and the state data of the robot, where the robot control strategy model is obtained through any one of claims 1 - 8; Control the robot to operate according to the control strategy.

10. An apparatus for determining a robot control strategy model, wherein, The device includes: A reinforcement learning policy module configured to train a reinforcement learning policy model for determining the control strategy of the simulation robot in the simulation environment; A policy transfer module configured to, in the process of iteratively training the student policy model with the reinforcement learning policy model as the teacher policy model, for multiple links involved in the execution process of the control strategy of the simulation robot by the reinforcement learning policy model, adopt the respective corresponding transfer algorithms of the multiple links to gradually adapt the student model during training to the real environment to obtain the robot control strategy model; Wherein, the policy transfer module is further configured to: For the motion control link of the simulation robot by the student policy model, determine the credibility of the state data and the sequence of simulation environment perception data of the simulation robot through a state credibility recursive encoder to obtain credibility data; through the student policy model, determine the control strategy of the simulation robot according to the credibility data to gradually adapt the motion control algorithm of the student policy model to the real environment.

11. An apparatus for determining a robot control strategy, wherein, The device includes: A policy determination module configured to determine the control strategy of the robot through the robot control strategy model according to the real environment perception data of the robot for the real environment and the state data of the robot, where the robot control strategy model is obtained through claim 10; A control strategy module configured to control the robot to operate according to the control strategy.

12. An electronic device, the electronic device includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1 to 9.

13. A computer-readable medium having computer program instructions stored thereon, the computer program instructions being executable by a processor to implement the method according to any one of claims 1 to 9.

14. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Action generation method and device, electronic equipment and storage medium

    CN116360472A