Robot control method, device, equipment and computer storage medium
By using backward experience replay to update reward values, the problems of insufficient generalization of robot control methods and slow training process in the prior art are solved, and efficient agent training and multi-task completion capabilities are achieved.
Patent Information
- Application Number
- CN202011271477.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-13
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2040-11-13
AI Technical Summary
Existing robot control methods require training multiple models when completing specific tasks, which are not very generalized and have a slow training process.
By obtaining environmental interaction data, updating reward values using backward experience playback, improving data utilization, accelerating agent training, and completing multiple goals at the same time through one model.
It improves the efficiency and generalization ability of agent training, allowing a model to complete multiple tasks and shortens the training time.
Smart Images

Figure CN112476424B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of Internet technology, and relate to but are not limited to a robot control method, device, equipment and computer storage medium. Background Art
[0002] At present, when controlling a robot, one implementation method is a deep reinforcement learning control algorithm based on a priority experience replay mechanism. The priority is calculated using the state information of the object operated by the robot, and the deep reinforcement learning method is used to complete the end-to-end robot control model, allowing the deep reinforcement learning agent to autonomously learn in the environment and complete the specified tasks; another implementation method is a kinematic self-grasping learning method based on a simulated industrial robot, which belongs to the field of computer-aided manufacturing. The robot grasping training is based on a simulation environment and uses reinforcement learning theory. The simulated robot automatically obtains the position information of the object through the image taken by the camera, and determines the grasping position of the robot's end grasping tool. At the same time, the image processing method based on reinforcement learning determines the posture of the grasping tool according to the shape and placement status of the grasped object in the observed image, and finally successfully grasps objects of various shapes and randomly placed objects.
[0003] However, the implementation method in the related technology usually requires training a model to complete a feature task, the generalization is not strong, and the training process of the robot task is slow. Summary of the invention
[0004] The embodiments of the present application provide a robot control method, device, equipment and computer storage medium, which relate to the field of artificial intelligence technology. Based on the determined reward value, the reward value in the environmental interaction data is updated, and the backward experience playback method is used to improve the utilization rate of data and accelerate the training of the intelligent agent. Since the environmental interaction data includes the target value, a large number of targets can be trained at the same time, and all tasks in a certain target space can be completed through one model.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] The present application provides a robot control method, including:
[0007] Acquire environmental interaction data, wherein the environmental interaction data at least includes state data, action data, reward value, and target value at two adjacent moments;
[0008] Obtaining an actual target value actually achieved after executing the action corresponding to the action data;
[0009] determining a reward value after performing the action according to the state parameter at a first moment of the two adjacent moments, the action data and the actual target value;
[0010] Using the reward value after executing the action to update the reward value in the environment interaction data, to obtain updated environment interaction data;
[0011] Using the updated environmental interaction data to train an intelligent agent corresponding to a robot control network;
[0012] The trained agent is used to control the actions of the target robot.
[0013] The present application provides a robot control device, comprising:
[0014] A first acquisition module, used to acquire environmental interaction data, wherein the environmental interaction data at least includes state data, action data, reward value and target value at two adjacent moments;
[0015] A second acquisition module is used to acquire an actual target value actually achieved after executing the action corresponding to the action data;
[0016] a determination module, configured to determine a reward value after executing the action according to the state parameter at the first moment of the two adjacent moments, the action data and the actual target value;
[0017] An updating module, configured to update the reward value in the environment interaction data using the reward value after executing the action, to obtain updated environment interaction data;
[0018] A training module, used to train an intelligent agent corresponding to a robot control network using the updated environment interaction data;
[0019] The control module is used to control the actions of the target robot using the trained intelligent agent.
[0020] In some embodiments, the training module is also used to: at each moment, according to the target value in the updated environmental interaction data, control the intelligent agent to execute the action data in the updated environmental interaction data to obtain the state data of the next moment, and obtain the reward value of the next moment; obtain the reward values of all future moments after the next moment; determine the cumulative reward value corresponding to the reward values of all future moments; and control the training process of the intelligent agent with maximizing the cumulative reward value as the control goal.
[0021] In some embodiments, the training module is also used to: determine the expected cumulative reward of the cumulative reward value; calculate the initial action value function based on the expected cumulative reward; through forward experience playback, use the environmental interaction data at multiple consecutive moments to expand the initial action value function to obtain the expanded action value function, so as to accelerate the learning of the action value function and realize control of the training process of the intelligent agent.
[0022] In some embodiments, the training module is also used to: obtain the expected reward value and a preset discount factor for each future moment in multiple consecutive future moments after the current moment; and obtain the expanded action value function based on the discount factor and the expected reward value for each future moment.
[0023] In some embodiments, the training module is also used to: obtain the weight of the action value function; wherein the value of the weight is greater than 0 and less than 1; through forward experience playback, using the environmental interaction data of multiple consecutive future moments, expand the initial action value function based on the weight to obtain the expanded action value function.
[0024] In some embodiments, expanding the initial action-value function based on the weight is implemented by the following formula:
[0025]
[0026] Among them, Q target (n) (λ) represents the action value function after expansion based on weight λ, Q target (i) represents the initial action-value function.
[0027] In some embodiments, the device also includes: an action data determination module, used to determine the action data at the next moment according to the expanded action value function; a second update module, used to use the action data at the next moment to update the action data in the environment interaction data to obtain updated environment interaction data; correspondingly, the training module is also used to use the updated environment interaction data to train the intelligent agent corresponding to the robot control network.
[0028] In some embodiments, the device further comprises: an execution strategy determination module, for determining the execution strategy of the agent according to the accumulated reward value when the current reward value is used to update the reward value in the environmental interaction data; a selection module, for selecting the action data at the next moment according to the execution strategy;
[0029] The third updating module is used to update the action data at the next moment into the environment interaction data to obtain the updated environment interaction data.
[0030] In some embodiments, after the agent executes the action, the state of the environment in which the agent is currently located is transferred to the state at the next moment, wherein the state at the next moment corresponds to the state parameters at the next moment; the device also includes: a fourth update module, used to update the state parameters at the next moment to the environmental interaction data to obtain the updated environmental interaction data.
[0031] In some embodiments, there are multiple target values, and the device also includes: a simultaneous determination module, used to simultaneously determine the multiple target values at the next moment when using the updated environmental interaction data to train the intelligent agent corresponding to the robot control network; a fifth update module, used to update the determined multiple target values at the next moment into the environmental interaction data.
[0032] An embodiment of the present application provides a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; wherein a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor is used to execute the computer instructions to implement the above-mentioned robot control method.
[0033] The present application provides a robot control device, including:
[0034] The memory is used to store executable instructions; the processor is used to implement the above-mentioned robot control method when executing the executable instructions stored in the memory.
[0035] An embodiment of the present application provides a computer-readable storage medium storing executable instructions for causing a processor to execute the executable instructions to implement the above-mentioned robot control method.
[0036] The embodiments of the present application have the following beneficial effects: obtaining environmental interaction data, wherein the environmental interaction data includes at least state data, action data, reward value and target value at two adjacent moments, determining the reward value after executing the action according to the state parameters, action data and actual target value of the action at the first moment of the two adjacent moments, and updating the reward value in the environmental interaction data, that is, utilizing the backward experience playback method to improve data utilization and accelerate the training of the intelligent agent, and since the environmental interaction data includes the target value, a large number of targets can be trained at the same time, and all tasks in a certain target space can be completed by one model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1A It is a schematic diagram of the implementation process of a robot control method in the related art;
[0038] Figure 1B It is a schematic diagram of the implementation flow of another robot control method in the related art;
[0039] Figure 2 is an optional architecture diagram of a robot control system provided in an embodiment of the present application;
[0040] Figure 3 It is a schematic diagram of the structure of the server provided in the embodiment of the present application;
[0041] Figure 4 It is an optional flowchart of the robot control method provided in the embodiment of the present application;
[0042] Figure 5 It is an optional flowchart of the robot control method provided in the embodiment of the present application;
[0043] Figure 6 It is an optional flowchart of the robot control method provided in the embodiment of the present application;
[0044] Figure 7 It is an optional flowchart of the robot control method provided in the embodiment of the present application;
[0045] Figure 8 is a flow chart of a method for combining backward experience playback provided in an embodiment of the present application;
[0046] Fig. 9 is a flow chart of a method for combining forward and backward experience playback provided in an embodiment of the present application;
[0047] FIG. 10A to FIG. 10H It is a schematic diagram of the test process under different tasks using the method of an embodiment of the present application. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0049] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments, but it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as those commonly understood by those skilled in the art of the technical field of the embodiments of the present application. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0050] Before explaining the solutions of the embodiments of the present application, the nouns and special terms involved in the embodiments of the present application are first explained:
[0051] 1) Reinforcement learning: It belongs to the category of machine learning and is usually used to solve sequential decision-making problems. It mainly includes two components: the environment and the agent. The agent selects actions to perform based on the state of the environment. The environment transfers to a new state based on the action of the agent and feeds back a numerical reward. The agent continuously optimizes its strategy based on the reward fed back by the environment.
[0052] 2) Offline strategy: This is a method different from the action strategy for collecting data and the updated target strategy in reinforcement learning. Offline strategy usually requires the use of experience replay technology.
[0053] 3) Experience replay: It is a technique used by offline policy algorithms in reinforcement learning. An experience pool is maintained to store data on the interaction between the agent and the environment. When training the policy, data is sampled from the experience pool to train the policy network. The experience replay method makes the data utilization efficiency of offline policy algorithms higher than that of online policy algorithms.
[0054] 4) Multi-objective reinforcement learning: The usual reinforcement learning task is to complete a specific task, but in robot control there are often a large number of tasks, such as moving a robotic arm to a position in space. It is hoped that the strategy learned by the agent can reach any target in the target space, so multi-objective reinforcement learning is introduced. Multi-objective reinforcement learning refers to completing multiple goals at the same time.
[0055] 5) Backward experience replay: refers to a method for multi-objective reinforcement learning. By modifying the expected goal of the data in the experience pool to the completed goal, backward experience replay can greatly improve the utilization efficiency of failed data.
[0056] 6) Forward experience replay: The idea of forward experience replay comes from Monte Carlo and temporal difference value function estimation, which accelerates the estimation of the value function by unfolding the expected cumulative rewards for multiple steps.
[0057] 7) Offline policy bias: When the forward experience replay method is used directly in the offline policy algorithm, due to the difference between the behavioral strategy and the target strategy, the usual forward experience replay will lead to the accumulation of offline policy bias, which may seriously affect the policy learning of the intelligent agent.
[0058] Before explaining the embodiments of the present application, the robot control method in the related art is first described:
[0059] Figure 1A It is a flowchart of the implementation of a robot control method in the related technology. The method is a deep reinforcement learning control algorithm based on a priority experience replay mechanism. It uses the state information of the objects operated by the robot to calculate the priority, and uses the deep reinforcement learning method to complete the end-to-end robot control model. This method allows the deep reinforcement learning agent to autonomously learn in the environment and complete the specified tasks. During the training process, the state information of the target object is collected in real time to calculate the priority of the experience replay, and then the data in the experience replay pool is sampled and learned by the reinforcement learning algorithm according to the priority to obtain the control model. This method maximizes the use of environmental information, improves the effect of the control model and accelerates the speed of learning convergence while ensuring the robustness of the deep reinforcement learning algorithm. Among them, such as Figure 1A As shown, the method comprises the following steps:
[0060] Step S11, building a virtual environment.
[0061] Step S12, obtaining sensor data during the robot mission process.
[0062] Step S13, obtaining the environmental state parameters during the robot task process and constructing a sample trajectory set.
[0063] Step S14, calculate the sample trajectory priority, which is composed of three parts: position change, angle change, and speed change of the material.
[0064] Step S15, performing sampling training according to the sample trajectory priority.
[0065] Step S16, determining whether the network update has reached a preset number of steps.
[0066] Step S17: if yes, complete the training process and obtain the reinforcement learning model.
[0067] Figure 1BIt is a schematic diagram of the implementation flow of another robot control method in the related technology. The method discloses a kinematic self-grasping learning method and system based on a simulated industrial robot, which belongs to the field of computer-aided manufacturing. The method is based on a simulation environment and uses reinforcement learning theory to perform robot grasping training. The simulated robot automatically obtains the position information of the object through the image taken by the camera, and determines the grasping position of the robot's end grasping tool; at the same time, the image processing method based on reinforcement learning determines the posture of the grasping tool according to the shape and placement of the grasped object in the observed image, and finally successfully grasps objects of various shapes and randomly placed; the grasping technology of this method can be applied to many industrial and life scenarios. It can simplify the complexity of the grasping work programming of traditional robots, improve the scalability of robot programs, and greatly improve the application range of robots and work efficiency in actual production. Among them, such as Figure 1B As shown, the entire robot control system includes a robot simulation environment 11, a value estimation network 12 and an action selection network 13. Through the interaction between the robot simulation environment 11, the value estimation network 12 and the action selection network 13, the training of the network in the entire system is realized.
[0068] However, the above two methods in the related technology have at least the following problems: the related technology usually requires training a model to complete a feature task, and the generalization is not strong; the related technology does not utilize the backward experience playback information, and often cannot learn from failed data; the related technology does not utilize the forward experience playback information, and often uses a single-step temporal difference method for training, which has low training efficiency and low accuracy of the trained intelligent agent.
[0069] The embodiment of the present application proposes a robot control method, which is a multi-objective reinforcement learning robot control technology that combines forward experience replay and backward experience replay. This method can greatly improve the utilization efficiency of the data for agent training, and at the same time can alleviate the impact of offline policy deviations. The method of the embodiment of the present application can train a large number of targets at the same time, and one model can complete all tasks in a certain target space; and backward experience replay is used to improve the utilization of failed data and accelerate the training of robot tasks; at the same time, the multi-step reward expansion using forward experience replay accelerates the learning of value functions and the training of agents.
[0070] The robot control method provided by the embodiment of the present application first obtains environmental interaction data, which includes at least state data, action data, reward value and target value at two adjacent moments; the actual target value actually completed after executing the action corresponding to the action data; then, according to the state parameters, action data and actual target value of the first moment of the two adjacent moments, the reward value after executing the action is determined; the reward value in the environmental interaction data is updated by the reward value after executing the action to obtain the updated environmental interaction data; then, the intelligent agent corresponding to the robot control network is trained by the updated environmental interaction data; finally, the action of the target robot is controlled by the trained intelligent agent. In this way, the backward experience playback method is used to improve the utilization rate of data and accelerate the training of the intelligent agent. Since the environmental interaction data includes the target value, a large number of targets can be trained at the same time, and all tasks in a certain target space can be completed by one model.
[0071] The following describes an exemplary application of the robot control device of the embodiment of the present application. In one implementation, the robot control device provided by the embodiment of the present application can be implemented as any electronic device or intelligent agent itself, such as a laptop computer, a tablet computer, a desktop computer, a mobile device (e.g., a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable gaming device), an intelligent robot, etc. In another implementation, the robot control device provided by the embodiment of the present application can also be implemented as a server. The following describes an exemplary application when the robot control device is implemented as a server. The intelligent agent can be trained through the server, and the action of the target robot can be controlled by the trained intelligent agent.
[0072] See also Figure 2 , Figure 2It is an optional architecture diagram of the robot control system 10 provided in the embodiment of the present application. In order to realize the training of the intelligent agent, the robot control system 10 provided in the embodiment of the present application includes a robot 100, an intelligent agent 200 and a server 300, wherein the server 300 obtains the environmental interaction data of the robot 100, and the environmental interaction data at least includes the state data, action data, reward value and target value of two adjacent moments, wherein the state data can be the state data of the robot obtained by the robot 100 through the sensor, the action data is the data corresponding to the action performed by the robot 100, the reward value is the return value obtained by the robot after performing the action, and the target value is the preset target to be achieved by the robot. The server 300 further obtains the actual target value actually completed by the robot 100 after performing the action corresponding to the action data. After acquiring the environmental interaction data and the actual target value, the server 300 determines the reward value after the robot 100 performs the action based on the state parameters, action data and actual target value at the first moment of two adjacent moments; uses the reward value after performing the action to update the reward value in the environmental interaction data to obtain updated environmental interaction data; uses the updated environmental interaction data to train the intelligent agent 200 corresponding to the robot control network; after completing the training of the intelligent agent 200, uses the trained intelligent agent 200 to control the action of the target robot.
[0073] The robot control method provided in the embodiment of the present application also relates to the field of artificial intelligence technology, and can be realized at least by computer vision technology and machine learning technology in artificial intelligence technology. Among them, computer vision technology (CV, Computer Vision) is a science that studies how to make a machine "see", and further, it refers to machine vision such as using cameras and computers to replace human eyes to identify and measure targets, and further do graphic processing to make computer processing become more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR, Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, synchronous positioning and map construction, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition. Machine Learning (ML) is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithmic complexity theory and other disciplines. It specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning generally include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning. In the embodiment of the present application, the response to the network structure search request is realized by machine learning technology to automatically search for the target network structure, and to realize the training and model optimization of the controller and the score model.
[0074] Figure 3 is a schematic diagram of the structure of the server 300 provided in the embodiment of the present application, Figure 3 The server 300 shown includes: at least one processor 310, a memory 350, at least one network interface 320 and a user interface 330. The various components in the server 300 are coupled together via a bus system 340. It is understood that the bus system 340 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 340 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 340 is not described in detail. Figure 3 Various buses are labeled as bus system 340 .
[0075] The processor 310 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0076] The user interface 330 includes one or more output devices 331 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 330 also includes one or more input devices 332, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0077] The memory 350 may be removable, non-removable or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drive, optical disk drive, etc. The memory 350 may optionally include one or more storage devices physically away from the processor 310. The memory 350 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 350 described in the embodiment of the present application is intended to include any suitable type of memory. In some embodiments, the memory 350 can store data to support various operations, and examples of these data include programs, modules, and data structures or subsets or supersets thereof, as exemplarily described below.
[0078] Operating system 351, including system programs for processing various basic system services and performing hardware-related tasks, such as framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0079] A network communication module 352, for reaching other computing devices via one or more (wired or wireless) network interfaces 320, exemplary network interfaces 320 include: Bluetooth, Wireless Compatibility Authentication (WiFi), and Universal Serial Bus (USB);
[0080] The input processing module 353 is used to detect one or more user inputs or interactions from one of the one or more input devices 332 and translate the detected inputs or interactions.
[0081] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 3 A robot control device 354 stored in the memory 350 is shown. The robot control device 354 may be a robot control device in the server 300. It may be software in the form of a program or a plug-in, and includes the following software modules: a first acquisition module 3541, a second acquisition module 3542, a determination module 3543, an update module 3544, a training module 3545, and a control module 3546. These modules are logical, and thus may be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.
[0082] In other embodiments, the device provided in the embodiments of the present application can be implemented in hardware. As an example, the device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the robot control method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components.
[0083] The robot control method provided by the embodiment of the present application will be described below in conjunction with the exemplary application and implementation of the server 300 provided by the embodiment of the present application. Figure 4 , Figure 4 is an optional flow chart of the robot control method provided in the embodiment of the present application, which will be combined with Figure 4 The steps shown are explained.
[0084] Step S401, obtaining environmental interaction data, where the environmental interaction data at least includes state data, action data, reward value and target value at two adjacent moments.
[0085] Here, the state data may be the state data of the robot acquired by the robot through sensors or the state data of the environment in which the robot is currently located. The action data is the data corresponding to the action performed by the robot, which may be the action that the robot has performed at a moment before the current moment, or the action that is about to be performed at the next moment after the current moment, wherein the action may be any action that the robot can perform, such as moving, grasping, sorting, etc.
[0086] It should be noted that there is a mapping relationship between the state set corresponding to the state data of the environment and the action set corresponding to the action data, that is, when the robot observes a certain state in the environment, it needs to issue a specific action, and in each state, the robot has different probabilities of issuing different actions. For example, in Go, the state set of the environment consists of all possible chess situations, and the action set of the robot (such as AlphaGo) is all the moves that AlphaGo can take that comply with the rules. At this time, the strategy is AlphaGo's behavior, that is, the chess plan that AlphaGo chooses when facing different situations.
[0087] The reward value is the reward value obtained by the robot after performing an action, that is, the reward value is the reward value obtained based on the robot's actions during the reinforcement learning process. The purpose of reinforcement learning is to find an optimal strategy so that the robot receives a higher cumulative reward value after issuing a series of actions.
[0088] The target value is a preset target that the robot is to achieve. In the embodiment of the present application, there can be multiple target values.
[0089] Step S402, obtaining the actual target value actually achieved after executing the action corresponding to the action data.
[0090] Here, the actual target value refers to the target reached by the robot after executing the action corresponding to the action data. The target is the actual target value at the current moment. There may be a certain deviation between the actual target value and the expected target value (that is, the target value in the environmental interaction data). When there is a deviation, it is necessary to continue learning to perform further actions to make the actual target value approach the expected target value.
[0091] Step S403, determining the reward value after executing the action according to the state parameters, action data and actual target value at the first moment of two adjacent moments.
[0092] Here, a preset reward function may be used to determine the reward value after executing the action based on the state parameters, action data, and actual target value at the first moment of two adjacent moments.
[0093] The first moment is the moment before the action is executed. According to the state parameters before the action is executed, the action to be performed by the robot and the actual target value actually completed after the action is executed, the deviation between the actual target value and the expected target value is determined, and then the current reward value after the action is executed is determined based on the deviation.
[0094] In some embodiments, when the goal is achieved, that is, there is no deviation between the actual target value and the expected target value or the deviation is less than a threshold, the current reward value is 0; when the goal is not achieved, that is, the deviation between the actual target value and the expected target value is greater than or equal to the threshold, the current reward value is -1.
[0095] Step S404: Use the reward value after executing the action to update the reward value in the environment interaction data to obtain updated environment interaction data.
[0096] Here, the reward value after executing the action is accumulated with the reward value corresponding to the robot's historical actions, so as to update the reward value in the environment interaction data and obtain the updated environment interaction data. The updated environment interaction data has new state data, new action data, new reward value and new target value, wherein the new state data in the updated environment interaction data is the state data of a new environment entered by the robot after executing the action. For example, when the action performed by the robot is translation, the new state data is the position and posture of the robot after translation. The new action data is the next action to be performed by the robot after the action is performed, determined according to the new reward value, wherein the result after the execution of multiple consecutive actions is to make the final result closer to the expected target value. The new reward value is the cumulative reward value between the reward value after the action is executed and the reward value corresponding to the robot's historical actions.
[0097] For example, when a robot is trying to complete a goal, the current action may not complete the given goal, but complete other goals. Then, after the robot completes this action, it will reselect a goal from a backward perspective. The second goal is different from the first goal. The second goal is basically the goal that the robot can achieve. Because the previous goal may be too high, a lower goal is set the second time, that is, through multiple executions, to achieve the expected target value that the robot wants to achieve in the end.
[0098] Step S405: Use the updated environment interaction data to train the intelligent agent corresponding to the robot control network.
[0099] In an embodiment of the present application, while using the updated environmental interaction data to train the intelligent agent, the environmental interaction data before the update can also be used to train the intelligent agent at the same time. In other words, the new and old data are used together to train the intelligent agent. In this way, backward experience playback is used to improve the utilization of failure data (i.e., historical environmental interaction data that did not successfully achieve the expected goal), thereby accelerating the training of robot tasks.
[0100] Step S406, using the trained intelligent agent to control the action of the target robot.
[0101] The robot control method provided in the embodiment of the present application obtains environmental interaction data, wherein the environmental interaction data includes at least state data, action data, reward value and target value at two adjacent moments, determines the reward value after executing the action according to the state parameters, action data and actual target value of the action at the first moment of the two adjacent moments, and updates the reward value in the environmental interaction data, that is, utilizes the backward experience playback method to improve data utilization and accelerate the training of the intelligent agent, and because the environmental interaction data includes the target value, a large number of targets can be trained at the same time, and all tasks in a certain target space can be completed by one model.
[0102] In some embodiments, the robot control system includes a robot, an agent and a server, wherein the robot can perform arbitrary actions, such as grasping and moving, and the agent can reach any target in the target space according to the learned strategy, that is, control the robot so that the robot achieves the action corresponding to the specific target.
[0103] Figure 5 is an optional flow chart of the robot control method provided in the embodiment of the present application, such as Figure 5 As shown, the method comprises the following steps:
[0104] Step S501: The robot collects and obtains environmental interaction data, which includes at least state data, action data, reward value and target value at two adjacent moments.
[0105] Here, the environmental interaction data can be collected by the sensors carried by the robot itself, or the robot can obtain the environmental interaction data collected by external sensors.
[0106] Step S502: The robot sends the collected environmental interaction data to the server.
[0107] Step S503: the server obtains the actual target value actually achieved by the robot after executing the action corresponding to the action data.
[0108] Step S504: The server determines the reward value after executing the action according to the state parameters, action data and actual target value at the first moment of two adjacent moments.
[0109] Step S505: The server updates the reward value in the environment interaction data using the reward value after the action is performed, and obtains updated environment interaction data.
[0110] Step S506: The server uses the updated environment interaction data to train the agent corresponding to the robot control network.
[0111] It should be noted that steps S503 to S506 are the same as the above-mentioned steps S402 to S405, and will not be repeated in the embodiment of the present application.
[0112] In the embodiment of the present application, the intelligent agent can be a software module in the server or a hardware structure independent of the server. The server trains the intelligent agent to obtain an intelligent agent that can effectively and accurately control the robot, and uses the trained intelligent agent to control the robot, thereby avoiding the problem of waste of network resources caused by the server controlling the robot in real time.
[0113] Step S507, using the trained intelligent agent to control the action of the target robot.
[0114] Step S508: The robot performs specific actions based on the control of the intelligent agent.
[0115] It should be noted that the embodiment of the present application is based on reinforcement learning technology to train the intelligent agent. Therefore, through gradual training and learning, the trained intelligent agent can accurately control the robot, and the robot can accurately achieve the user's desired goals, thereby improving the robot's work efficiency and work quality. In addition, since robots can be used to replace manual operations in many cases in industrial production, the intelligent agent trained by reinforcement learning can achieve robot control with the same actions as manual operations, thereby improving industrial production efficiency and production accuracy.
[0116] based on Figure 4 , Figure 6 is an optional flow chart of the robot control method provided in the embodiment of the present application, such as Figure 6 As shown, step S405 can be implemented by the following steps:
[0117] Step S601, at each moment, according to the target value in the updated environment interaction data, control the intelligent agent to execute the action data in the updated environment interaction data to obtain the state data of the next moment and obtain the reward value of the next moment.
[0118] Here, since the robot will perform the action corresponding to the action data in the current environment interaction data at each moment, it will get a reward value after performing the action at each moment, and the reward value will be added to the reward value in the current environment interaction data. At the same time, other data in the environment interaction data is updated. That is, as the robot continues to perform actions, it realizes the process of iterative optimization of different data in the environment interaction data.
[0119] Step S602, obtaining the reward values of all future moments after the next moment.
[0120] Here, the reward value at a future moment refers to the expected reward value that is expected to be obtained. The expected reward value at each future moment after the next moment can be preset, wherein the expected reward value corresponds to the expected target value.
[0121] Step S603, determining the cumulative reward value corresponding to all the reward values at future moments. Here, the cumulative reward value refers to the cumulative sum of the expected reward values at future moments.
[0122] Step S604, control the training process of the agent with maximizing the cumulative reward value as the control target. In the embodiment of the present application, the training of the agent is realized based on the forward experience playback technology, and the maximization of the cumulative reward value is to maximize the expected reward value at the future moment, thereby ensuring that the robot's action can be closer to the expected target value.
[0123] In some embodiments, step S604 may be implemented by the following steps:
[0124] Step S6041, determine the expected cumulative reward of the cumulative reward value. Step S6042, calculate the initial action value function according to the expected cumulative reward. Step S6043, use the environmental interaction data of multiple consecutive moments to expand the initial action value function to obtain the expanded action value function, so as to accelerate the learning of the action value function and realize the control of the training process of the intelligent agent.
[0125] Here, in step S6043, forward experience playback can be used to expand the initial action-value function to accelerate the learning of the action-value function and control the training process of the intelligent agent.
[0126] In some embodiments, the expected reward value and a preset discount factor of each future moment in multiple consecutive future moments after the current moment can be obtained; then, the expanded action-value function is obtained based on the discount factor and the expected reward value of each future moment.
[0127] In other embodiments, the weight of the action value function can also be obtained; wherein the value of the weight is greater than 0 and less than 1; then, by replaying the forward experience and utilizing the environmental interaction data of multiple consecutive future moments, the initial action value function is expanded based on the weight to obtain the expanded action value function.
[0128] Here, the weight-based initial action value function is implemented by the following formula (1-1):
[0129]
[0130] Among them, Q target (n) (λ) represents the action value function after expansion based on weight λ, Qtarget (i) represents the initial action-value function.
[0131] Figure 7 is an optional flow chart of the robot control method provided in the embodiment of the present application, such as Figure 7 As shown, the method comprises the following steps:
[0132] Step S701, obtaining environmental interaction data, where the environmental interaction data at least includes state data, action data, reward value and target value at two adjacent moments.
[0133] Step S702, obtaining the actual target value actually achieved after executing the action corresponding to the action data.
[0134] Step S703, determining the reward value after executing the action according to the state parameters, action data and actual target value at the first moment of two adjacent moments.
[0135] Step S704, determining the action data at the next moment according to the expanded action-value function.
[0136] Step S705: Use the action data at the next moment to update the action data in the environment interaction data to obtain updated environment interaction data.
[0137] In an embodiment of the present application, after obtaining the action value function, an action that can increase the reward value is selected from multiple actions as the target action, and the data corresponding to the target action is updated to the environmental interaction data as the action data at the next moment to achieve further updating of the action.
[0138] Step S706, when the current reward value is used to update the reward value in the environment interaction data, the execution strategy of the agent is determined according to the accumulated reward value.
[0139] Step S707, selecting action data at the next moment according to the execution strategy.
[0140] Step S708, updating the action data at the next moment into the environment interaction data to obtain updated environment interaction data.
[0141] In some embodiments, after the agent performs the action, the state of the environment in which the agent is currently located is transferred to the state at the next moment, wherein the state at the next moment corresponds to the state parameter at the next moment; correspondingly, the method further includes:
[0142] Step S709, updating the state parameters at the next moment into the environment interaction data to obtain updated environment interaction data.
[0143] Step S710, using the updated environment interaction data to train the intelligent agent corresponding to the robot control network.
[0144] In an embodiment of the present application, when the environmental interaction data is updated, each data in the environmental interaction data is updated at the same time. In this way, when the updated environmental interaction data is used to train the intelligent agent, it can be ensured that the action determined by the intelligent agent at the next moment is close to the expected target value.
[0145] In some embodiments, there are multiple target values in the environment interaction data, and correspondingly, the method further includes:
[0146] Step S711, determine multiple target values at the next moment.
[0147] Step S712, updating the determined multiple target values at the next moment into the environmental interaction data.
[0148] Step S713, using the trained intelligent agent to control the action of the target robot.
[0149] In the robot control method provided by the embodiment of the present application, there are multiple target values in the environmental interaction data, so that multiple targets can be trained at the same time, that is, a large number of targets can be trained at the same time, so that one model can complete all tasks in a certain target space. For example, multiple targets may include: moving in the direction of Y, moving an X distance, grabbing a specific object during movement, and lifting the specific object after grabbing the specific object. It can be seen that multiple targets can be coherent actions in a series of actions, that is, all tasks in the target space are achieved through one model, thereby completing the execution of a series of actions, making the robot more intelligent.
[0150] The following is an explanation of an exemplary application of the embodiments of the present application in a practical application scenario.
[0151] An embodiment of the present application provides a robot control method. The method of the embodiment of the present application can be applied to multi-objective robot tasks, such as the need to place specified items at different locations in space (logistics, robot sorting, etc.), the movement of robots (aircraft / unmanned vehicles) to specified locations, etc.
[0152] Before explaining the method of the embodiment of the present application, the symbols involved in the present application are first explained:
[0153] Reinforcement learning can usually be expressed as a Markov decision process (MDP). In the embodiment of the present application, a target-expanded MDP is used. The MDP contains a six-tuple (S, A, R, P, γ, G), where S represents the state space, A represents the action space, R represents the reward function, P represents the state transition probability matrix, γ represents the discount factor, and G represents the target space (it should be noted that the target space contains the set of all targets to be achieved, that is, the target space G includes multiple target values g, each target value g corresponds to a target, which is the target to be achieved through reinforcement learning). The agent observes the state s at each moment. t (where t represents the corresponding time), and executes action a according to the state t , the environment receives action a t Then transfer to the next state s t+1 And feedback reward r t , the goal of reinforcement learning optimization is to maximize the cumulative reward value The agent follows the strategy π(a t |s t ) selects an action, the action value function Q(s t ,a t ) represents the state s t Execute action a t The expected cumulative reward after .
[0154] in, E stands for expected value.
[0155] In multi-objective reinforcement learning, the agent's strategy and reward function are both regulated by the goal g. The reward function, value function, and strategy have the following representation: r(s t ,a t ,g),Q(s t ,a t ,g),π(s t ,g). In the embodiment of the present application, the reward function can be set based on success or failure, that is, the reward is 0 when the goal is completed, and the reward is -1 when the goal is not completed. φ represents the mapping from state to goal, and ε represents the threshold for reaching the goal. The reward function can be expressed by the following formula (2-1):
[0156]
[0157] In the embodiment of the present application, the Deep Deterministic Policy Gradient (DDPG) algorithm is implemented based on the Actor Critic architecture, where the Critic part evaluates the state action and the Actor part is the strategy for selecting the action. Under the setting of multi-objective reinforcement learning, the loss function L of the Actor part and the Critic part is actor , L critic Calculated by the following formulas (2-2) to (2-4):
[0158]
[0159] where Q target =r t +γQ(s t+1 ,π(s t+1 ,g),g)(2-4).
[0160] In the embodiment of the present application, forward experience playback refers to the use of continuous multi-step data to expand the value function on the basis of the general offline strategy algorithm update, accelerating the learning of the value function. Figuratively speaking, it allows the intelligent agent to have a forward-looking vision, and the calculation formula immediately replaces Q in the above formula target Formula (2-5) for n-step expansion:
[0161] Q target (n) =r t +γr t+1 +…+γ n Q(s t+1 ,π(s t+1 ,g),g)(2-5).
[0162] Although the method of the embodiment of the present application can accelerate the learning of the value function, if it is applied to an offline policy algorithm, such as the DDPG used here, it will cause offline policy deviation.
[0163] Backward experience playback refers to replacing the failed goal with the actually completed goal in multi-objective reinforcement learning. This is a hindsight approach that brings a backward-looking perspective and can greatly improve the efficiency of data utilization. Figure 8 , which is a flow chart of a method for combining backward experience playback provided in an embodiment of the present application, wherein the method comprises the following steps:
[0164] Step S801, obtaining interaction data with the environment (ie, environment interaction data) (s t ,a t ,r t ,s t+1 ,g).
[0165] Step S802: sampling the actually completed target g'.
[0166] Step S803: recalculate the reward value r′ according to the reward function t = r(s t ,a t ,g').
[0167] Step S804, using the calculated reward value r' t Update to get new environment interaction data (s t ,a t ,r′ t ,s t+1 ,g').
[0168] Step S805: Use the new environment interaction data and the old environment interaction data to train the offline strategy.
[0169] The embodiment of the present application provides a multi-objective reinforcement learning robot control technology that combines forward and backward, which can speed up the training speed and greatly improve the efficiency of data utilization, and can save a lot of unnecessary physical / simulation experimental data in the robot scene. Directly combining the forward technology n-step with the backward technology (HER, Hindsight ExperienceReplay) will be affected by the offline strategy deviation. The weighted average of the n-step with exponentially decreasing weights can be used to alleviate the impact of offline strategy deviation. The method Q weighted with λ weight provided in the embodiment of the present application target (n) (λ) is calculated by the following formula (2-6):
[0170]
[0171] In the method of the embodiment of the present application, when the weight λ is close to 0, Q target (n) (λ) is close to a single-step expansion, at which point Q target (n) (λ) has no offline bias but does not utilize forward information. When λ increases, Q target (n) (λ) contains more n-step forward information, but also brings more bias, so λ can play a role in weighing the forward reward information and offline bias. By adjusting λ and the number of steps n, the forward reward information can be better utilized.
[0172] Fig. 9 : is a flow chart of a method for combining forward and backward experience playback provided in an embodiment of the present application, wherein the method comprises the following steps:
[0173] Step S901, obtaining interaction data with the environment (ie, environment interaction data) (s t ,a t ,r t ,s t+1 ,g).
[0174] Step S902: sampling the actually completed target g'.
[0175] Step S903: recalculate the reward value r′ according to the reward function t = r(s t ,a t ,g').
[0176] Step S904, using the calculated reward value r' t Update to get new environment interaction data (s t ,a t ,r′ t ,s t+1 ,g').
[0177] Here, step S903 to step S904 is the backward technique.
[0178] Step S905: Calculate the multi-step expanded Q according to the new environment interaction data target .
[0179] Step S906, calculate Q target (n) (λ) to update the value function.
[0180] Here, step S905 to step S906 is the forward technique.
[0181] Step S907: Use the new environment interaction data and the old environment interaction data to train the offline strategy.
[0182] The robot control method provided in the embodiment of the present application can be applied to multi-target robot control. Compared with the methods in the related art, it will greatly improve the data utilization efficiency and accelerate the training speed. At the same time, it can learn the strategy to complete the entire target space, which is more generalized than the methods in the related art.
[0183] The following Table 1 compares the implementation results of the method of the embodiment of the present application with the existing method. The eight tasks of the simulation environment Fetch and Hand are used for testing respectively. Fetch represents the operation of the robot arm, and Hand represents the operation of the robot hand. DDPG represents the method in the related art, n-step DDPG represents the forward experience playback, HER represents the backward experience playback, and MHER represents the forward combined with backward method provided in the embodiment of the present application. The comparison result is the average success rate of completing the task after the same number of trainings (on Fetch). It can be seen from the expression that the performance of the method of the embodiment of the present application is the best under the same number of trainings:
[0184] Table 1 Comparison of the implementation results of the method of the present application embodiment and the existing method
[0185]
[0186]
[0187] FIG. 10A to FIG. 10H is a schematic diagram of the test process under different tasks using the method of the embodiment of the present application, wherein, Fig. 10A , is a schematic diagram of hand reaching (HandReach), in which the hand with the dark shadow 1001 must reach with its thumb and a selected finger until they meet at the target position above the palm. Fig. 10B As shown in FIG. 1 , it is a schematic diagram of a hand controlling a cube (HandBlock), where the hand must manipulate a block 1002 until it reaches a desired target position and rotation. Fig. 10C As shown, it is a schematic diagram of a hand operating an egg (HandEgg), the hand must manipulate an egg 1003 or a sphere until it reaches a desired target position and rotation. Fig. 10D As shown, it is a schematic diagram of a hand operating a pen (HandPen), the hand must manipulate a pen 1004 or a wooden stick until it reaches a desired target position and rotation. Fig.10E As shown in FIG. 1 , it is a schematic diagram of a robot reaching a certain position (FetchReach), and the end effector 1005 of the robot must be moved to the desired target position. Fig.10F As shown in Figure 1, it is a schematic diagram of the robot sliding (FetchSlide). The robot must move in a certain direction so that it will slide and rest on the desired target. Figure 10G As shown in FIG. 1 , it is a schematic diagram of a robot push (FetchPush), where the robot must move a box 1006 until the box 1006 reaches the desired target position. Fig. 10H, is a schematic diagram of robot picking (FetchPick), in which the robot must pick up a box 1007 from a table with its gripper and move the box 1007 to a target position above the table.
[0188] It should be noted that in addition to the exponentially decreasing weighted average multi-step expected reward used in the embodiments of the present application, the weights can also be manually designed, or the forward multi-step expected reward (n-step return) can be used directly.
[0189] The following is a description of an exemplary structure of the robot control device 354 provided in the embodiment of the present application implemented as a software module. In some embodiments, Figure 3 As shown, the software module stored in the robot control device 354 of the memory 350 may be the robot control device in the server 300, including:
[0190] The first acquisition module 3541 is used to acquire environmental interaction data, and the environmental interaction data at least includes state data, action data, reward value and target value at two adjacent moments; the second acquisition module 3542 is used to acquire the actual target value actually completed after executing the action corresponding to the action data; the determination module 3543 is used to determine the reward value after executing the action according to the state parameters of the first moment of the two adjacent moments, the action data and the actual target value; the update module 3544 is used to update the reward value in the environmental interaction data with the reward value after executing the action to obtain the updated environmental interaction data; the training module 3545 is used to train the intelligent agent corresponding to the robot control network with the updated environmental interaction data; the control module 3546 is used to control the action of the target robot with the trained intelligent agent.
[0191] In some embodiments, the training module is also used to: at each moment, according to the target value in the updated environmental interaction data, control the intelligent agent to execute the action data in the updated environmental interaction data to obtain the state data of the next moment, and obtain the reward value of the next moment; obtain the reward values of all future moments after the next moment; determine the cumulative reward value corresponding to the reward values of all future moments; and control the training process of the intelligent agent with maximizing the cumulative reward value as the control goal.
[0192] In some embodiments, the training module is also used to: determine the expected cumulative reward of the cumulative reward value; calculate the initial action value function based on the expected cumulative reward; use the environmental interaction data at multiple consecutive moments to expand the initial action value function to obtain the expanded action value function, so as to accelerate the learning of the action value function and realize control of the training process of the intelligent agent.
[0193] In some embodiments, the training module is also used to: obtain the expected reward value and a preset discount factor for each future moment in multiple consecutive future moments after the current moment; and obtain the expanded action value function based on the discount factor and the expected reward value for each future moment.
[0194] In some embodiments, the training module is also used to: obtain the weight of the action value function; wherein the value of the weight is greater than 0 and less than 1; through forward experience playback, using the environmental interaction data of multiple consecutive future moments, expand the initial action value function based on the weight to obtain the expanded action value function.
[0195] In some embodiments, expanding the initial action-value function based on the weight is implemented by the following formula:
[0196]
[0197] Among them, Q target (n) (λ) represents the action value function after expansion based on weight λ, Q target (i) represents the initial action-value function.
[0198] In some embodiments, the device also includes: an action data determination module, used to determine the action data at the next moment according to the expanded action value function; a second update module, used to use the action data at the next moment to update the action data in the environment interaction data to obtain updated environment interaction data; the training module is also used to use the updated environment interaction data to train the intelligent agent corresponding to the robot control network.
[0199] In some embodiments, the device also includes: an execution strategy determination module, which is used to determine the execution strategy of the agent according to the cumulative reward value when the current reward value is used to update the reward value in the environmental interaction data; a selection module, which is used to select the action data at the next moment according to the execution strategy; and a third update module, which is used to update the action data at the next moment into the environmental interaction data to obtain the updated environmental interaction data.
[0200] In some embodiments, after the agent executes the action, the state of the environment in which the agent is currently located is transferred to the state at the next moment, wherein the state at the next moment corresponds to the state parameters at the next moment; the device also includes: a fourth update module, used to update the state parameters at the next moment to the environmental interaction data to obtain the updated environmental interaction data.
[0201] In some embodiments, there are multiple target values, and the device also includes: a simultaneous determination module, used to simultaneously determine the multiple target values at the next moment when using the updated environmental interaction data to train the intelligent agent corresponding to the robot control network; a fifth update module, used to update the determined multiple target values at the next moment into the environmental interaction data.
[0202] It should be noted that the description of the device of the embodiment of the present application is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment, so it is not repeated. For technical details not disclosed in the embodiment of the device, please refer to the description of the method embodiment of the present application for understanding.
[0203] The embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned method of the embodiment of the present application.
[0204] The present application embodiment provides a storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a processor, the processor will execute the method provided by the present application embodiment, for example, Figure 4 The method shown.
[0205] In some embodiments, the storage medium can be a computer-readable storage medium, for example, a ferroelectric random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disk, or a compact disk read-only memory (CD-ROM), etc.; it can also be various devices including one or any combination of the above memories.
[0206] In some embodiments, executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.
[0207] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions). As an example, executable instructions may be deployed to be executed on one computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed at multiple locations and interconnected by a communication network.
[0208] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A robot control method, characterized in that: include: Acquire environmental interaction data, wherein the environmental interaction data at least includes state data, action data, reward value, and target value at two adjacent moments; Obtaining an actual target value actually achieved after executing the action corresponding to the action data; determining a reward value after performing the action according to the state parameter at a first moment of the two adjacent moments, the action data and the actual target value; Using the reward value after executing the action to update the reward value in the environment interaction data, to obtain updated environment interaction data; At each moment, according to the target value in the updated environment interaction data, the agent corresponding to the robot control network is controlled to execute the action data in the updated environment interaction data to obtain the state data at the next moment and obtain the reward value at the next moment; Obtaining reward values for all future moments after the next moment; Determine the cumulative reward value corresponding to the reward values of all future moments; determining an expected jackpot reward of the jackpot value; Calculating an initial action-value function according to the expected cumulative reward; Expanding the initial action-value function using the updated environmental interaction data at a plurality of consecutive moments to obtain an expanded action-value function, so as to accelerate the learning of the action-value function and control the training process of the intelligent agent; The trained agent is used to control the actions of the target robot.
2. The method according to claim 1, characterized in that The method of using the updated environmental interaction data at a plurality of consecutive moments to expand the initial action value function to obtain an expanded action value function includes: Obtain the expected reward value and the preset discount factor for each future moment in a plurality of consecutive future moments after the current moment; The expanded action-value function is obtained according to the discount factor and the expected reward value at each future moment.
3. The method according to claim 1, characterized in that: The method of using the updated environmental interaction data at a plurality of consecutive moments to expand the initial action value function to obtain an expanded action value function includes: Obtaining a weight of the action value function; wherein the value of the weight is greater than 0 and less than 1; By forward experience playback, the updated environmental interaction data at multiple consecutive future moments is used to expand the initial action-value function based on the weight to obtain the expanded action-value function.
4. The method according to claim 3, characterized in that The weight-based expansion of the initial action value function is implemented by the following formula: Among them, Q target (n) (λ) represents the action value function after expansion based on weight λ, Q target (i) represents the initial action-value function.
5. The method according to claim 1, characterized in that The method further comprises: Determining the action data at the next moment according to the expanded action value function; Using the action data at the next moment, updating the action data in the environment interaction data, to obtain updated environment interaction data; The updated environment interaction data is used to train an intelligent agent corresponding to the robot control network.
6. The method according to claim 1, characterized in that The method further comprises: When the reward value after executing the action is used to update the reward value in the environmental interaction data, determining the execution strategy of the agent according to the accumulated reward value; Selecting action data at the next moment according to the execution strategy; The action data at the next moment is updated into the environment interaction data to obtain the updated environment interaction data.
7. The method according to any one of claims 1 to 6, characterized in that: After the agent performs the action, the state of the environment in which the agent is currently located is transferred to the state at the next moment, wherein the state at the next moment corresponds to the state parameter at the next moment; the method further includes: The state parameters at the next moment are updated into the environment interaction data to obtain the updated environment interaction data.
8. The method according to any one of claims 1 to 6, characterized in that: The target value is multiple, and the method further includes: When using the updated environmental interaction data to train an agent corresponding to the robot control network, simultaneously determining a plurality of target values at the next moment; The determined multiple target values at the next moment are updated into the environmental interaction data.
9. A robot control device, characterized in that: include: A first acquisition module, used to acquire environmental interaction data, wherein the environmental interaction data at least includes state data, action data, reward value and target value at two adjacent moments; A second acquisition module is used to acquire an actual target value actually achieved after executing the action corresponding to the action data; a determination module, configured to determine a reward value after executing the action according to the state parameter at the first moment of the two adjacent moments, the action data and the actual target value; An updating module, configured to update the reward value in the environment interaction data using the reward value after executing the action, to obtain updated environment interaction data; A training module, used for controlling the agent corresponding to the robot control network to execute the action data in the updated environment interaction data at each moment according to the target value in the updated environment interaction data, so as to obtain the state data at the next moment and obtain the reward value at the next moment; Obtaining the reward values of all future moments after the next moment; determining the cumulative reward value corresponding to the reward values of all future moments; determining an expected jackpot reward of the jackpot value; Calculating an initial action value function according to the expected cumulative reward; using the updated environmental interaction data at a plurality of consecutive moments to expand the initial action value function to obtain an expanded action value function, so as to accelerate the learning of the action value function and control the training process of the intelligent agent; The control module is used to control the actions of the target robot using the trained intelligent agent.
10. A robot control device, characterized in that: include: A memory for storing executable instructions; A processor, configured to implement the robot control method according to any one of claims 1 to 8 when executing the executable instructions stored in the memory.
11. A computer-readable storage medium, characterized in that: Executable instructions are stored, which are used to cause a processor to execute the executable instructions to implement the robot control method described in any one of claims 1 to 8.
12. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the robot control method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Virtual object behavior strategy training method and device, electronic equipment and storage medium
CN111026272A
Deep reinforcement learning robot control method based on priority experience playback
CN111421538A
Cited By
Robot control method, apparatus and device, and storage medium and program product
EP4183531B1