Robot control method, device, equipment and storage medium
By training the intelligent model in a simulation environment, predicting the robot's next moment's actions, the problem that robots can only train trial and error in the actual physical environment in the prior art is solved, reducing costs and improving efficiency and accuracy.
Patent Information
- Application Number
- CN202411561144.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-04
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-04
AI Technical Summary
The existing robot control method can only be trained and tried and errors in the actual physical environment, and the next action of the robot cannot be predicted, resulting in excessive application scenarios and costs.
By building the robot's simulation environment and simulation control system, split operation tasks, collect joint action images and locations, build state space, use reward functions to train the intelligent model, optimize the model to predict the target action space, and control the robot to perform actions in the real environment.
It realizes training of intelligent models in a simulation environment to predict the robot's next moment's actions, reduces the equipment wear and experimental costs caused by training and trial and error in the actual physical environment, and improves the efficiency and accuracy of robot control.
Smart Images

Figure CN119188767B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robots, and in particular to a control method, device, equipment and storage medium of a robot. Background Art
[0002] At present, robots are widely used in industry, agriculture, medicine, military and other fields. The skill operation and motion planning of robots are the key research directions in the field of robotics. With the continuous development of society, the environment and scenes faced by robots are becoming more and more complex, and the advancement of science and technology has put forward more stringent requirements on the execution efficiency of robots. When traditional robots are learning tasks and skills, they usually collect data, train models and conduct trial and error based on the actual physical environment. In the process, the robot may suffer from problems such as abnormal parameters, wear and damage of the robotic arm. In the prior art, the control of the robot is directly realized by executing the corresponding action instructions issued by the control system, and when executing the action instructions, only the corresponding instructions are executed and the next action cannot be determined based on the current action instructions, resulting in excessive application scenarios and application costs of the robot. Summary of the invention
[0003] The main purpose of the present invention is to solve the technical problem that the existing robot control method can only be trained and tried in an actual physical environment to determine the current action of the robot, and the next action of the robot cannot be predicted.
[0004] A first aspect of the present invention provides a control method for a robot, the method comprising: building a simulation environment and a simulation control system of the robot according to a task environment of the robot and a control system based on joint control; in the simulation environment, performing action decomposition on an operation task based on the simulation control system to obtain an action set of the robot in the simulation environment; controlling the robot to execute the action set according to the simulation control system, and collecting joint action images and joint action positions of the robot to construct a state space of the robot in the simulation environment, wherein the state space includes an action space, a color image and a depth image of a double-arm gripper of the robot; using the state space as training data, and using a reward function to score the operation task to obtain an intelligent agent model, and optimizing the intelligent agent model using a preset loss function, wherein the input of the intelligent agent model is the state space of the robot at the current moment, the output of the intelligent agent model is the action space of the robot at the next moment, and the reward function is used to determine the completion degree of the robot for the operation task, including the product of a distance reward and a corresponding weight and the product of a result reward and a corresponding weight; predicting the target action space of the robot based on the current state space of the robot in the task environment using the intelligent agent model, and controlling the robot to perform corresponding actions based on the target action space.
[0005] Optionally, in a first implementation of the first aspect of the present invention, the state space specifically includes the robot's motion space, a color image of the left arm gripper based on the gripper, a depth image of the left arm gripper based on the gripper, a color image of the right arm gripper based on the gripper, a depth image of the right arm gripper based on the gripper, a depth image of the double-arm gripper based on the side, and a color image of the double-arm gripper based on the side, and the state space is in the form of:
[0006] ],
[0007] in, represents the joint action position of the robot at the tth moment, Represent the left and right arms of the robot respectively, represents a color image, represents the depth image, t represents the image corresponding to the tth moment, l , r , s Respectively represent the left arm gripper, the right arm gripper and the double arm gripper at the side position; the action space specifically includes the left arm joint action position and the right arm joint action position of the robot, and the form of the action space is: ].
[0008] Optionally, in a second implementation method of the first aspect of the present invention, the operation task is scored using a reward function to obtain an intelligent body model, and the intelligent body model is optimized using a preset loss function, including: calculating the robot's completion of the operation task, determining the score of the operation task based on the completion and a preset reward function, and selecting the one with the largest score as the intelligent body model to obtain an intelligent body model; calculating the error value between the predicted joint action position and the collected joint action position based on the preset loss function, and optimizing the intelligent body model based on the error value.
[0009] Optionally, in a third implementation of the first aspect of the present invention, the degree of completion includes a distance reward and a result reward, and the distance reward is: , the result reward is: , the reward function is: ,in, represents the position of the end of the robot arm, Indicates left arm and right arm, Indicates the position of the target object to be gripped. Indicates the target object to be gripped. Indicates successful completion of the task, the corresponding reward is 5, Indicates that the task was not completed successfully, and the corresponding reward is 0. represents the total reward of the task, , Represents the weight of the reward, the distance reward is the distance between the robot's gripper and the target object in the operation task, and the result reward is the result of the robot's completion of the operation task.
[0010] Optionally, in a fourth implementation of the first aspect of the present invention, the loss function is:
[0011] , represents the predicted joint action position, represents the joint action position input to the intelligent model, is a hyperparameter, and .
[0012] Optionally, in a fifth implementation of the first aspect of the present invention, predicting the target action space of the robot using the intelligent agent model based on the current state space of the robot in the task environment includes: acquiring the current state space of the robot in the task environment, inputting the current state space into the intelligent agent model, and obtaining the target action space of the robot in the task environment, wherein the prediction formula of the intelligent agent model is:
[0013]
[0014] , The predicted left arm joint action position and right arm joint action position of the robot in the task environment at the next moment are the target action space. are decoder parameters, represents the action space of the task environment at the current moment of input, The color image and depth image of the robot's double-arm gripper in the task environment at the current moment of the input are represented as the current state space.
[0015] Optionally, in a sixth implementation method of the first aspect of the present invention, after controlling the robot to perform a corresponding action based on the target action space, it also includes: obtaining the execution result of the robot after completing the execution based on the target action space to obtain the first action space of the robot, and collecting the color image and depth image of the robot after the execution to obtain the first state space of the robot; using the first state space as input, predicting the second action space of the robot after completing the execution based on the target action space based on the intelligent body model; and sequentially performing input and output operations on the intelligent body model until the operation task is completed.
[0016] A second aspect of the present invention provides a control device for a robot, the device comprising:
[0017] A building module, used to build a simulation environment and a simulation control system of the robot based on the robot's task environment and control system;
[0018] A splitting module, used for performing action splitting on the operation task in the simulation environment based on the simulation control system to obtain an action set of the robot in the simulation environment;
[0019] An acquisition module, used for controlling the robot to execute the action set according to the simulation control system, and acquiring joint action images and joint action positions of the robot, and constructing a state space of the robot in the simulation environment, wherein the state space includes an action space, a color image and a depth image of the robot's double-arm grippers;
[0020] A training module, used to use the state space as training data, and use a reward function to score the operation task to obtain an intelligent agent model, and use a preset loss function to optimize the intelligent agent model, wherein the input of the intelligent agent model is the state space of the robot at the current moment, the output of the intelligent agent model is the action space of the robot at the next moment, and the reward function is the completion degree of the robot for the operation task, including the product of the distance reward and the corresponding weight and the product of the result reward and the corresponding weight;
[0021] A prediction module is used to predict the target action space of the robot based on the current state space of the robot in the task environment using the intelligent agent model, and control the robot to perform corresponding actions based on the target action space.
[0022] Optionally, in a first implementation of the second aspect of the present invention, the acquisition module includes:
[0023] The state space specifically includes the robot's action space, a color image of the left arm gripper based on the gripper, a depth image of the left arm gripper based on the gripper, a color image of the right arm gripper based on the gripper, a depth image of the right arm gripper based on the gripper, a depth image of the double-arm gripper based on the side, and a color image of the double-arm gripper based on the side. The state space is in the form of:
[0024] ]
[0025] ,in, represents the joint action position of the robot at the tth moment, Represent the left and right arms of the robot respectively, represents a color image, represents the depth image, t represents the image corresponding to the tth moment, l , r , s Respectively represent the left arm gripper, the right arm gripper and the double arm gripper at the side position; the action space specifically includes the left arm joint action position and the right arm joint action position of the robot, and the form of the action space is: ].
[0026] Optionally, in a second implementation of the second aspect of the present invention, the training module includes:
[0027] A scoring unit, used to calculate the robot's degree of completion of the operation task, determine the score of the operation task based on the degree of completion and a preset reward function, and select the one with the largest score as the intelligent agent model to obtain the intelligent agent model;
[0028] The optimization unit is used to calculate the error value between the predicted joint action position and the collected joint action position based on a preset loss function, and optimize the intelligent body model based on the error value.
[0029] Optionally, in a third implementation of the second aspect of the present invention, the scoring unit includes:
[0030] The completion degree includes distance reward and result reward, and the distance reward is: , the result reward is: , the reward function is: ,in, represents the position of the end of the robot arm, Indicates left arm and right arm, Indicates the position of the target object to be gripped. Indicates the target object to be gripped. If you successfully complete the task, you will get a reward of 5. If the task is not completed successfully, the reward is 0. represents the total reward of the task, , Represents the weight of the reward, the distance reward is the distance between the robot's gripper and the target object in the operation task, and the result reward is the result of the robot's completion of the operation task.
[0031] Optionally, in a fourth implementation of the second aspect of the present invention, the optimization unit includes:
[0032] The loss function is:
[0033] , represents the predicted joint action position, represents the joint action position input to the intelligent model, is a hyperparameter, and .
[0034] Optionally, in a fifth implementation of the second aspect of the present invention, the prediction module includes:
[0035] An input unit is used to obtain the current state space of the robot in the task environment, input the current state space into the agent model, and obtain the target action space of the robot in the task environment, wherein the prediction formula of the agent model is:
[0036]
[0037] , The predicted left arm joint action position and right arm joint action position of the robot in the task environment at the next moment are the target action space. are decoder parameters, represents the action space of the task environment at the current moment of input, The color image and depth image of the robot's double-arm gripper in the task environment at the current moment of the input are represented as the current state space.
[0038] Optionally, in a sixth implementation of the second aspect of the present invention, the prediction module further includes:
[0039] The loop unit is used to obtain the execution result of the robot after the execution based on the target action space is completed, to obtain the first action space of the robot, and to collect the color image and depth image of the robot after the execution is completed, to obtain the first state space of the robot; using the first state space as input, based on the intelligent agent model, predicting the second action space of the robot after the execution based on the target action space is completed; and sequentially executing input and output operations on the intelligent agent model until the operation task is completed.
[0040] The third aspect of the present invention provides a robot control device, which includes a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the robot control device executes the robot control method as described above.
[0041] A fourth aspect of the present invention provides a computer-readable storage medium having instructions stored thereon, and when the instructions are executed by a processor, the control method of the robot as described above is implemented.
[0042] In the technical solution provided by the present invention, a simulation environment and a simulation control system of the robot are built, and the operation task is split into actions to obtain an action set, the joint action images and joint action positions of the robot in the process of executing the action set are collected to generate a state space, and an intelligent body model is obtained based on state space and reward function training and reinforcement, and the intelligent body model is optimized using a preset loss function. The intelligent body model is used to predict the target action space of the robot based on the current state space of the robot in the task environment, and the robot is controlled to perform the corresponding action. This solution trains the intelligent body model in a simulation environment, predicts the action of the robot at the next moment, and realizes the control of the robot in a real environment, avoiding the experimental cost problem such as equipment wear caused by continuous training and trial and error in the actual physical environment, reducing the control cost of the robot, and improving the efficiency and accuracy of the robot control. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A schematic diagram of a first embodiment of a robot control method provided by an embodiment of the present invention;
[0044] Figure 2 A schematic diagram of a second embodiment of a control system of a robot provided by an embodiment of the present invention;
[0045] Figure 3 A schematic diagram of a flow chart of a robot control method provided by an embodiment of the present invention;
[0046] Figure 4 A schematic diagram of a training network of an intelligent agent model provided in an embodiment of the present invention;
[0047] Figure 5 A schematic diagram of the structure of a robot control device provided by an embodiment of the present invention;
[0048] Figure 6 Another schematic diagram of the structure of the robot control device provided by the embodiment of the present invention;
[0049] Figure 7 A schematic diagram of the structure of a robot control device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0050] In view of the existing robot control methods, this application builds a simulation environment and simulation control system for the robot, and decomposes the operation task into an action set, collects the robot's joint action images and joint action positions in the process of executing the action set to generate a state space, obtains an intelligent body model based on state space and reward function training and reinforcement, optimizes the intelligent body model using a preset loss function, and uses the intelligent body model to predict the robot's target action space based on the robot's current state space in the task environment, and controls the robot to perform the corresponding action. This solution trains the intelligent body model in a simulation environment, predicts the robot's action at the next moment, and realizes the control of the robot in a real environment, avoiding the experimental cost problem such as equipment wear caused by continuous training and trial and error in the actual physical environment, reducing the robot's control cost, and improving the efficiency and accuracy of the robot's control.
[0051] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, device, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0052] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 A schematic diagram of a first embodiment of a robot control method provided by an embodiment of the present invention, the method specifically comprises the following steps:
[0053] 101. Build the robot's simulation environment and simulation control system according to the robot's task environment and the control system based on joint control.
[0054] This solution does not limit the number of robot arms. When the corresponding number of robot arms is added, the corresponding number of positions and images need to be collected. This embodiment takes a dual-arm robot as an example. First, determine the task environment and robot model, analyze the type of task to be performed by the robot, determine the physical properties of the task environment, select or design the mechanical structure of the robot according to the task requirements, such as the number of joints, joint type (rotation or translation), joint motion range, etc., and define the physical parameters of the robot. Select software suitable for robot simulation, import or create a robot model, and configure the physical parameters of the model. Build a simulation environment similar to the real task environment in the simulation software, and add the same sensors as the real robot to the virtual robot. Select a suitable control algorithm according to the task requirements and integrate it into the simulation system, write a program to control the robot, implement joint-based robot control in the control program, and build a simulation environment and simulation control system for the robot.
[0055] 102. In the simulation environment, the operation task is decomposed based on the simulation control system to obtain the action set of the robot in the simulation environment.
[0056] For performing operation tasks in a simulation environment, first define the operation tasks and clarify the specific tasks that the robot needs to perform, such as grabbing objects, walking, etc., set the initial state of the robot in the simulation environment, including position, posture, speed, etc., start the simulation software, send control instructions to the robot through the simulation control system, and observe the movement state of the robot in the simulation environment. During the simulation process, record the robot's motion data in real time, including joint angles, speeds, accelerations, etc., organize the recorded motion data to form a set of robot actions in the simulation environment, and the action set includes a series of timestamps and corresponding state information of the robot's arms.
[0057] 103. According to the simulation control system, the robot is controlled to execute a set of actions, and the joint action images and joint action positions of the robot are collected to construct the state space of the robot in the simulation environment.
[0058] The state space includes an action space, a color image and a depth image of the robot's double-arm grippers. The joint action position is the action space, including the robot's left arm joint action position and right arm joint action position. The joint action image includes a color image and a depth image, which are respectively a color image of the left arm gripper based on the gripper, a depth image of the left arm gripper based on the gripper, a color image of the right arm gripper based on the gripper, a depth image of the right arm gripper based on the gripper, a depth image of the double-arm gripper based on the side, and a color image of the double-arm gripper based on the side.
[0059] 104. Using the state space as training data and using the reward function to score the operation task, the intelligent agent model is obtained, and the preset loss function is used to optimize the intelligent agent model.
[0060] Among them, the input of the agent model is the state space of the robot at the current moment, the output of the agent model is the action space of the robot at the next moment, and the reward function is used to determine the robot's completion of the operation task, including the product of the distance reward and the corresponding weight and the product of the result reward and the corresponding weight. For scoring the operation task based on the reward function, first clarify the scoring target. The scoring target in this scheme is the completion of the operation task, and then select or design the scoring function. The scoring function is a mathematical function that maps task attributes to score values, that is, the reward function. The input of the reward function is a parameter or feature that can quantify the task performance. Different parameters or features may have different effects on the task performance, and the difference is reflected by assigning different weights. In addition, the reward function should have a clear output range, usually an integer or floating point number between 0 and 100.
[0061] 105. Based on the current state space of the robot in the task environment, the intelligent agent model is used to predict the robot's target action space, and the robot is controlled to perform corresponding actions based on the target action space.
[0062] Specifically, the position of the robot's mechanical arm and joints and the state of the gripper are obtained, and according to the motion instructions issued by the robot, the amount of deviation from the initial position of the robot after executing the corresponding instructions according to the target action space is obtained, and the robot is controlled to move the amount of deviation from the initial position. In this scheme, the transform.localEulerAngles() function is used to realize the movement of the robot in the joint coordinate system, and the end control point is kept horizontally moving in the axial direction through the transform.Translate() method. The posture information is obtained by using the transformed.localEulerAngles() and transform.localPosition() functions of the control point after the movement, and the inverse dynamics is used to realize the movement of the robot. For controlling the robot to perform corresponding actions in the task environment, first determine the specific task that the robot needs to perform. In this step, the robot's mechanical arm is controlled to reach the corresponding position according to the target action space, and the robot's gripper is controlled to adjust the state, and the environmental information is obtained through the sensor, and the sensor data is processed and analyzed, and obstacles and target objects in the environment are identified. According to the task requirements and the perceived environmental information, the robot's action path and strategy are planned, and the decision results are converted into control signals, and the robot's actions are controlled by the actuator.
[0063] Based on the robot control method recorded in this scheme, after predicting the robot's action space at the next moment, the robot's joints and grippers are controlled to perform corresponding actions and reach corresponding positions, and the prediction is repeated to achieve control of the robot. This scheme builds a simulation environment and simulation control system for the robot, and splits the operation task into action sets, collects joint action images and joint action positions of the robot in the process of executing the action set to generate a state space, and obtains an intelligent body model based on state space and reward function training and reinforcement, optimizes the intelligent body model using a preset loss function, and predicts the robot's target action space based on the robot's current state space in the task environment using the intelligent body model, and controls the robot to perform corresponding actions, avoiding experimental cost issues such as equipment wear caused by continuous training and trial and error in the actual physical environment.
[0064] See also Figure 2 , a schematic diagram of a second embodiment of a robot control method provided by an embodiment of the present invention, the method specifically comprises the following steps:
[0065] 201. Obtain the robot's task environment and operation tasks, build a simulation environment based on the task environment, obtain the robot's three-dimensional data and a control system based on joint control, and build a robot model and a simulation control system.
[0066] The robot control method recorded in this scheme is applied to a robot simulation migration real scene system, wherein the robot simulation migration real scene system includes an action planning design module, a simulation data set acquisition module, a reinforcement learning algorithm training module and a simulation migration real environment robot module, wherein the action planning design module is used to design simulation experiment tasks, and according to the action planning design method, the action of the robot in the simulation environment to complete the simulation experiment task is designed; the simulation data set acquisition module is used to drive the robot in the simulation environment to complete the experimental task according to the planned action of the simulation experiment task, and save the joint action position, color image and depth image during the task; the reinforcement learning algorithm training module is used to train the reinforcement learning algorithm according to the simulation data set, and input the prediction result to the robot in the simulation environment, drive the robot in the simulation environment to complete the task, and check the training effect; the simulation migration real environment robot module is used to migrate the trained strategy to the robot in the real environment, and verify the performance of the strategy in performing tasks in real scenes.
[0067] This embodiment takes a dual-arm robot as an example. Figure 3, a flow chart of the robot control method provided by an embodiment of the present invention, firstly, an experimental application scenario of a dual-arm robot is designed, and data such as the dual-arm robot, task scenario, object size, color, etc. of a real experimental environment are obtained; in a simulation environment, a simulation environment based on joint control is built, and the simulation environment is kept consistent with the real environment as much as possible, and then, according to the motion planning design method, the joint action positions of the dual-arm robot when performing the task are designed; in the simulation environment, the dual-arm robot is driven to perform the task, the joint action positions and visual images of the dual-arm robot are obtained, and a task data set for teaching operation is constructed; a reinforcement learning algorithm is used to perform model training and learning, and an intelligent agent model of the dual-arm robot performing the task is obtained; the intelligent agent model trained in the simulation environment is transferred to the real environment, the joint action positions and images in the real environment are input, the inference output results are output, the dual-arm robot in the real environment is driven to perform the corresponding task, and the experimental results are verified.
[0068] In this solution, the left arm and the right arm of the dual-arm robot are both mechanical arms including a number of joints and grippers. Specifically, the left arm and the right arm are both mechanical arms with 6 degrees of freedom and two-finger grippers. The left arm and the right arm both include a base, 6 joints and 1 two-finger gripper. The angular motion limit range of joints 1, 4, and 6 is 360 degrees, while the angular motion limit range of joints 2, 3, and 5 is 180 degrees. Joints with different angular motion limit ranges are alternately connected, and the opening and closing limit of the two-finger grippers of the left and right arms is 180 degrees. It should be noted that the camera used by the dual-arm robot is a color-depth camera, which can obtain color images and depth images of the robot during the execution of the operation task, involving three cameras of the left arm, the right arm and the side. The side of the dual-arm robot refers to a position similar to the human eye, which is used to capture the image of the overall motion state of the two grippers. The image of the gripper position will move with the gripper movement, and then obtain the picture of the gripper movement, while the side position is fixed, and the image of the two grippers in the motion process can be observed.
[0069] For building simulation environment and simulation control system, obtain the robot's task environment and operation tasks, define the task environment, clarify the environment in which the robot needs to work, including physical space, obstacles, etc., determine the operation tasks, describe in detail the specific tasks that the robot needs to perform, such as handling, grasping, etc., build simulation environment based on task environment, select appropriate robot simulation software according to task requirements, create or import a 3D model similar to the real task environment in the simulation software, including the ground, walls, obstacles, etc., configure environmental parameters to simulate the physical characteristics of the real environment, select 3D modeling software to create a 3D model of the robot, create a detailed model of the robot in the 3D modeling software according to the design parameters and size of the robot, including mechanical structure, sensor position, etc., and export the created robot model to a format supported by the simulation software. Then build the robot model and simulation control system, import the exported robot 3D data into the simulation software, configure the control system parameters in the simulation software, including the parameters of the joint controller, sensor settings, etc., write simulation scripts according to task requirements, and control the movement and behavior of the robot in the simulation environment.
[0070] Specifically, the operation task performed by the dual-arm robot in this example is to lift the lid of the box on the table with the left arm, then pick up the ballpoint pen with the right arm, put it into the box, and finally close the lid with the left arm. According to the task flow of the operation task, the information of the physical environment obtained includes the simulation model of the dual-arm robot, the size of the actual scene, the shape and size of the box, the shape and size of the pencil, etc.
[0071] 202. In the simulation environment, the robot model is controlled to perform operation tasks based on the simulation control system to obtain the action set of the robot in the simulation environment.
[0072] In the simulation environment, a simulation environment based on joint control is built in proportion according to the tasks performed by the dual-arm robot and the information obtained, and the center point of the desktop in the simulation environment is used as the origin of the world coordinate system. Simulation cameras similar to color-depth images are set at the left arm, right arm and side of the dual-arm robot, and corresponding target object models are set, including boxes, ballpoint pens, lids of boxes, etc. It should be noted that the simulation environment based on joint control built in the simulation environment is basically consistent with the parameters of the real environment. In this embodiment, the specific simulation environment can be a MuJoCo simulation environment. For performing operation tasks in the simulation environment, first define the operation tasks, clarify the specific tasks that the robot needs to perform, such as grabbing objects, walking, etc., set the initial state of the robot in the simulation environment, including position, posture, speed, etc., start the simulation software, send control instructions to the robot through the simulation control system, observe the motion state of the robot in the simulation environment, and during the simulation process, record the robot's action data in real time, including joint angles, speeds, accelerations, etc., sort out the recorded action data, and form an action set of the robot in the simulation environment, the action set includes a series of timestamps and corresponding robot arms state information.
[0073] In this embodiment, the action trajectory design method is used to design an execution strategy for a dual-arm robot when performing an operation task: using the left arm to lift the lid of a box on the table, then using the right arm to pick up a ballpoint pen and put it into the box, and finally the left arm to close the lid. The action set of the robot in the simulation environment is determined by the execution strategy.
[0074] 203. When the robot model is executing an action set, a joint action image and a joint action position of each joint of the robot model are obtained to obtain a state space in a simulation environment.
[0075] When the robot model is executing an action set, the dual-arm robot in the simulation environment is controlled and driven to execute the entire action set. In the process of executing the operation task, the joint action positions of the left and right arms when performing the task, including the information of the joints and grippers and the image visual information of the left arm gripper, the right arm gripper and the side positions (including color images and depth images) are continuously saved as a data set in a preset format according to a complete task process as one piece of data. A total of 200 pieces of data need to be obtained.
[0076] The state space consists of the joint action positions, color images and depth images of the dual-arm robot, which are used to be input into the intelligent model for learning. The state space specifically includes the joint action positions of the left and right arms of the dual-arm robot and the positions of the left arm, right arm and side. The format of the state space is as follows:
[0077] ]
[0078] in, represents the joint action position of the dual-arm robot corresponding to the tth time step, Represent the left and right arms of the dual-arm robot, represents a color image, represents the depth image, t represents the image corresponding to the t-th time step, l, r, and s represent the images of the left arm gripper position, the right arm gripper position, and the side position, respectively.
[0079] The action space is the output of the state space predicted by the agent model, and the action space at the current moment is input into the interactive environment to obtain the state information and reward information at the next moment. The action space specifically includes the joint action positions of the left arm and the right arm of the dual-arm robot. The format of the action space is as follows: ].
[0080] 204. The state space of the robot model at the current moment in the simulation environment is used as input data, and the action space at the next moment is used as output data. The preset output formula is trained, and the operation task is scored using the reward function to obtain the intelligent agent model.
[0081] The distance between the end gripper and the target object to be gripped is calculated according to the process of the dual-arm robot performing a task, and the distance reward is determined. The formula for calculating the distance reward is as follows:
[0082]
[0083] In the above formula, represents the position of the end of the robot arm, where It means left arm and right arm; represents the position of the target object to be gripped, where Indicates the target object to be gripped.
[0084] The calculation is performed based on the task result of the dual-arm robot during a task. The task goal in this embodiment is that the ballpoint pen is successfully placed in the box and the lid of the box is successfully closed. The formula for calculating the result reward is as follows:
[0085]
[0086] In the above formula, If you successfully complete the task, you will get a reward of 5; If the task is not completed successfully, the reward will be 0.
[0087] The final reward function is calculated as:
[0088]
[0089] In the above formula, Represents the total reward of the task; , Represents the weight of the task reward.
[0090] The distance reward is the distance between the robot's gripper and the target object in the operation task, the result reward is the result of the robot's completion of the operation task, and the reward function is used to determine the robot's completion degree of the operation task, including the product of the distance reward and the corresponding weight and the product of the result reward and the corresponding weight.
[0091] For the intelligent model to predict output data based on input data, first build an image convolution network, input the color images and depth images of three simulated cameras, output the image feature encoding data of the color images and depth images after feature extraction, and then use the encoder and decoder to input the image feature encoding data of the color images and depth images and the joint action positions, and predict and output the joint action positions of the next time step. Among them, the prediction output formula of the training network of the intelligent model is expressed as:
[0092]
[0093] In the above formula, It is expressed as the predicted joint action position at time t+1, that is, the joint action position at the next moment; is the decoder parameter; Indicates the joint action position at time t of the input, including the left arm and right arm , Represents the color image and depth image at time t. Specifically, for extracting image feature encoding data, the color images of the three simulated cameras are used to construct an image encoder using an image convolutional network to extract the features of the color images. The formula for extracting the features of the color images can be expressed as:
[0094]
[0095] In the above formula, Indicates the extracted Features of the tth color image from the camera; represents the input of the previous layer of the image convolutional network,
[0096] Represents the residual mapping result of the current layer, Indicates the number of layers of the network; the depth images of the three simulated cameras are used to construct an image encoder using an image convolutional network to extract the features of the depth images. The formula for extracting the features of the depth images can be expressed as:
[0097]
[0098] In the above formula, Indicates the extracted Features of the tth depth image of the camera; Represents the input of the previous layer of RestNet50; Represents the residual mapping result of the current layer, Indicates the number of layers in the network.
[0099] In this embodiment, reinforcement learning technology is used to train the model of the imitation learning algorithm, in which the state space and reward are the input of the model, and the model output is the action space. The action space at the current moment will enter the interactive environment, and the state space and reward at the next moment will be obtained through the interaction with the environment. The cycle is repeated to obtain the model parameter value with the maximum reward, and the optimal model is obtained through training. The task performed this time is to use the left arm to lift the lid of the box on the table, then the right arm to pick up the ballpoint pen and put it into the box, and finally the left arm to close the lid. The reinforcement learning algorithm is used to simulate the process of interaction between the intelligent body of the dual-arm robot and the environment, and obtain the action position and image information of each joint of the robot arm. Please refer to Figure 4 , a training network diagram of the intelligent model provided by an embodiment of the present invention, respectively inputs a color image and a depth image into an image convolutional network to extract image features, and then inputs the image features and the joint action positions at the current moment into the encoder and decoder to obtain the joint action positions at the next moment. Among them, the Transformer encoder is one of the core components in the Transformer model, which is used to convert the input sequence into a context vector representation.
[0100] 205. Use the preset loss function to optimize the intelligent agent model.
[0101] The loss function of the training network of the agent model is calculated as:
[0102]
[0103] In the above formula, represents the predicted joint action position; represents the input joint action position, It is a hyperparameter. , which is used to splice into piecewise functions and can control the connection point of the two loss functions.
[0104] 206. Obtain the current state space of the robot in the task environment, and use the current state space as input data of the intelligent agent model, obtain the output data of the intelligent agent model, and obtain the target action space of the robot in the task environment at the next moment.
[0105] Specifically, in a real environment, the initial joint action positions (left arm, right arm) of the dual-arm robot and the color images and depth images obtained by three color-depth image cameras are obtained, input into the intelligent body model, and the joint action positions of the next time step (step 1) are predicted and output, which are applied to the dual-arm robot in the real environment to drive the dual-arm robot to perform the corresponding actions. Then, the predicted joint action positions of the first step plus the color images and depth images obtained by the three color-depth image cameras in the next time step are used as inputs for the new time step (step 2), and input into the strategy for prediction and output. The above steps are continuously repeated until the dual-arm robot in the real environment scene is driven to complete the task.
[0106] Furthermore, the intelligent agent model trained in the simulation environment is transferred to the real environment to drive the dual-arm robot in the real environment to perform the corresponding tasks. The verification experimental results include:
[0107]
[0108] In the above formula, Represents the predicted joint action position of the real environment at time t+1, that is, the joint action position at the next moment; is the decoder parameter; Represents the joint action position of the real environment at input time t, including the left arm and the right arm, Represents the color image and depth image of the real environment at input time t.
[0109] 207. Based on the target action space, the robot is controlled to perform corresponding actions.
[0110] The target action space is mapped to specific instructions or control signals that the robot can execute, including converting the target position into a control signal for the motor, or transmitting the grasping instructions to the robot's hand actuator, etc., and the relevant signals or instructions are sent to the robot to perform the corresponding actions.
[0111] This solution trains the intelligent model in a simulation environment, predicts the robot's next action, and controls the robot in a real environment, avoiding experimental cost issues such as equipment wear caused by continuous training and trial and error in the actual physical environment.
[0112] The control method of the robot in the embodiment of the present invention is described above. The control device of the robot in the embodiment of the present invention is described in detail from the perspective of modular functional entities. Figure 5 , a schematic diagram of a structure of a control device of a robot provided by an embodiment of the present invention, the device comprises:
[0113] A building module 510 is used to build a simulation environment and a simulation control system of the robot based on the task environment and control system of the robot;
[0114] A splitting module 520 is used to split the operation task into actions in the simulation environment based on the simulation control system to obtain an action set of the robot in the simulation environment;
[0115] The acquisition module 530 is used to control the robot to execute the action set according to the simulation control system, and to acquire the joint action images and joint action positions of the robot, and to construct the state space of the robot in the simulation environment, wherein the state space includes the action space, the color image and the depth image of the robot's double-arm gripper;
[0116] A training module 540 is used to use the state space as training data, and use a reward function to score the operation task to obtain an intelligent agent model, and use a preset loss function to optimize the intelligent agent model, wherein the input of the intelligent agent model is the state space of the robot at the current moment, the output of the intelligent agent model is the action space of the robot at the next moment, and the reward function is the completion degree of the robot for the operation task, including the product of the distance reward and the corresponding weight and the product of the result reward and the corresponding weight;
[0117] The prediction module 550 is used to predict the target action space of the robot based on the current state space of the robot in the task environment using the intelligent agent model, and control the robot to perform corresponding actions based on the target action space.
[0118] This solution trains the intelligent model in a simulation environment, predicts the robot's next action, and controls the robot in a real environment. This avoids experimental cost issues such as equipment wear caused by continuous training and trial and error in the actual physical environment, reduces the robot's control cost, and improves the efficiency and accuracy of robot control.
[0119] See also Figure 6 , another structural schematic diagram of a robot control device provided by an embodiment of the present invention, the device comprises:
[0120] A building module 610 is used to build a simulation environment and a simulation control system of the robot based on the task environment and control system of the robot;
[0121] A splitting module 620 is used to split the operation task into actions in the simulation environment based on the simulation control system to obtain an action set of the robot in the simulation environment;
[0122] The acquisition module 630 is used to control the robot to execute the action set according to the simulation control system, and to acquire the joint action images and joint action positions of the robot, and to construct the state space of the robot in the simulation environment, wherein the state space includes the action space, the color image and the depth image of the robot's double-arm gripper;
[0123] A training module 640 is used to use the state space as training data, and use a reward function to score the operation task to obtain an intelligent agent model, and use a preset loss function to optimize the intelligent agent model, wherein the input of the intelligent agent model is the state space of the robot at the current moment, the output of the intelligent agent model is the action space of the robot at the next moment, and the reward function is the completion degree of the robot for the operation task, including the product of the distance reward and the corresponding weight and the product of the result reward and the corresponding weight;
[0124] The prediction module 650 is used to predict the target action space of the robot based on the current state space of the robot in the task environment using the intelligent agent model, and control the robot to perform corresponding actions based on the target action space.
[0125] In this embodiment, the acquisition module 630 includes:
[0126] The state space specifically includes the robot's action space, a color image of the left arm gripper based on the gripper, a depth image of the left arm gripper based on the gripper, a color image of the right arm gripper based on the gripper, a depth image of the right arm gripper based on the gripper, a depth image of the double-arm gripper based on the side, and a color image of the double-arm gripper based on the side. The state space is in the form of:
[0127] ]
[0128] ,in, represents the joint action position of the robot at the tth moment, Represent the left and right arms of the robot respectively, represents a color image, represents the depth image, t represents the image corresponding to the tth moment, l , r , s Respectively represent the left arm gripper, the right arm gripper and the double arm gripper at the side position; the action space specifically includes the left arm joint action position and the right arm joint action position of the robot, and the form of the action space is: ].
[0129] In this embodiment, the training module 640 includes:
[0130] A scoring unit 641 is used to calculate the robot's degree of completion of the operation task, determine the score of the operation task based on the degree of completion and a preset reward function, and select the one with the largest score as the intelligent agent model to obtain the intelligent agent model;
[0131] The optimization unit 642 is used to calculate the error value between the predicted joint action position and the collected joint action position based on a preset loss function, and optimize the intelligent body model based on the error value.
[0132] In this embodiment, the scoring unit 641 includes:
[0133] The completion degree includes distance reward and result reward, and the distance reward is: , the result reward is: , the reward function is: ,in, represents the position of the end of the robot arm, Indicates left arm and right arm, Indicates the position of the target object to be gripped. Indicates the target object to be gripped. If you successfully complete the task, you will get a reward of 5. If the task is not completed successfully, the reward is 0. represents the total reward of the task, , Represents the weight of the reward, the distance reward is the distance between the robot's gripper and the target object in the operation task, and the result reward is the result of the robot's completion of the operation task.
[0134] In this embodiment, the optimization unit 642 includes:
[0135] The loss function is:
[0136] , represents the predicted joint action position, represents the joint action position input to the intelligent model, is a hyperparameter, and .
[0137] In this embodiment, the prediction module 650 includes:
[0138] The input unit 651 is used to obtain the current state space of the robot in the task environment, input the current state space into the agent model, and obtain the target action space of the robot in the task environment, wherein the prediction formula of the agent model is:
[0139]
[0140] , The predicted left arm joint action position and right arm joint action position of the robot in the task environment at the next moment are the target action space. are decoder parameters, represents the action space of the task environment at the current moment of input, The color image and depth image of the robot's double-arm gripper in the task environment at the current moment of the input are represented as the current state space.
[0141] In this embodiment, the prediction module 650 further includes:
[0142] The loop unit 652 is used to obtain the execution result of the robot after completing the execution based on the target action space, obtain the first action space of the robot, and collect the color image and depth image of the robot after the execution, to obtain the first state space of the robot; using the first state space as input, predicting the second action space of the robot after completing the execution based on the target action space based on the intelligent agent model; and sequentially executing input and output operations on the intelligent agent model until the operation task is completed.
[0143] This solution trains the intelligent model in a simulation environment to predict the robot's next action and achieve control of the robot in a real environment, avoiding experimental cost issues such as equipment wear caused by continuous training and trial and error in the actual physical environment.
[0144] above Figure 5-6 The control device of the robot in the embodiment of the present invention is described in detail from the perspective of modular functional entities, and the control device of the robot in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0145] See also Figure 7 As shown, the control device of the robot includes a processor 700 and a memory 701. The memory 701 stores machine executable instructions that can be executed by the processor 700. The processor 700 executes the machine executable instructions to implement the above-mentioned robot control method.
[0146] Further, Figure 7 The control device of the robot shown further includes a bus 702 and a communication interface 703 , and the processor 700 , the communication interface 703 and the memory 701 are connected via the bus 702 .
[0147] Among them, the memory 701 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), for example, at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 703 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 702 can be an ISA bus, a PCI bus or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0148] The processor 700 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 700. The above processor 700 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly embodied as a hardware decoding processor to be executed, or a combination of hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 701 , and the processor 700 reads the information in the memory 701 and completes the method steps of the above-mentioned embodiment in combination with its hardware.
[0149] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the various steps of the robot control method provided in the above-mentioned embodiments.
[0150] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described equipment, devices, and units can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.
[0151] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the whole or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0152] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A robot control method, characterized in that: The control method of the robot comprises: Building a simulation environment and a simulation control system of the robot according to the robot's task environment and a control system based on joint control; In the simulation environment, the operation task is split into actions based on the simulation control system to obtain an action set of the robot in the simulation environment; Controlling the robot to execute the action set according to the simulation control system, collecting joint action images and joint action positions of the robot, and constructing a state space of the robot in the simulation environment, wherein the state space includes an action space, a color image and a depth image of the robot's double-arm grippers; The state space is used as training data, and the operation task is scored using a reward function to obtain an intelligent agent model, and the intelligent agent model is optimized using a preset loss function, wherein the input of the intelligent agent model is the state space of the robot at the current moment, and the output of the intelligent agent model is the action space of the robot at the next moment. The reward function is used to determine the degree of completion of the robot for the operation task, including the product of the distance reward and the corresponding weight and the product of the result reward and the corresponding weight; Predicting a target action space of the robot using the agent model based on a current state space of the robot in the task environment, and controlling the robot to perform a corresponding action based on the target action space; The method of using a reward function to score the operation task to obtain an intelligent body model and using a preset loss function to optimize the intelligent body model includes: calculating the robot's completion of the operation task, determining the score of the operation task based on the completion and a preset reward function, and selecting the task with the largest score as the intelligent body model to obtain the intelligent body model; calculating the error value between the predicted joint action position and the collected joint action position based on a preset loss function, and optimizing the intelligent body model based on the error value.
2. The robot control method according to claim 1, characterized in that: The state space specifically includes the robot's action space, a color image of the left arm gripper based on the gripper, a depth image of the left arm gripper based on the gripper, a color image of the right arm gripper based on the gripper, a depth image of the right arm gripper based on the gripper, a depth image of the double-arm gripper based on the side, and a color image of the double-arm gripper based on the side. The state space is in the form of: ], in, represents the joint action position of the robot at the tth moment, Represent the left and right arms of the robot respectively, represents a color image, represents the depth image, t represents the image corresponding to the tth moment, l , r , s Representing the left arm gripper, right arm gripper and double arm gripper in side position respectively; The action space specifically includes the action position of the robot's left arm joint and the right arm joint, and the form of the action space is: ].
3. The robot control method according to claim 1, characterized in that: The completion degree includes distance reward and result reward, and the distance reward is: , the result reward is: , the reward function is: ,in, represents the position of the end of the robot arm, Indicates left arm and right arm, Indicates the position of the target object to be gripped. Indicates the target object to be gripped. Indicates successful completion of the task, the corresponding reward is 5, Indicates that the task was not completed successfully, and the corresponding reward is 0. represents the total reward of the task, , Represents the weight of the reward, the distance reward is the distance between the robot's gripper and the target object in the operation task, and the result reward is the result of the robot's completion of the operation task.
4. The robot control method according to claim 1, characterized in that: The loss function is: , represents the predicted joint action position, represents the joint action position input to the intelligent model, is a hyperparameter, and .
5. The robot control method according to claim 1, characterized in that: The method of predicting the target action space of the robot by using the agent model based on the current state space of the robot in the task environment includes: The current state space of the robot in the task environment is obtained, and the current state space is input into the agent model to obtain the target action space of the robot in the task environment, wherein the prediction formula of the agent model is: , The predicted left arm joint action position and right arm joint action position of the robot in the task environment at the next moment are the target action space. are decoder parameters, represents the action space of the task environment at the current moment of input, The color image and depth image of the robot's double-arm gripper in the task environment at the current moment of the input are represented as the current state space.
6. The robot control method according to claim 1, characterized in that: After controlling the robot to perform a corresponding action based on the target action space, the method further includes: Acquire the execution result of the robot after the execution based on the target action space to obtain the first action space of the robot, and collect the color image and depth image of the robot after the execution to obtain the first state space of the robot; Taking the first state space as input, predicting a second action space of the robot after executing based on the target action space based on the agent model; The input and output operations on the agent model are performed sequentially until the operation task is completed.
7. A robot control device, characterized in that: The control device of the robot comprises: A building module, used to build a simulation environment and a simulation control system of the robot based on the robot's task environment and control system; A splitting module, used for performing action splitting on the operation task in the simulation environment based on the simulation control system to obtain an action set of the robot in the simulation environment; An acquisition module, used for controlling the robot to execute the action set according to the simulation control system, and acquiring joint action images and joint action positions of the robot, and constructing a state space of the robot in the simulation environment, wherein the state space includes an action space, a color image and a depth image of the robot's double-arm grippers; A training module, used to use the state space as training data, and use a reward function to score the operation task to obtain an intelligent agent model, and use a preset loss function to optimize the intelligent agent model, wherein the input of the intelligent agent model is the state space of the robot at the current moment, the output of the intelligent agent model is the action space of the robot at the next moment, and the reward function is the completion degree of the robot for the operation task, including the product of the distance reward and the corresponding weight and the product of the result reward and the corresponding weight; A prediction module, configured to predict a target action space of the robot using the agent model based on a current state space of the robot in the task environment, and control the robot to perform a corresponding action based on the target action space; The method of using a reward function to score the operation task to obtain an intelligent body model and using a preset loss function to optimize the intelligent body model includes: calculating the robot's completion of the operation task, determining the score of the operation task based on the completion and a preset reward function, and selecting the task with the largest score as the intelligent body model to obtain the intelligent body model; calculating the error value between the predicted joint action position and the collected joint action position based on a preset loss function, and optimizing the intelligent body model based on the error value.
8. A robot control device, characterized in that: The control device of the robot includes a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor calls the instructions in the memory so that the control device of the robot executes the control method of the robot as described in any one of claims 1-6.
9. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by a processor, a control method for a robot as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Robot stirring and grabbing combination method based on deep reinforcement learning
CN112102405A