Robot skill simulation method, device, equipment and storage medium

By constructing a gripper and joint simulation environment and optimizing the prediction formula, the problem of insufficient environmental adaptability in traditional robot skill learning is solved, realizing autonomous action planning and flexible response of robots in complex environments, and reducing the cost of skill simulation.

CN119514337BActive Publication Date: 2025-11-28SOUTH CHINA UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411561143.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-11-28
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

Traditional robot skill learning suffers from problems such as limited action execution and inability to adapt to environmental changes.

Method used

By constructing a gripper simulation environment and a joint simulation environment, the gripper state and joint movement position of the robot operation task are obtained. The robot's prediction formula is generated by optimizing the preset initial prediction formula and loss function, and the robot is driven to perform the corresponding action.

Benefits of technology

It improves the intelligence level of robots and the generalization ability of skill simulation, enabling them to autonomously plan motion trajectories and respond flexibly in complex environments, reducing the cost of skill simulation and improving simulation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119514337B_ABST
    Figure CN119514337B_ABST
Patent Text Reader

Abstract

The application relates to the field of robots and discloses a skill simulation method, device and equipment of a robot and a storage medium. The method comprises the following steps: acquiring the gripper state when the robot completes an operation task and inputting the gripper state into a gripper simulation environment to obtain a first joint action position through simulation simulation; inputting the first joint action position into a joint simulation environment to obtain a second joint action position and a visual image through simulation simulation, taking the second joint action position and the visual image as input data, training and optimizing a preset initial prediction formula through a preset loss function, obtaining a prediction formula, inputting the initial joint action position and the initial visual image of the robot in the joint simulation environment into the prediction formula, obtaining a target joint action sequence, and driving the robot to perform corresponding actions. According to the scheme, the joint action information and the visual information are learned, the action sequence of the robot is predicted, the technical difficulty of the robot in operation behavior learning is reduced, and the intelligent degree and the generalization ability of skill simulation of the robot are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of robots, in particular to a skill simulation method, device and equipment of a robot and a storage medium. BACKGROUND

[0002] With the rapid development of robot technology, robots have shown great application potential in various industries due to their high flexibility and adaptability. However, traditional robot skill learning has many limitations, such as: traditional robots can only repeat the planned motion trajectory, cannot obtain environmental information, can only perform a single task type and are inefficient, and traditional robots lack the ability to adapt and learn autonomously, making it difficult to autonomously plan motion trajectories and make flexible responses and adjustments when facing complex and variable environments or similar tasks. That is, in the prior art, the skill learning of the robot is realized by issuing specific work instructions by the system, and cannot adapt to the actual environment. SUMMARY

[0003] The main purpose of the present application is to solve the technical problems of single action execution and inability to adapt to the environment in the existing skill simulation method of the robot.

[0004] The first aspect of the present application provides a skill simulation method of a robot, the method comprising: constructing a simulation environment based on an operation task, and parsing and splitting the operation task to obtain a gripper state of the robot when completing the operation task, wherein the simulation environment comprises a gripper simulation environment based on gripper control and a joint simulation environment based on joint control; inputting the gripper state into the gripper simulation environment and performing simulation based on the gripper state to obtain a first joint action position of the robot in the gripper simulation environment; inputting the first joint action position into the joint simulation environment and performing simulation based on the first joint action position to obtain a second joint action position and a visual image of the robot in the joint simulation environment; training a preset initial prediction formula using the second joint action position and the visual image as input data, and optimizing the preset initial prediction formula using a preset loss function to obtain a prediction formula of the robot; obtaining an initial joint action position and an initial visual image of the robot in the joint simulation environment, and inputting the initial joint action position and the initial visual image into the prediction formula to obtain a target joint action sequence of the robot in the joint simulation environment, and driving the robot to perform a corresponding action based on the target joint action sequence.

[0005] Optionally, in the first implementation manner of the first aspect, the gripper state includes a position, a rotation vector, and a gripping state of the gripper, and the simulation simulation of the gripper simulation environment is:

[0006]

[0007] wherein, represents the gripper simulation environment based on the gripper control, represents the first joint action position output in the gripper simulation environment based on the gripper control, represents the position and the rotation vector of the gripper, represents the gripping state of the gripper, 0 represents closed, and 1 represents open, respectively represent the left arm and the right arm of the dual-arm robot, represents the time.

[0008] Optionally, in the second implementation manner of the first aspect, the visual image includes a left arm image, a right arm image, and a side image of the robot, and the simulation simulation of the joint simulation environment is:

[0009]

[0010] wherein, represents the joint simulation environment based on the joint control, represents the second joint action position output in the joint simulation environment based on the joint control, represents the visual image of the RGB, respectively represent the left arm, the right arm, and the side of the robot, represents the time.

[0011] Optionally, in the third implementation manner of the first aspect, the training of the preset initial prediction formula with the second joint action position and the visual image as input data includes: obtaining the second joint action position and the visual image output by the joint simulation environment, generating input data based on the second joint action position and the visual image, taking the input data as training data, and training the preset initial prediction formula.

[0012] Optionally, in the fourth implementation manner of the first aspect, the form of the input data is:

[0013]

[0014] wherein, represents the input at time t, represents the combination of the joint action position, ​​combined visual image, d represents a dimension of joint action position, i.e., a dimension of left arm + right arm, represents the i-th camera, ; represents a width x height x channel number of a visual image, represents a time point; the preset initial prediction formula is:

[0015]

[0016] wherein, represents a predicted joint action sequence from time t to time t+k, represents input data at time t, is a decoder parameter, and L represents a sample number, = 0, 1, 2,..., L, is a z latent variable corresponding to the i-th target action sequence.

[0017] Optionally, in a fifth implementation manner of the first aspect, the optimization of the preset initial prediction formula by using the preset loss function to obtain the prediction formula of the robot comprises: calculating an error value between the predicted joint action sequence and the collected joint action sequence based on the preset loss function, and optimizing the preset initial prediction formula based on the error value to obtain the prediction formula of the robot, wherein the preset loss function is:

[0018]

[0019] , represents a predicted joint action sequence from time t to time t+k, represents an input joint action sequence, and the prediction formula is:

[0020]

[0021] wherein, represents a joint action position at time t represents a visual image of the i-th camera at time t; is a decoder parameter, and L represents a sample number, = 0, 1, 2,..., L, is a z latent variable corresponding to the i-th target action sequence, and .

[0022] ​Optionally, in a sixth implementation form of the first aspect of the present application, the driving the robot to perform a corresponding action based on the target joint action sequence comprises: calculating control torques of each joint of the robot based on the target joint action sequence by a preset control calculation algorithm, and driving the robot to perform a corresponding action based on each control torque, wherein the preset control calculation algorithm is:

[0023]

[0024] , denotes a control torque of the joint J at a current time, denotes a proportional gain coefficient of a controller, denotes a derivative gain coefficient of the controller, denotes a predicted action position of the joint J at a next time, denotes an actual action position of the joint J at the next time, denotes an actual speed of the joint J at the next time.

[0025] The second aspect of the present application provides a skill simulation device of a robot, the device comprising:

[0026] a construction module configured to construct a simulation environment based on an operation task, and to parse and split the operation task to obtain a gripper state of the robot when completing the operation task, wherein the simulation environment comprises a gripper simulation environment based on gripper control and a joint simulation environment based on joint control;

[0027] a first simulation module configured to input the gripper state into the gripper simulation environment, and to perform simulation based on the gripper state to obtain a first joint action position of the robot in the gripper simulation environment;

[0028] a second simulation module configured to input the joint action position into the joint simulation environment, and to perform simulation based on the first joint action position to obtain a second joint action position and a visual image of the robot in the joint simulation environment;

[0029] a generation module configured to train a preset initial prediction formula by taking the second joint action position and the visual image as input data, and to optimize the preset initial prediction formula by using a preset loss function to obtain a prediction formula of the robot;

[0030] a driving module, configured to acquire an initial joint action position of the robot in the joint simulation environment and an initial visual image, input the initial joint action position and the initial visual image into the prediction formula, obtain a target joint action sequence of the robot in the joint simulation environment, and drive the robot to perform a corresponding action based on the target joint action sequence.

[0031] Optionally, in the first implementation manner of the second aspect of the present application, the first simulation module comprises: the gripper state comprises a position, a rotation vector and a gripping state of the gripper, and the simulation simulation of the gripper simulation environment is:

[0032]

[0033] wherein, represents the gripper simulation environment based on the gripper control, represents the first joint action position output in the gripper simulation environment based on the gripper control, represents the position and the rotation vector of the gripper, represents the gripping state of the gripper, 0 represents closed, and 1 represents open, respectively represent the left arm and the right arm of the dual-arm robot, represents the time.

[0034] Optionally, in the second implementation manner of the second aspect of the present application, the second simulation module comprises: the visual image comprises a left arm image, a right arm image and a side image of the robot, and the simulation simulation of the joint simulation environment is:

[0035]

[0036] wherein, represents the joint simulation environment based on the joint control, represents the second joint action position output in the joint simulation environment based on the joint control, represents the color image, respectively represent the left arm, the right arm and the side of the robot, represents the time.

[0037] Optionally, in the third implementation manner of the second aspect of the present application, the generation module comprises:

[0038] a training unit, configured to acquire the second joint action position output by the joint simulation environment and the visual image, generate input data based on the second joint action position and the visual image, take the input data as training data, and train a preset initial prediction formula.

[0039] Optionally, in a fourth implementation of the second aspect of the present invention, the training unit includes:

[0040] The input data is in the following format:

[0041]

[0042] ,in, This represents the input at time t. Indicates will The combined joint movement position, for The combined visual image d The dimension representing the position of joint movement, i.e., the dimension of the left arm plus the right arm. Indicates the first One camera, ; This represents the width × height × number of channels of a visual image. Indicates the time; the preset initial prediction formula is:

[0043]

[0044] ,in, This represents the predicted sequence of joint movements from time t to t+k. This represents the input data at time t. Here are the decoder parameters, where L represents the number of samples. =0, 1, 2, ..., L, Let z be the latent variable corresponding to the i-th target action sequence.

[0045] Optionally, in a fifth implementation of the second aspect of the present invention, the generation module further includes:

[0046] An optimization unit is used to calculate the error value between the predicted joint motion sequence and the acquired joint motion sequence based on a preset loss function, and to optimize the preset initial prediction formula based on the error value to obtain the prediction formula of the robot, wherein the preset loss function is:

[0047]

[0048] , This represents the predicted sequence of joint movements from time t to t+k. The input joint motion sequence is represented by the prediction formula:

[0049]

[0050] ,in, Indicates the joint position at time t. represents a visual image of the i-th camera at time t; L represents a number of samples, and = 0, 1, 2, …, L, is a z latent variable corresponding to the i-th target action sequence, and .

[0051] Optionally, in a sixth implementation manner of the second aspect of the present application, the driving module comprises:

[0052] a torque calculation unit, configured to calculate control torques of joints of the robot based on the target joint action sequence through a preset control calculation algorithm, and drive the robot to perform corresponding actions based on the control torques, wherein the preset control calculation algorithm is:

[0053]

[0054] , represents a control torque of joint J at the current time, represents a proportional gain coefficient of the controller, represents a derivative gain coefficient of the controller, represents a predicted action position of joint J at the next time, represents an actual action position of joint J at the next time, represents an actual speed of joint J at the next time.

[0055] The third aspect of the present application provides a skill simulation device of a robot, comprising a memory and at least one processor, the memory storing instructions; the at least one processor invokes the instructions in the memory to enable the skill simulation device of the robot to perform the skill simulation method of the robot as described above.

[0056] The fourth aspect of the present application provides a computer readable storage medium, the computer readable storage medium storing instructions, the instructions being executed by a processor to implement the skill simulation method of the robot as described above.

[0057] The technical solution provided by this invention obtains the gripper state of the robot when completing an operation task and inputs it into a gripper simulation environment for simulation to obtain the first joint action position. The first joint action position is then input into a joint simulation environment for simulation to obtain the second joint action position and a visual image. These are used as input data. A preset initial prediction formula is trained and optimized using a preset loss function to obtain a prediction formula. The robot's initial joint action position and initial visual image in the joint simulation environment are then input into the prediction formula to obtain the target joint action sequence, which drives the robot to execute the corresponding action. This solution predicts the robot's action sequence by learning joint action information and visual information, reducing the technical difficulty of robot learning operation behavior, improving the robot's intelligence and the generalization ability of skill simulation, and enabling the robot to learn operation behavior in a simulation environment. This allows the robot to have the ability to autonomously adapt and learn, thus enabling it to autonomously plan action trajectories, make flexible responses and adjustments when facing complex and changing environments or tasks, improve the robot's fine operation capabilities, reduce the cost of robot skill simulation, and improve the efficiency and convenience of robot skill simulation. Attached Figure Description

[0058] Figure 1 A schematic diagram of the first embodiment of the robot skill simulation method provided in this invention;

[0059] Figure 2 A schematic diagram of a second embodiment of the robot skill simulation method provided in this invention;

[0060] Figure 3 A flowchart illustrating the robot skill simulation method provided in an embodiment of the present invention;

[0061] Figure 4 A schematic diagram of the input and output of the simulation environment and prediction formula provided in the embodiments of the present invention;

[0062] Figure 5 A schematic diagram of a robot skill simulation device provided in an embodiment of the present invention;

[0063] Figure 6 Another structural schematic diagram of the robot skill simulation device provided in an embodiment of the present invention;

[0064] Figure 7 This is a schematic diagram of the robot skill simulation device provided in an embodiment of the present invention. Detailed Implementation

[0065] Compared with the skill simulation method of the existing robot, the joint action information and the visual information are learned, the action sequence of the robot is predicted, the technical difficulty of the robot in operation behavior learning is reduced, the intelligent degree and the generalization ability of the skill simulation of the robot are improved, the learning of the operation behavior of the robot in the simulation environment is realized, the robot can have the self-adaptive and learning ability, so that when facing the complex and changeable environment or task, the robot can independently plan the action track, make flexible response and adjustment, improve the fine operation ability of the robot, reduce the skill simulation cost of the robot, improve the skill simulation efficiency and convenience of the robot, and improve the skill simulation efficiency of the robot.

[0066] The terms "first", "second", "third", "fourth" and the like in the description, claims, and drawings of the present application, and those above and below (if any) are used for distinguishing between similar objects and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of the terms so termed herein is to be taken in context and is not a necessary limitation, unless otherwise understood from the presentation of the application. Moreover, the term "comprising" or "having" and any variations thereof, is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of steps or units not necessarily limited to those explicitly stated, but can include other not expressly listed or inherent to such processes, methods, articles, or apparatus.

[0067] For the sake of understanding, the specific flow of the embodiments of the present application is described below, please refer to Figure 1 The first embodiment of the skill simulation method of the robot provided by the embodiments of the present application is shown in the figure, and the method specifically comprises the following steps:

[0068] 101, a simulation environment is constructed based on an operation task, and the operation task is parsed and split to obtain the gripper state of the robot when completing the operation task.

[0069] The simulation environment includes a gripper simulation environment based on gripper control and a joint simulation environment based on joint control. The number of mechanical arms of the robot is not limited by the scheme, and when the corresponding number of mechanical arms is increased, the corresponding number of positions and images need to be collected. The present embodiment takes a dual-arm robot as an example.

[0070] The differences between gripper-based simulation environments and joint-based simulation environments lie in the following aspects: Controlled objects and functions: The gripper simulation environment controls the operation and control of the gripper, simulating actions such as grasping and releasing, as well as the interaction between the gripper and the object. The joint simulation environment controls the robot's joints, simulating the movement of the robot's joints, including rotation and displacement, to achieve the robot's overall movement and posture adjustment. Modeling and simulation requirements: The gripper simulation environment focuses on parameters such as the gripper's mechanical structure, gripping force, and gripping stability, simulating the gripping effect under different conditions, such as gripping force, gripping speed, and gripping accuracy. The joint simulation environment requires building an overall model of the robot, including the kinematics and dynamics of the joints, to simulate the robot's trajectory, speed, acceleration, and interaction with the environment in different scenarios.

[0071] 102. Input the gripper state into the gripper simulation environment, and perform simulation based on the gripper state to obtain the first joint movement position of the robot in the gripper simulation environment.

[0072] The gripper state is input into the gripper simulation environment. Based on the gripper state when the robot completes the operation task, the gripper's motion trajectory is determined, and the motion trajectory is converted into joint motion positions, obtaining the robot's first joint motion position in the gripper simulation environment. The gripper-based simulation environment is mainly used to convert the motion trajectory of the end effector gripper of the dual-arm robot in the simulation environment into joint motion positions, obtaining the joint motion positions used for task teaching operations.

[0073] Specifically, start the gripper simulation software or platform, load the relevant models of the gripper and the robot, configure the simulation environment to simulate the actual working environment and physical conditions of the gripper, determine the initial state of the gripper, such as the degree of opening and closing, clamping force, etc., and pass the initial state of the gripper as input data to the simulation environment. Based on the results of the actual robot or simulation experiment, define a mapping relationship, which describes the association between the gripper state and the robot joint movement position. Using the defined mapping relationship, calculate the movement position of the first joint based on the input gripper state.

[0074] In another implementable manner, a desired trajectory that a double-arm robot gripper needs to follow in a simulation environment is determined, the desired trajectory being a series of pose points, each point containing position and attitude information of the gripper in a three-dimensional space, a state of the double-arm robot is initialized in the simulation environment, including initial positions and attitudes of all joints, for each pose point on the trajectory, an inverse kinematics algorithm is used to calculate angles or positions that the joints of the robot should be in when reaching the pose point, since the inverse kinematics calculation can produce discontinuous or jittering joint trajectories, the calculated joint trajectories need to be smoothed, and the smoothing method can include interpolation, filtering, spline curve fitting, etc. Further, each first joint action position in the calculated joint trajectory is input into the simulation environment, and whether the end gripper of the double-arm robot can move according to the desired trajectory is observed, if there is deviation or problem, the inverse kinematics algorithm or the smoothing method can be adjusted, and calculation and verification are performed again.

[0075] 103、inputting the first joint action position into the joint simulation environment and performing simulation based on the first joint action position to obtain a second joint action position and a visual image of the robot in the joint simulation environment.

[0076] The second joint action position is a state position of each joint of the robot based on joint control output by the robot after simulation in the joint simulation environment, and the visual image includes a color image of the left arm, the right arm and the side of the robot.

[0077] A joint simulation software or platform is started and a model of the robot is loaded, a simulation environment is configured to simulate physical conditions in the real world, the first joint action position is input, the initial action position of the first joint is input as input data to the simulation environment, and the simulation environment simulates the movement of the robot according to the input first joint action position. During the simulation process, the robot model calculates the action positions of the joints in the joint simulation environment according to physical laws and kinematics equations, and a virtual camera is configured to capture the movement process of the robot and shoot visual images of the robot at different time points.

[0078] 104、using the second joint action position and the visual image as input data to train a preset initial prediction formula, and using a preset loss function to optimize the preset initial prediction formula to obtain a prediction formula of the robot.

[0079] collecting a dataset containing second joint action positions and corresponding visual images, labeling each sample in the dataset to contain a desired or target joint action sequence, e.g., joint angles for the next time step, pre-processing the second joint action position data, e.g., normalization or standardization, pre-processing the visual images, including scaling, cropping, graying, binarization, feature extraction, etc., using the prepared dataset to train the prediction formula, in each iteration, passing the second joint action positions and visual images as input data to the prediction formula and calculating the predicted target joint action sequence, calculating the error between the predicted value and the actual value using a pre-set loss function, calculating the gradient of the loss function with respect to the parameters of the prediction formula using a backpropagation algorithm, and updating the parameters using an optimization algorithm to minimize the loss function.

[0080] 105、obtaining the initial joint action position of the robot in the joint simulation environment and the initial visual image, and inputting the initial joint action position and the initial visual image into the prediction formula to obtain the target joint action sequence of the robot in the joint simulation environment, and driving the robot to perform the corresponding action based on the target joint action sequence.

[0081] In practical applications, first, initialize the robot and the simulation environment, start the joint simulation environment, and load the model of the robot, set the initial conditions, such as the initial joint action position of the robot (initial pose), obtain the initial joint action position and the visual image, obtain the initial joint action position of the robot (usually a set of joint angles) from the simulation environment, if there is a visual sensor simulation, at the same time, obtain the visual image corresponding to the initial state of the robot, pre-process the initial joint action position and the visual image as necessary, such as normalization, denoising, etc., extract features useful for prediction, which may include joint angles, angular velocities, specific target positions in the image, etc., input the pre-processed initial joint action position and visual image features into the prediction formula, use the prediction formula to calculate the target joint action sequence based on the input data, the target joint action sequence is a number of joint action positions, convert the target joint action sequence into instructions that can be understood by the robot controller in the simulation environment, send the instructions to the robot controller in the simulation environment, and drive the robot to perform the corresponding action.

[0082] The scheme obtains the gripper state when the robot completes the operation task, inputs the gripper state into the gripper simulation environment for simulation to obtain the first joint action position, inputs the first joint action position into the joint simulation environment for simulation to obtain the second joint action position and the visual image, takes the second joint action position and the visual image as input data, trains and optimizes the preset initial prediction formula through the preset loss function, obtains the prediction formula, inputs the initial joint action position and the initial visual image of the robot in the joint simulation environment into the prediction formula, obtains the target joint action sequence, and drives the robot to perform the corresponding action. Through learning of the joint action information and the visual information, the robot action sequence is predicted, the technical difficulty of the robot in operation behavior learning is reduced, and the intelligent degree and the generalization ability of skill simulation of the robot are improved.

[0083] Please refer to Figure 2 The second embodiment of the robot skill simulation method provided by the embodiment of the present application is shown in the figure, and the method specifically comprises the following steps:

[0084] 201, a simulation environment is constructed based on an operation task, the gripper state of the robot when completing the operation task is obtained, and a gripper simulation environment based on gripper control and a joint simulation system based on joint control of the robot are built.

[0085] In the scheme, a simulation environment and a robot model are built based on the action environment of the robot; in the simulation environment, a gripper trajectory is determined based on an operation task, and a joint action is determined by inverse kinematics; the joint action is learned and trained based on the joint action, the joint action of the operation task is determined, and the robot action is controlled.

[0086] This embodiment takes a dual-arm robot as an example, please refer to Figure 3 The flowchart of the robot skill simulation method provided by the embodiment of the present application is shown in the figure, first, the task to be performed by the dual-arm robot in the simulation environment and the simulation model of the dual-arm robot are designed and determined; the simulation environment based on gripper control and the two simulation environments based on joint control of the dual-arm robot are built, which are used for teaching operation and obtaining simulation data sets; the end gripper action trajectory calculation method of the dual-arm robot when performing the task in the simulation environment is defined, which is used for calculating and obtaining the complete end gripper action trajectory when performing the task; in the simulation environment based on gripper control, the action trajectory of the end gripper when completing the task is input, the joint action position of the dual-arm robot is output, then the joint action position is input into the simulation environment based on joint control, teaching operation is performed, the joint action position and the visual image are output, and the simulation data set of task teaching is composed; in the simulation environment based on joint control, the simulation data set is input to train the imitation learning algorithm; when the model reasoning is performed in the simulation environment, the initial joint action position and the visual image of the dual-arm robot in the simulation environment are input into the model, the joint action sequence is output through reasoning, and the dual-arm robot is driven to perform the action.

[0087] In this example, the left arm and the right arm of the dual-arm robot are both mechanical arms containing a certain number of joints and grippers. Specifically, the left arm and the right arm of the dual-arm robot are each a 6-DOF mechanical arm, each including a base, 6 joints, and a gripper. The angle motion range of joints 1, 4, and 6 is 360 degrees, and the angle motion range of joints 2, 3, and 5 is 180 degrees. The joints with different angle motion ranges are connected alternately. The opening and closing range of the gripper of the left arm and the right arm is 180 degrees. It should be noted that the camera used by the dual-arm robot is a color camera that can obtain color images of the robot during the execution of an operation task. The left arm, the right arm, and the side camera are involved. The side of the dual-arm robot refers to a position similar to the human eye, which is used to capture images of the overall motion state of the two grippers. The image of the gripper position moves with the motion of the gripper, and then the image during the motion of the gripper is obtained. The side position is fixed and can be used to observe the images of the two grippers during the motion. Specifically, in the simulation environment, a gripper simulation environment based on gripper control and a joint simulation environment based on joint control are built according to the proportion. The center of the rectangular desktop is used as the origin of the world coordinate system, and the dual-arm robot and the corresponding object simulation model are placed. Simulation cameras are arranged at the end positions of the left arm and the right arm of the dual-arm robot. The simulation cameras are fixed at the end positions of the left arm and the right arm and move with the movement of the left arm and the right arm. A camera is arranged at a position slightly higher than the length of the mechanical arm on the side of the left arm and the right arm. The resolution of the images obtained by the three simulation cameras is 480x640x3.

[0088] Specifically, the operation task performed by the dual-arm robot in this example is a process in which the right arm of the dual-arm robot picks up an object between the two arms, transfers it to the left arm, and returns to the initial state, and the left arm picks up the object transferred by the right arm, places it at a specified position, and returns to the initial state. For the simulation environment based on the real physical environment in which the operation task is located, first, the simulation target is determined, the real physical environment that needs to be simulated by the simulation environment is determined, such as mechanical structure, physical characteristics, environmental interaction, etc. Real physical data is collected and modeled, a virtual physical environment is constructed, an accurate model of the robot is constructed, including mechanical structure, sensing system and control system, the robot model is imported into the simulation environment, and physical parameters and sensor data are configured. According to the sensors used in the real physical environment, the corresponding sensor data is simulated.

[0089] 202. The gripper state is input into the gripper simulation environment, the action trajectory of the gripper is determined based on the gripper state when the robot completes the operation task, and the action trajectory is converted into a joint action position to obtain the first joint action position of the robot in the gripper simulation environment.

[0090] The simulation environment based on gripper control is mainly used to convert the motion trajectory of the end gripper of the dual-arm robot in the simulation environment into joint motion positions, so as to obtain the joint motion positions for performing the task demonstration operation. Specifically, the current state of the gripper is determined, including the opening and closing degree of the gripper, the clamping force, etc., the corresponding gripper state parameters are set in the gripper simulation environment, the operation tasks that the robot needs to complete are analyzed, such as grabbing, placing, carrying, etc., the key states that the gripper needs to go through in the task are determined, such as the gripper starting to close, clamping the object, the gripper starting to open, etc., the motion trajectory of the gripper is planned according to the operation task and the gripper state, and the trajectory planning includes the speed, acceleration, path, etc. of the opening and closing of the gripper. According to the position, speed and acceleration of the gripper, the angle, speed and acceleration of each joint of the robot are calculated by using the inverse kinematics calculation method, and the motion trajectory of the gripper is converted into the joint motion position of the robot. The above processing is sequentially performed on each joint to determine the first joint motion position of each joint.

[0091] Further, a trajectory planning method using spatial straight line interpolation is used to define all gripper motion positions of the task for the end gripper position. Taking the end joint positions A (xA, yA, zA) and B (xB, yB, zB) of two key points of the right arm of the dual-arm robot in the gripper simulation environment based on gripper control as an example, A is the previous key point and B is the next key point, and the corresponding time intervals are tA and tB. Here, N=tB-tA is the number of interpolation points, and the end gripper position calculation method of each time step interpolation point is:

[0092]

[0093] Specifically, first, the positions and times of the key points are determined, the number of interpolation points and the time interval are calculated, linear interpolation is performed between points A and B to obtain the position of each interpolation point, an array of size N+1 (including points A and B) is created to store the positions of the interpolation points, and the interpolation point positions are calculated.

[0094] The formula for outputting the first joint motion position of the dual-arm robot is as follows:

[0095]

[0096] In the above formula, indicates that the simulation environment is a simulation environment based on gripper control; indicates the joint motion position output in the simulation environment based on gripper control; indicates the position and rotation vector of the end; indicates the gripping state of the gripper, which is 0 when closed and 1 when opened; indicates the left arm and right arm of the dual-arm robot, respectively; Indicates the time.

[0097] 203. Input the first joint motion position into the joint simulation environment for simulation to obtain the second joint motion position and visual image of the robot in the joint simulation environment.

[0098] The joint-based simulation environment is primarily used to perform task teaching operations on the joint-based simulation environment based on the joint motion positions output from the gripper-based simulation environment. It outputs joint motion positions and visual images, forming a simulation dataset for task teaching. Specifically, the task to be performed is defined in the simulation environment, such as grasping an object or moving an object to a specified position. Joint motion position data is obtained from the gripper-based simulation environment and imported into the joint-based simulation environment for task teaching. During the simulation, if the robot's movements do not match expectations, the joint motion positions are adjusted or optimized. After task teaching is completed, the simulation environment generates a simulation dataset containing joint motion positions, visual images, and other information.

[0099] The calculated complete end-effector motion trajectory during task execution is input into the joint motion positions obtained in the gripper-based simulation environment. Then, the joint motion positions are input into the joint-based simulation environment for teaching operations, and the joint motion positions of the left and right arms of the dual-arm robot and the visual images from the simulation cameras are output, namely, 480×640×3 images from 14D + 3 cameras.

[0100] The formula for outputting the joint position and visual image is:

[0101]

[0102] In the above formula This indicates that the simulation environment is a joint-controlled simulation environment; This represents the joint motion position output in a joint-controlled simulation environment. Represents a visual image of RGB. These represent the left arm, right arm, and side (the side position between the two arms) of the dual-arm robot, respectively. This indicates the time. The side view of a dual-arm robot refers to a position similar to a human eye, used to capture images of the overall movement of the two grippers. The image at the gripper position moves with the grippers, thus acquiring images of the grippers during movement. The side view, however, is fixed, allowing observation of the images of the two grippers during their movement.

[0103] 204. Generate input data based on the second joint movement position and visual image, and input the input data into the initial prediction formula to train the initial prediction formula.

[0104] The joint action position and visual image output by the joint control-based simulation environment are combined into input data; based on the ACT imitation learning algorithm, the joint action position and visual image are input, and a joint action sequence is predicted and output.

[0105] The second joint action position and visual image output by the joint control-based simulation environment are combined into input data as follows:

[0106]

[0107] In the above formula, represents the input at time t; represents the joint action position after combination; is the joint action position after combination; is the visual image after combination; d represents the dimension of the joint action position, i.e., the dimension of the left arm + the right arm; represents the i-th camera, ; represents the width (W) x height (H) x channel number (C) of the visual image, represents the time. The formula for predicting and outputting the joint action sequence is:

[0108] The formula for predicting and outputting the joint action sequence is:

[0109]

[0110] In the above formula, represents the predicted joint action sequence from time t to time t+k; represents the input data at time t; is the decoder parameter, L represents the number of samples, and l=0, 1, 2, …, L; is the z latent variable corresponding to the i-th target action sequence.

[0111] 205. Calculate the error value between the predicted joint action sequence obtained based on the initial prediction formula and the input joint action sequence collected using a preset loss function, and optimize the initial prediction formula based on the error value to obtain the prediction formula of the robot.

[0112] The loss function is used to measure the difference between the predicted value and the actual value. In the prediction of robot joint action, common loss functions include mean square error, mean absolute error, etc. In this scheme, the calculation formula of the loss function of the predicted output joint action sequence is:

[0113] ​​​​

[0114] The difference between the predicted value and the collected data is calculated using the mean square error as the loss function in the above formula, represents the predicted joint action sequence at time t to t+k; represents the input joint action sequence.

[0115] The initial prediction formula and input data are used to calculate the predicted joint action sequence, the error between the predicted value and the actual value is calculated using the defined loss function, and the parameters of the prediction formula are updated using an optimization algorithm (such as gradient descent, stochastic gradient descent, Adam, etc.) to minimize the loss value. This usually involves calculating the gradient of the loss function with respect to the model parameters, and using these gradients to update the parameters.

[0116] Please refer to Figure 4 The simulation environment and input / output schematic diagram of the prediction formula provided by the embodiment of the application first inputs the simulation environment based on end control, and inputs the data output by the simulation environment based on end control into the simulation environment based on joint control for joint teaching operation, and outputs the joint action based on joint control as position and visual image, so as to realize acquisition of simulation data, and input the joint action based on joint control as position and visual image into the ACT algorithm for encoding and decoding, and finally output the joint action sequence. The ACT algorithm, Specifically, the joint action position in the simulation data set is acquired for action blocking, the target action sequence is obtained, the current joint action position, the target action sequence, and the visual information are constructed as the model training input, the CVAE encoder (conditional variational autoencoder) in the ACT imitation learning algorithm is constructed, the joint action position and the target action sequence are input, and the z hidden variable is calculated; the CVAE decoder in the ACT imitation learning algorithm is constructed, including a residual network image encoder, an encoder and a decoder, the joint action position, the visual image of the camera and the z hidden variable are input, the visual image of the camera is first processed using the residual network image encoder to obtain image features, then the joint action position, the z hidden variable data are combined and input into the encoder for calculation and output of comprehensive features for prediction, and then the decoder is used for decoding to output the joint action sequence. It should be noted that the state space of the CVAE encoder in the ACT model learning algorithm, i.e. the input includes the joint action position of the double arms of the double arm robot, which is 7+7=14, and the visual image of the three simulation cameras, which is a color image of 480x640x3.

[0117] In practical applications, data preprocessing and motion segmentation are first performed to obtain a simulation dataset containing joint motion positions and visual information. The continuous sequence of joint motion positions is divided into smaller motion blocks or target motion sequences. The initial joint motion position of each motion block is selected as the current position, and the sequence of joint motion positions within each motion block is selected as the target motion sequence. For each current joint motion position, the corresponding visual image or video frame is acquired. The current joint motion position and the target motion sequence are used as input to the CVAE encoder, encoding the input data into one or more latent variables z. These latent variables capture the underlying structure and relationships of the input data. The decoder's input includes the latent variables z, the current joint motion position, and visual information. An image feature extractor is used to encode the visual image, extracting image features. The image features, the current joint motion position, and the latent variables z are combined and input into the encoder. The encoder uses a self-attention mechanism to capture the dependencies between inputs and outputs a comprehensive feature. The decoder receives the encoder's output as input and uses self-attention and encoder-decoder attention mechanisms to generate the predicted joint motion sequence.

[0118] 206. Input the robot's initial joint motion positions and initial visual images in the joint simulation environment into the prediction formula to obtain the target joint motion sequence.

[0119] In a joint-controlled simulation environment, the initial joint position information (7+7=14D) of the two arms of a dual-arm robot and visual images (480×640×3) from three simulation cameras are acquired and input into the model. This predicts the joint motion sequence (14D) for the next time step (step 1). Using the Stable PD control calculation method, the control torque for each joint of the dual-arm robot in the MuJoCo simulation is calculated to control the robot and drive it to perform corresponding actions. Then, the predicted joint position information (14D) from step 1, along with the visual images (480×640×3) from the three simulation cameras, is used as the input for step 2. This input is fed into the model for inference output, and the joint motion information is then applied to the dual-arm robot in the simulation environment to drive it to perform corresponding actions. This process is repeated continuously to drive the dual-arm robot in the simulation environment to complete the task.

[0120] During inference, the CVAE encoder needs to be removed. Then, in the CVAE decoder of the ACT imitation learning algorithm, the joint position information of the current dual-arm robot and the visual information from the camera are input, and the z-hidden variable is set to 0. The inference outputs the joint action sequence, and the final prediction formula is expressed as follows:

[0121]

[0122] in the above formula, denotes the joint action position at time t denotes the visual image of the i-th camera at time t; is a decoder parameter; L represents the number of samples, l = 0, 1, 2, …, L; is the z latent variable corresponding to the i-th target action sequence, and .

[0123] In practical applications, a dual-arm robot is run in a simulation environment, initial joint position information (14 dimensions) and corresponding 3 camera visual images (each image is 480x640x3) are collected, and the visual images are preprocessed, such as cropping, scaling, normalization, etc., in order to facilitate input into the neural network. A CVAE model is constructed, which includes an encoder and a decoder. The encoder receives the current joint position information (14 dimensions) and the conditional information (such as target position or task description), and outputs the latent variable z. The decoder receives the latent variable z, the current joint position information and the conditional information, and predicts the joint action sequence (14 dimensions) at the next time step. A ResNet image encoder (Residual Network, ResNet image encoder is an image feature extractor based on residual network, used to convert input image into high-level semantic features) is constructed to predict the joint action sequence at the next time step. The CVAE model and the ResNet image encoder are trained using the collected data set. During the training process, the model will learn how to predict the joint action sequence at the next time step from the current state (including joint position and visual information). In the simulation environment, the initial joint position information and the visual image are input into the trained model, and the model predicts the joint action sequence at the next time step.

[0124] 207、Based on the target joint action sequence, the control torque of each joint of the robot is calculated through a preset control calculation algorithm, and the robot is driven to perform the corresponding action based on the control torque of each joint.

[0125] The control torque of each joint is calculated using the control calculation method according to the predicted joint action sequence, and the control torque is applied to the robot to drive the robot to perform the corresponding action. The simulation dual-arm robot in the gripper control and joint control simulation environment uses proportional-differential control to control the dual-arm robot joints in the simulation environment to perform actions, and the formula is as follows:

[0126]

[0127] in the above formula, denotes the control torque of joint J at the current time; denotes the proportional gain coefficient of the controller; This represents the derivative gain coefficient of the controller; This indicates the predicted position of joint J in the next instant; This indicates the actual position of joint J at the next moment; This indicates the actual velocity of joint J at the next moment.

[0128] In practical applications, the required joint torques are calculated using the robot's dynamic model and the target joint motion sequence. The target joint angles, velocities, and accelerations are obtained by differentiating the target motion sequence and used as input to calculate the required torque for each joint. Besides directly using inverse dynamics to calculate torque, the preset control calculation algorithm can also combine other control algorithms to obtain smoother or more stable torque outputs. For example, proportional-derivative (PD) control and proportional-integral-derivative (PI) control can be used. It is understood that robot joint torques usually have certain limitations. After calculating the torque, it is necessary to ensure that it does not exceed the joint torque limit. If the calculated torque exceeds the limit, saturation processing is usually required, that is, limiting the torque to the maximum or minimum allowable value. After calculating and adjusting the torque values, each torque value is sent to the robot's controller, which drives the motors according to the received torque commands, thereby generating the required joint torques and applying them to the robot's links and joints, causing them to move according to the target motion sequence, thus driving the robot to perform the corresponding actions.

[0129] This solution predicts the robot's action sequence by learning joint motion information and visual information, which reduces the technical difficulty of the robot in learning operational behaviors and improves the robot's intelligence and the generalization ability of skill simulation.

[0130] The skill simulation method for robots in the embodiments of the present invention has been described above. The skill simulation device for robots in the embodiments of the present invention will be described in detail below from the perspective of modular functional entities. Please refer to [link / reference]. Figure 5 A schematic diagram of a robot skill simulation device provided in this embodiment of the invention, the device comprising:

[0131] The construction module 510 is used to construct a simulation environment based on the operation task, and to parse and decompose the operation task to obtain the gripper state of the robot when completing the operation task. The simulation environment includes a gripper simulation environment based on gripper control and a joint simulation environment based on joint control.

[0132] The first simulation module 520 is used to input the gripper state into the gripper simulation environment and perform simulation based on the gripper state to obtain the first joint movement position of the robot in the gripper simulation environment.

[0133] The second simulation module 530 is configured to input the joint action position into the joint simulation environment, and perform simulation simulation based on the first joint action position to obtain a second joint action position and a visual image of the robot in the joint simulation environment.

[0134] The generating module 540 is configured to train a preset initial prediction formula by taking the second joint action position and the visual image as input data, and optimize the preset initial prediction formula by using a preset loss function to obtain the prediction formula of the robot.

[0135] The driving module 550 is configured to obtain an initial joint action position and an initial visual image of the robot in the joint simulation environment, input the initial joint action position and the initial visual image into the prediction formula to obtain a target joint action sequence of the robot in the joint simulation environment, and drive the robot to perform a corresponding action based on the target joint action sequence.

[0136] The scheme predicts the action sequence of the robot by learning the joint action information and the visual information, reduces the technical difficulty of the robot in learning the operation behavior, improves the intelligent degree and the generalization ability of the skill simulation of the robot, and realizes the learning of the operation behavior of the robot in the simulation environment, so that the robot can have the ability of autonomous adaptation and learning, thereby being capable of autonomously planning an action trajectory, making flexible response and adjustment, improving the fine operation ability of the robot, reducing the skill simulation cost of the robot, and improving the skill simulation efficiency and convenience of the robot.

[0137] Please refer to Figure 6 The robot skill simulation device provided by the embodiment of the present application has another structure diagram, which comprises:

[0138] The constructing module 610 is configured to construct a simulation environment based on an operation task, analyze and split the operation task, and obtain a gripper state of the robot when completing the operation task, wherein the simulation environment comprises a gripper simulation environment based on gripper control and a joint simulation environment based on joint control.

[0139] The first simulation module 620 is configured to input the gripper state into the gripper simulation environment, and perform simulation simulation based on the gripper state to obtain a first joint action position of the robot in the gripper simulation environment.

[0140] The second simulation module 630 is configured to input the joint action position into the joint simulation environment, and perform simulation simulation based on the first joint action position to obtain a second joint action position and a visual image of the robot in the joint simulation environment.

[0141] The generating module 640 is configured to train a preset initial prediction formula by taking the second joint action position and the visual image as input data, and optimize the preset initial prediction formula by using a preset loss function, to obtain a prediction formula of the robot.

[0142] The driving module 650 is configured to acquire an initial joint action position and an initial visual image of the robot in the joint simulation environment, input the initial joint action position and the initial visual image into the prediction formula, obtain a target joint action sequence of the robot in the joint simulation environment, and drive the robot to perform a corresponding action based on the target joint action sequence.

[0143] In this embodiment, the first simulation module 620 includes that the gripper state includes a position, a rotation vector, and a gripping state of the gripper, and the simulation simulation of the gripper simulation environment is:

[0144]

[0145] wherein, represents a gripper simulation environment based on gripper control, represents a first joint action position output in the gripper simulation environment based on the gripper control, represents a position and a rotation vector of the gripper, represents a gripping state of the gripper, 0 represents closed, and 1 represents open, respectively represent a left arm and a right arm of the dual-arm robot, represents a time.

[0146] In this embodiment, the second simulation module 630 includes that the visual image includes a left arm image, a right arm image, and a side image of the robot, and the simulation simulation of the joint simulation environment is:

[0147]

[0148] wherein, represents a joint simulation environment based on joint control, represents a second joint action position output in the joint simulation environment based on the joint control, represents a color image, respectively represent a left arm, a right arm, and a side of the robot, represents a time.

[0149] In this embodiment, the generating module 640 includes:

[0150] The training unit 641 is configured to acquire the second joint action position and the visual image output by the joint simulation environment, and generate input data based on the second joint action position and the visual image, and train a preset initial prediction formula based on the input data as training data.

[0151] In this embodiment, the training unit 641 comprises:

[0152] The input data is in the form of:

[0153]

[0154] wherein, represents the input at time t, represents the joint action position after combination, represents the visual image after combination, represents the dimension of the joint action position, i.e., the dimension of the left arm + the right arm, represents the i-th camera, d ; represents the width x height x channel number of the visual image, represents the time; and the preset initial prediction formula is:

[0155]

[0156] wherein, represents the predicted joint action sequence from time t to time t+k, represents the input data at time t, is a decoder parameter, and L represents the number of samples, = 0, 1, 2, …, L, is the z latent variable corresponding to the i-th target action sequence.

[0157] In this embodiment, the generation module 640 further comprises:

[0158] The optimization unit 642 is configured to calculate an error value between the predicted joint action sequence and the collected joint action sequence based on a preset loss function, and optimize the preset initial prediction formula based on the error value to obtain the prediction formula of the robot, wherein the preset loss function is:

[0159]

[0160] , represents the predicted joint action sequence from time t to time t+k, represents the input joint action sequence, and the prediction formula is:​​​

[0161]

[0162] wherein, represents the joint action position at time t represents the visual image of the i-th camera at time t; is a decoder parameter, and L represents the number of samples, = 0, 1, 2, …, L, is the z latent variable corresponding to the i-th target action sequence, and .

[0163] In this embodiment, the driving module 650 includes:

[0164] a torque calculation unit 651, configured to calculate control torques of joints of the robot based on the target joint action sequence by a preset control calculation algorithm, and drive the robot to perform corresponding actions based on the control torques, wherein the preset control calculation algorithm is:

[0165]

[0166] , represents the control torque of the joint J at the current time, represents the proportional gain coefficient of the controller, represents the derivative gain coefficient of the controller, represents the predicted action position of the joint J at the next time, represents the actual action position of the joint J at the next time, represents the actual speed of the joint J at the next time.

[0167] The scheme predicts the action sequence of the robot by learning joint action information and visual information, reduces the technical difficulty of the robot in learning operation behaviors, improves the intelligent degree and generalization ability of skill simulation of the robot, and realizes learning of operation behaviors of the robot in a simulation environment.

[0168] The above Figures 5-6 The skill simulation device of the robot in the embodiment of the application is described in detail from the perspective of a modular functional entity, and the skill simulation device of the robot in the embodiment of the application is described in detail from the perspective of hardware processing.

[0169] Referring to Figure 7 , the skill simulation device of the robot includes a processor 700 and a memory 701, the memory 701 stores machine executable instructions capable of being executed by the processor 700, and the processor 700 executes the machine executable instructions to implement the above-mentioned skill simulation method of the robot.

[0170] Further, Figure 7 The skill simulation device of the robot shown further includes a bus 702 and a communication interface 703, and the processor 700, the communication interface 703 and the memory 701 are connected through the bus 702.

[0171] The memory 701 can include a high-speed random access memory (RAM), and can also include a non-volatile memory, for example, at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 703 (which can be wired or wireless), and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used. The bus 702 can be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0172] The processor 700 can be an integrated circuit chip with signal processing capability. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 700 or the instruction in the form of software. The processor 700 described above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. Each method, step and logic block disclosed in the embodiment of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiment of the present disclosure can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium in the art. The storage medium is located in the memory 701, and the processor 700 reads the information in the memory 701, and combines the hardware to complete the method steps of the above embodiment.

[0173] The application further provides a computer readable storage medium, which can be a non-volatile computer readable storage medium or a volatile computer readable storage medium, and the computer readable storage medium stores instructions, and the instructions, when executed on a computer, cause the computer to perform the steps of the robot skill simulation method provided by each of the embodiments.

[0174] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the devices or apparatuses, units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.

[0175] The integrated units, if implemented in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the application or the whole or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each of the embodiments of the application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0176] The above-described embodiments are only used to illustrate the technical solutions of the application, rather than limit them. Although the application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not cause the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the application.

Claims

1. A method for simulating robot skills, characterized in that, The robot's skill simulation method includes: A simulation environment is constructed based on the operation task, and the operation task is parsed and decomposed to obtain the gripper state of the robot when completing the operation task; wherein, the simulation environment includes a gripper simulation environment based on gripper control and a joint simulation environment based on joint control; the gripper state includes the gripper position, rotation vector, and gripping state, and the simulation of the gripper simulation environment is as follows: ,in, This represents a gripper simulation environment based on gripper control. This represents the output position of the first joint movement in a gripper simulation environment based on gripper control. This indicates the position and rotation vector of the gripper. This indicates the gripping state of the grippers; 0 means closed, and 1 means open. These represent the left and right arms of the dual-arm robot. Indicates time; The gripper state is input into the gripper simulation environment, and simulation is performed based on the gripper state to obtain the first joint movement position of the robot in the gripper simulation environment; The first joint motion position is input into the joint simulation environment, and simulation is performed based on the first joint motion position to obtain the second joint motion position and visual image of the robot in the joint simulation environment; the visual image includes the robot's left arm image, right arm image, and side view image, and the simulation of the joint simulation environment is as follows: ,in, This represents a joint simulation environment based on joint control. This represents the output position of the second joint motion in a joint simulation environment based on joint control. Represents a color image. These represent the robot's left arm, right arm, and side, respectively. Indicates time; The robot's prediction formula is obtained by training a preset initial prediction formula with the second joint movement position and the visual image as input data, and by optimizing the preset initial prediction formula with a preset loss function. The robot's initial joint motion position and initial visual image in the joint simulation environment are obtained, and the initial joint motion position and the initial visual image are input into the prediction formula to obtain the target joint motion sequence of the robot in the joint simulation environment. The robot is then driven to perform corresponding actions based on the target joint motion sequence.

2. The robot skill simulation method according to claim 1, characterized in that, The step of training a preset initial prediction formula using the second joint movement position and the visual image as input data includes: The second joint motion position and visual image output by the joint simulation environment are obtained, and input data is generated based on the second joint motion position and the visual image. The input data is used as training data to train a preset initial prediction formula.

3. The robot skill simulation method according to claim 2, characterized in that, The input data is in the following format: ,in, This represents the input at time t. Indicates will The combined joint movement position, The combined visual image d The dimension representing the position of joint movement, i.e., the dimension of the left arm plus the right arm. Indicates the first One camera; It represents the width × height × number of channels of a visual image.

4. The robot skill simulation method according to claim 3, characterized in that, The step of optimizing the preset initial prediction formula using a preset loss function to obtain the robot's prediction formula includes: The error between the predicted joint motion sequence and the acquired joint motion sequence is calculated based on a preset loss function. The preset initial prediction formula is then optimized based on this error value to obtain the robot's prediction formula. The preset loss function is: , This represents the mean square error. This represents the predicted sequence of joint movements from time t to t+k. This represents the input sequence of joint movements.

5. The robot skill simulation method according to claim 1, characterized in that, The step of driving the robot to perform corresponding actions based on the target joint motion sequence includes: Based on the target joint motion sequence, a preset control calculation algorithm is used to calculate the control torque of each joint of the robot, and the robot is driven to perform corresponding actions based on each control torque. The preset control calculation algorithm is as follows: , This represents the control torque of joint J at the current moment. This represents the proportional gain coefficient of the controller. This represents the derivative gain coefficient of the controller. This indicates the predicted position of joint J in the next moment. This indicates the actual position of joint J at the next moment. This indicates the actual velocity of joint J at the next moment.

6. A robot skill simulation device, characterized in that, The robot's skill simulation device includes: A construction module is used to build a simulation environment based on an operation task, and to parse and decompose the operation task to obtain the gripper state of the robot when completing the operation task. The simulation environment includes a gripper simulation environment based on gripper control and a joint simulation environment based on joint control. The gripper state includes the gripper's position, rotation vector, and grasping state. The simulation of the gripper simulation environment is as follows: ,in, This represents a gripper simulation environment based on gripper control. This represents the output position of the first joint movement in a gripper simulation environment based on gripper control. This indicates the position and rotation vector of the gripper. This indicates the gripping state of the grippers; 0 means closed, and 1 means open. These represent the left and right arms of the dual-arm robot. Indicates time; The first simulation module is used to input the gripper state into the gripper simulation environment and perform simulation based on the gripper state to obtain the first joint movement position of the robot in the gripper simulation environment. The second simulation module is used to input the joint motion position into the joint simulation environment, and perform simulation based on the first joint motion position to obtain the second joint motion position and visual image of the robot in the joint simulation environment; the visual image includes the robot's left arm image, right arm image, and side image, and the simulation of the joint simulation environment is as follows: ,in, This represents a joint simulation environment based on joint control. This represents the output position of the second joint motion in a joint simulation environment based on joint control. Represents a color image. These represent the robot's left arm, right arm, and side, respectively. Indicates time; The generation module is used to train a preset initial prediction formula with the second joint movement position and the visual image as input data, and to optimize the preset initial prediction formula using a preset loss function to obtain the prediction formula of the robot. The driving module is used to acquire the initial joint motion position and initial visual image of the robot in the joint simulation environment, input the initial joint motion position and the initial visual image into the prediction formula to obtain the target joint motion sequence of the robot in the joint simulation environment, and drive the robot to perform corresponding actions based on the target joint motion sequence.

7. A robot skill simulation device, characterized in that, The robot skill simulation device includes a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the robot skill simulation device to execute the robot skill simulation method as described in any one of claims 1-5.

8. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement the skill simulation method for the robot as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Robot grabbing planning method and device, electronic equipment, storage medium and computer program product

    CN117885101A

  • Action selection neural network training using imitation learning in latent space

    US20200104680A1