Control method, training method, control component, and robot
By determining interactive poses and stitching features based on environmental images, and combining attribute prediction and reinforcement learning modules, the problem of insufficient generalization ability of robots to operate on unseen objects is solved, and real-world operations with a high success rate are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING GALBOT AI CO LTD
- Filing Date
- 2024-10-29
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies lack generalization ability when robots manipulate articulated objects they have never seen before, resulting in a low success rate in the real world and a large gap between simulation and reality.
By determining interactive poses based on environmental images, splicing features are generated. Then, by utilizing attribute prediction and reinforcement learning modules, object attribute data and historical output instructions are considered to reduce visual dependence and improve the understanding of the object's intrinsic characteristics.
It exhibits high generalization ability when dealing with unseen objects, demonstrates a high success rate in the real world, and reduces the gap between simulation and reality.
Smart Images

Figure CN119369394B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot control technology, and in particular to a control method, training method, control components, and robot. Background Technology
[0002] Manipulating articulated objects with robots is a challenging task, requiring a deep understanding of the object and compliant movement to avoid damage to both the object and the robot. Related techniques include predicting a capability graph from point cloud input and then predicting the manipulation action using reinforcement learning strategies; or estimating part poses using multi-view RGB images (Red, Green, Blue, RGB) or partial point clouds and then executing planned actions through a heuristic manipulation module. However, these methods lack high generalization ability when dealing with unseen objects and have not demonstrated high success rates in the real world. Summary of the Invention
[0003] In view of this, embodiments of this application provide at least one control method, a training method, a control component, and a robot.
[0004] The technical solution of this application embodiment is implemented as follows:
[0005] On one hand, embodiments of this application provide a control method applied to a robot. The method includes: determining a first interaction pose for an object to be interacted with based on an image of the robot's environment; generating a spliced feature for the current time step based on the output command of the previous time step and the first interaction pose; inputting the spliced feature of the current time step and the spliced feature of at least one historical time step into a trained attribute prediction module to obtain attribute features; the attribute features are used to characterize attribute data of the object to be interacted with that affect the interaction; and inputting the attribute features and the spliced feature of the current time step into a trained reinforcement learning module to obtain the output command for the current time step.
[0006] On the other hand, embodiments of this application provide a training method for a control model, the control model including an attribute prediction module and a reinforcement learning module; the training method includes: generating simulation stitching features for a second time step based on simulation output commands at a first time step and observation data at a second time step; the first time step being the previous time step of the second time step; inputting the simulation stitching features of the second time step and simulation stitching features from at least one historical time step into the attribute prediction module to obtain predicted simulation attribute features; the simulation attribute features are used to characterize attribute data related to the interaction of a sample object to be interacted with in the simulation environment; inputting the simulation output commands and target attribute features of the second time step into the reinforcement learning module to obtain simulation output commands for the second time step; the target attribute features being the predicted simulation attribute features or simulation attribute features generated based on the attribute data of the sample object; obtaining motion rewards after controlling robot movement based on the simulation output commands of the second time step; training the attribute prediction module based on the predicted simulation attribute features and the generated simulation attribute features, and training the reinforcement learning module through reinforcement learning based on the motion rewards.
[0007] In another aspect, embodiments of this application provide a control component disposed in a robot. The control component includes: a determining unit, configured to determine a first interactive pose for an object to be interacted with based on an image of the robot's environment; a generating unit, configured to generate a splicing feature for the current time step based on the output command of the previous time step and the first interactive pose; a first input unit, configured to input the splicing feature of the current time step and the splicing feature of at least one historical time step into a trained attribute prediction module to obtain attribute features; the attribute features are used to characterize attribute data of the object to be interacted that are related to the interaction; and a second input unit, configured to input the attribute features and the splicing feature of the current time step into a trained reinforcement learning module to obtain the output command of the current time step.
[0008] In another aspect, embodiments of this application provide a robot, which includes the control components described above.
[0009] In this embodiment, the reinforcement learning module can consider the attribute data, interaction pose, and historical output instructions of the object to be interacted with when outputting instructions, thereby reducing the reliance on vision, better understanding the intrinsic characteristics of the object, having a high generalization ability when dealing with unseen objects, exhibiting a high success rate in the real world, and reducing the gap between simulation and reality.
[0010] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this application. Attached Figure Description
[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.
[0012] Figure 1 A schematic diagram of the implementation process of a control method provided in this application embodiment. Figure 1 ;
[0013] Figure 2 A schematic diagram of the implementation process of a control method provided in this application embodiment. Figure 2 ;
[0014] Figure 3 A schematic diagram of the implementation process of a control method provided in this application embodiment. Figure 3 ;
[0015] Figure 4 A schematic diagram of the implementation process of a control method provided in this application embodiment. Figure 4 ;
[0016] Figure 5 A schematic diagram of the implementation process of a control method provided in this application embodiment. Figure 5 ;
[0017] Figure 6 A schematic diagram of the implementation process of a control model training method provided in this application embodiment. Figure 1 ;
[0018] Figure 7 A schematic diagram of the implementation process of a control model training method provided in this application embodiment. Figure 2 ;
[0019] Figure 8 Illustration of reinforcement learning strategy application provided in embodiments of this application Figure 1 ;
[0020] Figure 9 Illustration of reinforcement learning strategy application provided in embodiments of this application Figure 2 ;
[0021] Figure 10 This is a schematic diagram of the experimental setup for an embodiment of this application;
[0022] Figure 11 This is a schematic diagram of the experimental results of an embodiment of this application;
[0023] Figure 12 A schematic diagram of the composition structure of a control component provided in an embodiment of this application;
[0024] Figure 13 This is a schematic diagram of the composition structure of a training device provided in an embodiment of this application;
[0025] Figure 14 This is a schematic diagram of the hardware entity of a computer device provided in an embodiment of this application. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0027] In the following description, references to "some embodiments" refer to a subset of all possible embodiments. It is understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other without conflict. The terms "first / second / third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used herein is for descriptive purposes only and is not intended to limit the scope of this application.
[0029] This application provides a control method and a control model training method, which can be executed by a processor of a computer device. The computer device refers to a device with data processing capabilities, such as a server, laptop, tablet, desktop computer, smart TV, set-top box, or mobile device (e.g., mobile phone, portable video player, personal digital assistant, dedicated messaging device, portable gaming device).
[0030] In the field of robot control, for manipulating articulated objects, one approach is to predict the availability graph using point cloud input and then predict the maneuver through a reinforcement learning policy. This method uses point cloud features as policy input, leading to a significant gap between simulation and reality. Furthermore, the availability graph can be occluded by the robot arm during execution. Another approach uses multi-view RGB images to estimate part pose and then executes the planned action through a heuristic manipulation module. Multi-view RGB requires more time for part pose estimation, and the predicted waypoints need adjustment at each local point, resulting in an uneven trajectory. A third approach uses partial point cloud data to estimate part pose and then executes the planned action through a heuristic manipulation module. However, heuristic manipulation lacks a feedback mechanism to adjust the planned action and ignores the object's intrinsic characteristics, potentially leading to unsafe behavior. None of these methods demonstrate high generalization ability when handling unseen objects, nor do they show high success rates in the real world.
[0031] This application provides a control method for a robot, comprising: determining a first interaction pose for an object to be interacted with based on an image of the robot's environment; generating a spliced feature for the current time step based on the output command of the previous time step and the first interaction pose; inputting the spliced feature of the current time step and the spliced feature of at least one historical time step into a trained attribute prediction module to obtain attribute features; the attribute features are used to characterize the attribute data of the object to be interacted that are related to the interaction; and inputting the attribute features and the spliced feature of the current time step into a trained reinforcement learning module to obtain the output command for the current time step. In this way, the reinforcement learning module can consider the attribute data of the object to be interacted with, the interaction pose, and historical output commands when outputting commands, reducing reliance on vision, better understanding the intrinsic characteristics of objects, exhibiting high generalization ability when handling unseen objects, demonstrating a high success rate in the real world, and reducing the gap between simulation and reality.
[0032] In some embodiments, a hinged object refers to an object composed of multiple movable parts connected by hinges, slides, or other joints. Typical hinged objects include doors, drawers, and robotic arms; the connection between the joints allows the object to perform certain movements, such as rotation and sliding. Reinforcement learning is a machine learning method that optimizes action strategies through interaction with the environment, based on a trial-and-error process, to maximize cumulative rewards. A reinforcement learning strategy refers to the action selection rules that a reinforcement learning algorithm continuously adjusts during training to make optimal decisions in different situations. A point cloud is a three-dimensional data representation composed of a large number of points, each with its own three-dimensional coordinates (X, Y, Z). Point clouds are typically acquired using devices such as LiDAR or RGB-D cameras and are widely used in tasks such as 3D modeling, object detection, and scene understanding.
[0033] Figure 1 A schematic diagram of the implementation process of a control method provided in this application embodiment. Figure 1 Control methods are applied to robots, such as Figure 1 As shown, the method includes the following steps S101 to S104:
[0034] Step S101: Determine the first interaction pose for the object to be interacted with based on the environmental image of the robot.
[0035] The environmental image refers to an image of the robot's environment, which must contain at least the object to be interacted with. The environmental image can be obtained by capturing images of the robot's environment using an image acquisition module (e.g., a camera) carried by the robot.
[0036] In this context, the object to be interacted with is the object that the robot manipulates. For example, if the robot needs to perform the task of opening a door, it does so by grasping the doorknob and opening the door; the object to be interacted with is the doorknob.
[0037] The first interactive pose refers to the position and orientation that the robot's end effector (such as a gripper or suction cup) is expected to reach when performing an interactive task. For example, when the robot performs a door-opening task, the first interactive pose is the position and orientation that the robot's gripper is expected to reach.
[0038] The robot's interaction with the object to be interacted with can include grasping, lifting, translating, and pushing / pulling operations. This application does not limit the specific type of interaction.
[0039] In some embodiments, the environment image includes a depth image and an RGB image. Object recognition is performed on the RGB image using machine learning to obtain a mask image of the object to be interacted with. The depth image is transformed to obtain environment point cloud data. Multiple second interaction poses are determined using the environment point cloud data. A first interaction pose for the object to be interacted with is determined using the mask image among the multiple second interaction poses.
[0040] In some embodiments, a model for generating a first interaction pose is pre-trained. This model is capable of recognizing objects to be interacted with in environmental images and generating a first interaction pose for those objects. The training samples can be multiple environmental images containing objects to be interacted with, and the labels can be the objects to be interacted with in the environmental images and their corresponding first interaction poses. The trained model can recognize objects in the environmental images...
[0041] Step S102: Based on the output command of the previous time step and the first interactive pose, generate the splicing features of the current time step.
[0042] When the robot performs an operation task, the entire motion execution process is divided into multiple time steps. At each time step, the corresponding action is executed based on the output instructions of the reinforcement learning module. The output instructions contain information related to the robot's action.
[0043] In some embodiments, the output commands include target displacement, target orientation, gripper action, and impedance control parameters. Target displacement is the target position the robot needs to move to. Target orientation refers to the robot's complete orientation or posture information in three-dimensional space relative to a reference coordinate system (such as the world coordinate system or camera coordinate system). Gripper action refers to whether the robot performs the action of grasping an object. Impedance control parameters are used to control the impedance of the output action.
[0044] The stitched features of the current time step include the output instructions from the previous time step and the observation data from the current time step. Observation data refers to the environmental state information of the robot's environment. The stitched features are used to characterize the robot's historical actions and the current environmental state.
[0045] In some embodiments, the observation data includes a first interaction pose, robot joint configuration, relative distance between the robot and the object, end effector pose, and graspability signal. Robot joint configuration refers to the selection and configuration of the robot's joint type, number, arrangement, and actuation method during the design and manufacturing process. Relative distance between the robot and the object refers to the distance between the robot's end effector and the object to be interacted with. End effector pose refers to the position and orientation of the end effector in a specified coordinate system. Graspability signal indicates whether the robot's end effector grasps the object. The graspability signal is based on distance and contact perception conditions, rather than direct commands controlling the opening and closing of the gripper.
[0046] In some embodiments, at each time step, the robot acquires observation data for the current time step, including a first interactive pose, and combines the observation data of the current time step with the output instructions of the previous time step to obtain the splicing features of the current time step.
[0047] Step S103: Input the concatenated features of the current time step and the concatenated features of at least one historical time step into the trained attribute prediction module to obtain attribute features; the attribute features are used to characterize the attribute data of the object to be interacted with that are related to the interaction.
[0048] The attribute prediction module is used to infer the attribute characteristics of the object to be interacted with from historical observation data and output instructions. Attribute characteristics are used to characterize the attribute data of the object to be interacted with that are related to the interaction.
[0049] In some embodiments, to better understand the environment, it is necessary to obtain attribute data of the object to be interacted with that is relevant to the interaction. However, this data is difficult to obtain directly in the real world. Therefore, by training an attribute prediction module and inputting the concatenated features of the current time step and the concatenated features of at least one historical time step into the trained attribute prediction module, attribute features can be obtained. These attribute features can approximately reflect the attribute data of the object to be interacted with that is relevant to the interaction.
[0050] In some embodiments, at least one historical time step is a historical time step adjacent to the current time step. For example, when the current time step is t, the three historical time steps include time steps t-1, t-2, and t-3.
[0051] In some embodiments, the attribute data includes at least one of the following: pivot center, pivot radius, object stiffness, object mass, object joint position, and handle grip signal. The pivot center refers to the center point or axis of rotation of an object or system. The pivot radius is the distance from the pivot center to a point on the rotating object. Object stiffness refers to the object's ability to resist deformation when subjected to external forces. Object mass is a physical quantity describing the magnitude of an object's inertia; it is a measure of the amount of matter contained in the object. Object joint position is the current orientation of an object's joints. The handle grip signal indicates whether the handle is currently gripped.
[0052] In some embodiments, during the simulation process, at any time step, the simulation software can directly acquire the attribute data of the object to be interacted with, and simulation attribute features can be generated based on the attribute data. The attribute prediction module can obtain the predicted simulation attribute features from the concatenated features of the current time step and the concatenated features of at least one historical time step. The parameters of the attribute prediction module are adjusted based on the difference between the generated simulation attribute features and the predicted simulation attribute features, and finally, the trained attribute prediction module is obtained. In the application process, the concatenated features of the current time step and the concatenated features of at least one historical time step are input into the trained attribute prediction module to obtain the attribute features.
[0053] Step S104: Input the attribute features and the concatenated features of the current time step into the trained reinforcement learning module to obtain the output instructions of the current time step.
[0054] In some embodiments, for tasks where the robot manipulates objects, a reinforcement learning module is trained that can output the actions the robot should perform at each time step based on the current environmental state and historical actions. The robot then performs the actions based on the action instructions output by the reinforcement learning module to complete the task.
[0055] In this module, at each time step, the reinforcement learning module determines the output command for the current time step based on the attribute features output by the attribute prediction module and the concatenated features of the current time step. The robot can then execute actions based on these output commands.
[0056] In some embodiments, during the simulation training of the reinforcement learning module, at each time step, a generated simulation attribute feature and a predicted simulation attribute feature are obtained. The generated or predicted simulation attribute feature at the current time step, along with the concatenated feature, is input into the reinforcement learning module to obtain an output command. The robot receives a motion reward after executing the output command. Based on the motion reward, the parameters of the reinforcement learning module are adjusted to obtain the trained reinforcement learning module. In application, the attribute feature and the concatenated feature at the current time step are input into the trained reinforcement learning module to obtain the output command for the current time step.
[0057] For example, when a robot performs a door-opening task, the object to be interacted with is the door handle. Based on the environmental image, the robot determines the first interaction pose g for the door handle. t This allows us to obtain the observation index o at the current time step t. t Observation indicator o t Including the first interactive pose g t The output instruction a at time step t-1. t-1 Compared with the observed indicators o t By splicing, the spliced feature p is obtained. t =(o t ⊕a t-1 ), splicing feature p t splicing features p at time step t-1 t-1 Input the trained attribute prediction module to obtain the predicted attribute features. Predicted attribute features The concatenation feature p of the current time step t Input the trained reinforcement learning module to obtain the output instruction 'a' at the current time step. t .
[0058] In this embodiment, a first interactive pose for the object to be interacted with is determined based on an image of the robot's environment. Based on the output command of the previous time step and the first interactive pose, a spliced feature for the current time step is generated. The spliced feature of the current time step and the spliced feature of at least one historical time step are input into a trained attribute prediction module to obtain attribute features. The attribute features are used to characterize the attribute data of the object to be interacted with that influences the interaction. The attribute features and the spliced feature of the current time step are input into a trained reinforcement learning module to obtain the output command for the current time step. In this way, the reinforcement learning module can consider the attribute data of the object to be interacted with that influences the interaction, the interactive pose, and historical output commands when outputting commands, reducing reliance on vision, better understanding the intrinsic characteristics of objects, exhibiting high generalization ability when handling unseen objects, demonstrating a high success rate in the real world, and reducing the gap between simulation and reality.
[0059] In some embodiments, the trained attribute prediction module and the trained reinforcement learning module are obtained after training using attribute data related to the interaction of sample objects to be interacted with in a simulation environment.
[0060] In this process, at any time step, the attribute prediction module adjusts its parameters based on the difference between the simulated attribute features generated at that time step and the simulated attribute features predicted by the attribute prediction module; the generated simulated attribute features are obtained based on the attribute data of the sample object.
[0061] The reinforcement learning module adjusts parameters based on motion rewards, which are obtained after the robot responds to the simulation output command. The simulation output command is output by the reinforcement learning module at any time step based on the generated simulation attribute features or the predicted simulation attribute features.
[0062] In some embodiments, the attribute prediction module needs to be trained in a simulation environment. The simulation environment can directly acquire attribute data related to the interaction of the sample object at any time step. Based on the acquired attribute data of the sample object, simulation attribute features can be generated. The concatenated features of the current time step and at least one concatenated feature from a historical time step are input into the attribute prediction module to obtain the predicted simulation attribute features. The parameters of the attribute prediction module can be adjusted based on the difference between the generated simulation attribute features and the simulation attribute features predicted by the attribute prediction module.
[0063] For example, the attribute data of the sample object The input to the encoding module generates the simulation attribute features z. The encoding module can be a multilayer perceptron (MLP). The concatenated features from the current time step and the concatenated features from at least one historical time step are input to the attribute prediction module to obtain the predicted simulation attribute features. based on Adjust the parameters of the attribute prediction module. sg[.] represents the stopping gradient operator, and λ is a preset parameter.
[0064] In some embodiments, the reinforcement learning module needs to be trained in a simulation environment. At any time step, the simulation environment can directly acquire attribute data related to the interaction with the sample object to be interacted with. Based on the acquired attribute data, simulation attribute features can be generated. The simulation attribute features generated at the current time step and the concatenated features are input into the reinforcement learning module, or the simulation attribute features predicted at the current time step and the concatenated features are input into the reinforcement learning module. The reinforcement learning module outputs simulation output commands. After the robot moves in response to the simulation output commands, it receives a motion reward, and the reinforcement learning module adjusts its parameters based on the motion reward.
[0065] In this embodiment, the attribute prediction module and reinforcement learning module are trained using attribute data of the sample object to be interacted with in the simulation environment that are related to the interaction. Thus, the trained attribute prediction module can accurately predict the attribute characteristics of the object to be interacted with in the current time step, and the reinforcement learning module can determine the output command for the current time step based on the intrinsic attributes of the object to be interacted with, thereby better understanding the object's inherent characteristics and demonstrating a high success rate in the real world.
[0066] Figure 2 This is a schematic diagram of the implementation flow of a control method provided in an embodiment of this application. Figure 2 This method can be executed by the processor of a computer device. Based on Figure 1 The method further includes steps S201 and S202.
[0067] Step S201: Store the splicing features of the current time step in the history cache.
[0068] In some embodiments, after acquiring the splicing features at any time step, the splicing features of that time step can be stored in a history cache. The history cache stores the splicing features of all historical time steps.
[0069] In some embodiments, the history cache stores a preset number of recent concatenation features, such as the concatenation features of the last 10 historical time steps. After obtaining the concatenation features of the current time step, the concatenation features in the history cache are updated based on the concatenation features of the current time step.
[0070] Step S202: Obtain the splicing features of the at least one historical time step from the historical cache.
[0071] In some embodiments, based on the requirements of the attribute prediction module, the concatenated features of at least one historical time step are obtained from the historical cache. By reading the historical cache, the concatenated features of at least one historical time step can be obtained.
[0072] In this embodiment, the splicing features of the current time step are stored in a historical cache; at least one splicing feature of a historical time step is retrieved from the historical cache. This allows for efficient storage and access to the splicing features of historical time steps through the historical cache, improving the efficiency of splicing feature access.
[0073] In some embodiments, the method further includes scaling the output instruction to obtain a target control instruction.
[0074] The output instructions obtained by the reinforcement learning module are not standard action instructions. The robot cannot directly execute actions based on the output instructions. Therefore, the output instructions need to be scaled to obtain target control instructions. The robot executes actions based on the target control instructions. Adjusting or scaling the output instructions to ensure they can be executed safely and effectively by the robot may involve adjusting parameters such as the speed, force, and range of the action. The scaled actions become "robot instructions" that can be understood and executed by the robot. These instructions are then sent to the robot's control system to drive the robot to perform the corresponding actions.
[0075] In some embodiments, the output command is scaled based on a preset formula to obtain the target control command. For example, the target control command c t The calculation formula is formula (1):
[0076] c t =clip(a t ,-1,1)*40+100 (1);
[0077] Among them, a t The `clip` function is a range limiting function that restricts the output command to the range [-1, 1].
[0078] In some embodiments, the output command is scaled by a pre-trained scaling module, the output command is input into the scaling module, and the target control command is output.
[0079] In this embodiment, the output command is scaled to obtain the target control command. Thus, by scaling the output command, a target control command that the robot can accurately recognize can be obtained.
[0080] In some embodiments, the method further includes: controlling the interactive device based on the target control command to perform an interactive action at the current time step.
[0081] In this system, the robot interacts with the object to be interacted with through an interactive device, which is a component of the robot. For example, the interactive device is the robot's gripper.
[0082] The target control instructions include data related to the interactive device's execution of interactive actions. Based on the target control instructions, the interactive device can be controlled to execute interactive actions at the current time step. The target control instructions are action instructions that the robot can understand and effectively execute. For example, the target control instructions include target displacement, target posture, gripper action, and impedance control parameters.
[0083] In this embodiment, the interactive device is controlled based on the target control command to execute the interactive action at the current time step. In this way, the interactive device can accurately execute the interactive action according to the target control command.
[0084] In some embodiments, the method further includes: generating a stiffness matrix of the robot based on the output command; constructing a dynamic model for impedance control based on the robot's mass inertia matrix, damping matrix, and stiffness matrix; the dynamic model is used to control the balance between the robot's motion and contact forces.
[0085] Impedance control, a force-based control method, treats the robot as a programmable spring-mass-damped system. By adjusting the robot's impedance (its resistance to external forces), it controls its interaction with the environment. Impedance control is a strategy that regulates the dynamic relationship between the robot and its environment (impedance parameters, including stiffness, damping, and inertia) to control the robot's interaction. This gives the robot better adaptability and flexibility when interacting with its environment.
[0086] Stiffness represents a robot's ability to resist external deformation; by adjusting stiffness, the degree of deformation of the robot when subjected to external forces can be controlled. Damping represents the friction or energy dissipation capacity between the robot and its environment; by adjusting damping, the velocity change of the robot when subjected to external forces can be controlled. Inertia represents the robot's response speed to external forces; by adjusting inertia, the acceleration change of the robot when subjected to external forces can be controlled.
[0087] The dynamic model of impedance control is a mathematical model constructed based on the stiffness matrix, mass inertia matrix, and damping matrix. This model describes the dynamic response characteristics of the robot when subjected to external forces, that is, the balance between the robot's motion state (position, velocity, acceleration) and the contact force. By adjusting the parameters in these matrices, the robot's behavior in the contact environment can be controlled, such as maintaining a stable contact force and avoiding excessive deformation.
[0088] In some embodiments, the output instruction of the reinforcement learning module includes stiffness. The stiffness is scaled to obtain a target stiffness, and the target stiffness is expanded to obtain a stiffness matrix. A dynamic model is constructed based on the adjusted stiffness matrix, the unadjusted mass-inertia matrix, and the damping matrix. For example, the stiffness output by the reinforcement learning module is... The stiffness is converted into the target stiffness using formula (2):
[0089]
[0090] The `clip` function is a range limiting function that limits the scope of a function's output. The target stiffness is limited to [-1, 1]. The stiffness matrix K is obtained by expansion.
[0091] For example, a dynamic model is constructed based on the mass inertia matrix M, the damping matrix D, and the stiffness matrix K, as shown in equation (3):
[0092]
[0093] Among them, F ext x is the external force generated by the interaction between the robot and its environment. d For the expected trajectory, For prefetching speed, For the expected acceleration, x c For impedance-controlled output trajectory This represents impedance-controlled acceleration. This indicates the speed controlled by impedance. Based on a dynamic model, the balance between the robot's motion (position, velocity, acceleration) and contact forces can be controlled.
[0094] The expected trajectory serves as a target or reference point to guide the robot's movement, while the output trajectory is the result of the robot adjusting according to the expected trajectory and real-time environmental information during actual operation.
[0095] Figure 3 This is a schematic diagram of the implementation flow of a control method provided in an embodiment of this application. Figure 3 This method can be executed by the processor of a computer device. Based on Figure 1 The environmental images include depth images and color images. Figure 1 S101 in the middle can be updated to S301 to S304, which will combine Figure 3 The steps shown are explained.
[0096] Step S301: Perform object recognition on the color image to obtain the mask image of the object to be interacted with.
[0097] In this context, depth images primarily represent the distance information from each point in the scene to the image acquisition device. The grayscale value or a specific numerical value of each pixel represents the distance from that point to the image acquisition device. Color images primarily represent the color information of objects in the scene. The RGB value of each pixel represents the color component of that point, thus forming a rich and colorful image. For example, a color image is an RGB image.
[0098] A mask image is a special type of image used to specify the regions of the original image to be manipulated. Mask images are typically binary images (each pixel has only two possible values, such as 0 and 255), but can also be grayscale or multi-channel images. In a mask image, regions of interest are set to white (or higher grayscale values), while regions of no interest remain black (or lower grayscale values).
[0099] In some embodiments, an image segmentation model is pre-trained. The color image to be processed is input into the model, which processes the input image and outputs a segmentation result. Based on the segmentation result, a binary mask image is generated. In this mask image, pixels belonging to the object to be interacted with are set to white (or a specific value), while pixels not belonging to the object to be interacted with are set to black (or another specific value).
[0100] Step S302: Convert the depth image to obtain environmental point cloud data.
[0101] Environmental point cloud data is a form of three-dimensional spatial data representation, mainly composed of a large number of three-dimensional spatial points. These points represent sampling points on the surface of an object or specific locations in space. Each point typically contains three-dimensional coordinate information (X, Y, Z), and sometimes also includes other attribute information such as color, intensity, and reflectivity.
[0102] This process involves acquiring depth images and camera parameters, including intrinsic and extrinsic parameters. Intrinsic parameters describe the camera's internal properties, such as focal length, principal point coordinates, image resolution, and distortion parameters. Extrinsic parameters describe the camera's position and orientation in the world coordinate system.
[0103] In some embodiments, for each pixel in the depth image, it is converted to normalized planar coordinates using camera intrinsics; the coordinates of the point cloud in three-dimensional space are calculated using the depth value, the corrected normalized planar coordinates, and the depth scaling factor.
[0104] Step S303: Using the environmental point cloud data, determine multiple second interactive poses.
[0105] Among them, environmental point cloud data reflects the shape of object surfaces and can represent multiple objects in the robot's environment. Using a pre-trained interaction pose determination module, the second interaction pose of each object in the environment can be determined. Specifically, the environmental point cloud data is input into the interaction pose determination module, which outputs multiple second interaction poses.
[0106] Step S304: Determine the first interaction pose from the plurality of second interaction poses using the mask image.
[0107] The mask image can represent the spatial location of the object to be interacted with in the environment. The mask image can be used to determine the first interaction pose corresponding to the object to be interacted with from multiple second interaction poses.
[0108] In this embodiment, object recognition is performed on the color image to obtain a mask image of the object to be interacted with; the depth image is converted to obtain environmental point cloud data; multiple second interaction poses are determined using the environmental point cloud data; and the first interaction pose is determined from the multiple second interaction poses using the mask image. In this way, the first interaction pose of the object to be interacted with can be accurately determined using both the depth image and the color image.
[0109] Figure 4 This is a schematic diagram of the implementation flow of a control method provided in an embodiment of this application. Figure 4 This method can be executed by the processor of a computer device. Based on Figure 1 The training method for the attribute prediction module includes steps S401 to S404.
[0110] Step S401: Based on the simulation output command of the first time step and the observation data of the second time step, generate the simulation splicing feature of the second time step; the first time step is the previous time step of the second time step.
[0111] In this system, simulation stitching features represent the robot's historical actions and current environmental state. Historical actions are represented by simulation output commands at the first time step, while the current environmental state is represented by observation data at the second time step. Therefore, the simulation output commands at the first time step and the observation data at the second time step are combined into a data pair, which serves as the simulation stitching feature for the second time step. The first time step is the previous time step of the second time step. In this way, the simulation stitching feature for the second time step can represent the latest environmental state and the most recent historical actions.
[0112] In some embodiments, the simulation output commands include target displacement, target attitude, gripper action, and impedance control parameters.
[0113] In some embodiments, the observation data includes a first interactive pose, robot joint configuration, relative distance between the robot and the object, end effector pose, and graspability signal.
[0114] Step S402: Input the simulation splicing features of the second time step and the simulation splicing features of at least one historical time step into the attribute prediction module to obtain the predicted simulation attribute features; the simulation attribute features are used to characterize the attribute data of the sample objects to be interacted with in the simulation environment that are related to the interaction.
[0115] The attribute prediction module is used to predict simulation attribute features from simulation stitched features at least two time steps. Simulation attribute features characterize attribute data related to the interaction of sample objects in the simulation environment. Therefore, by inputting the simulation stitched features from the second time step and the simulation stitched features from at least one historical time step into the attribute prediction module, the predicted simulation attribute features can be obtained. At least one historical time step includes the first time step.
[0116] In some embodiments, the attribute data includes at least one of the following: pivot center, pivot radius, object stiffness, object mass, object joint position, and handle gripping signal.
[0117] Step S403: Input the simulation splicing features and target attribute features of the second time step into the reinforcement learning module to obtain the simulation output command of the second time step; the target attribute features are the predicted simulation attribute features or simulation attribute features generated based on the attribute data of the sample object.
[0118] The reinforcement learning module comprehensively considers the current environment state, historical actions, and attribute data of the object to be interacted with to determine the simulation output command. The current environment state and historical actions correspond to simulation splicing features, and the attribute data of the object to be interacted with corresponds to target attribute features. Therefore, the simulation splicing features and target attribute features of the second time step are input into the reinforcement learning module to obtain the simulation output command of the second time step.
[0119] The target attribute features are either predicted simulation attribute features or simulation attribute features generated based on the attribute data of the sample objects. Predicted simulation attribute features are the attribute features output by the attribute prediction module, while generated simulation attribute features are attribute features obtained based on the attribute data of the sample objects. By considering two different methods to determine the target attribute features, and inputting the simulation splicing features from the second time step and the target attribute features into the reinforcement learning module, the adaptability of the reinforcement learning module to different attribute features can be improved, thus enhancing its generalization ability.
[0120] In some embodiments, inputting the attribute data of the sample object into the encoding module yields generated simulation attribute features. The encoding module can be a multilayer perceptron, trained based on the difference between the predicted simulation attribute features and the generated simulation attribute features.
[0121] Step S404: Train the attribute prediction module based on the predicted simulation attribute features and the generated simulation attribute features.
[0122] Among them, the attribute prediction module is based on the predicted simulation attribute features. The difference between the generated simulation attribute features z and the training loss is used. For example, a supervised regularization loss is formulated. based on The parameters of the attribute prediction module are adjusted, where sg[.] represents the stopping gradient operator and λ is a preset parameter. The goal of training is to minimize the loss.
[0123] In this embodiment, simulation stitching features for the second time step are generated based on the simulation output command at the first time step and the observation data at the second time step. These features, along with simulation stitching features from at least one historical time step, are input into an attribute prediction module to obtain predicted simulation attribute features. The simulation stitching features and target attribute features are then input into a reinforcement learning module to obtain the simulation output command for the second time step. The attribute prediction module is trained based on the predicted and generated simulation attribute features. This allows for the training of an accurate attribute prediction module based on the simulation output command and observation data.
[0124] Figure 5 This is a schematic diagram of the implementation flow of a control method provided in an embodiment of this application. Figure 5 This method can be executed by the processor of a computer device. Based on Figure 4 The training method for the reinforcement learning module includes steps S501 to S502.
[0125] Step S501: After controlling the robot's movement based on the simulation output command of the second time step, obtain the motion reward.
[0126] In this process, after the reinforcement learning module outputs the simulation output command for the second time step, the robot can execute the corresponding action according to the simulation output command. After the robot moves, it will receive a motion reward. The motion reward is used to evaluate the quality of the robot's action.
[0127] In some embodiments, motion rewards include task-aware rewards and motion-aware rewards. Task-aware rewards are used to encourage the robot to perform tasks in the correct order, while motion-aware rewards are used to encourage the robot to generate smooth motion while performing tasks.
[0128] Step S502: Train the reinforcement learning module using reinforcement learning based on the exercise reward.
[0129] In some embodiments, each time step corresponds to a simulation output command. After the robot executes the corresponding action, it receives a motion reward, and the environmental state changes. The environmental state, simulation output command, corresponding motion reward, and new environmental state after the action execution at each time step are saved. Based on this data, the policy of the reinforcement learning module is updated at each time step using a reinforcement learning algorithm. For example, the policy of the reinforcement learning module is updated using the Proximal Policy Optimization (PPO) algorithm. Thus, at each time step, the policy is updated using the reinforcement learning algorithm based on the motion reward. The training process of the reinforcement learning module stops when the iteration stopping condition is met or the policy reaches the optimization objective.
[0130] In some embodiments, a round comprises multiple time steps. The environment state, simulation output instructions, corresponding motion rewards, and the new environment state after the action are executed at each time step are saved, and the reinforcement learning policy is updated once in each round. When the accumulated number of time steps reaches the threshold corresponding to the round, the policy is updated based on motion rewards using a reinforcement learning algorithm.
[0131] In this embodiment, after controlling the robot's movement based on the simulation output command at the second time step, a motion reward is obtained; a reinforcement learning module is then trained using this motion reward. This allows the reinforcement learning module to be trained based on the motion reward corresponding to the simulation output command, resulting in an accurate reinforcement learning module.
[0132] Figure 6 This is a schematic diagram of the implementation flow of a control model training method provided in an embodiment of this application. Figure 1 ,like Figure 6 As shown, the control model includes an attribute prediction module and a reinforcement learning module; the training method includes steps S601 to S605.
[0133] Step S601: Based on the simulation output command of the first time step and the observation data of the second time step, generate the simulation splicing feature of the second time step; the first time step is the previous time step of the second time step.
[0134] The attribute prediction module is used to infer the attribute characteristics of the object to be interacted with from historical observation data and output instructions. Attribute characteristics are used to characterize the attribute data of the object to be interacted with that are related to the interaction.
[0135] The reinforcement learning module can output the action that the robot should perform at each time step based on the current environmental state and historical actions. The robot then performs the action based on the action instructions output by the reinforcement learning module to complete the task.
[0136] In this system, simulation stitching features represent the robot's historical actions and current environmental state. Historical actions are represented by simulation output commands at the first time step, while the current environmental state is represented by observation data at the second time step. Therefore, the simulation output commands at the first time step and the observation data at the second time step are combined into a data pair, which serves as the simulation stitching feature for the second time step. The first time step is the previous time step of the second time step. In this way, the simulation stitching feature for the second time step can represent the latest environmental state and the most recent historical actions.
[0137] In some embodiments, the simulation output commands include target displacement, target attitude, gripper action, and impedance control parameters.
[0138] In some embodiments, the observation data includes a first interactive pose, robot joint configuration, relative distance between the robot and the object, end effector pose, and graspability signal.
[0139] Step S602: Input the simulation splicing features of the second time step and the simulation splicing features of at least one historical time step into the attribute prediction module to obtain the predicted simulation attribute features; the simulation attribute features are used to characterize the attribute data of the sample objects to be interacted with in the simulation environment that are related to the interaction.
[0140] The attribute prediction module is used to predict simulation attribute features from simulation stitched features at least two time steps. Simulation attribute features characterize attribute data related to the interaction of sample objects in the simulation environment. Therefore, by inputting the simulation stitched features from the second time step and the simulation stitched features from at least one historical time step into the attribute prediction module, the predicted simulation attribute features can be obtained. At least one historical time step includes the first time step.
[0141] In some embodiments, the attribute data includes at least one of the following: pivot center, pivot radius, object stiffness, object mass, object joint position, and handle gripping signal.
[0142] Step S603: Input the simulation output command and target attribute features of the second time step into the reinforcement learning module to obtain the simulation output command of the second time step; the target attribute features are the predicted simulation attribute features or simulation attribute features generated based on the attribute data of the sample object.
[0143] The reinforcement learning module comprehensively considers the current environment state, historical actions, and attribute data of the object to be interacted with to determine the simulation output command. The current environment state and historical actions correspond to simulation splicing features, and the attribute data of the object to be interacted with corresponds to target attribute features. Therefore, the simulation splicing features and target attribute features of the second time step are input into the reinforcement learning module to obtain the simulation output command of the second time step.
[0144] The target attribute features are either predicted simulation attribute features or simulation attribute features generated based on the attribute data of the sample objects. Predicted simulation attribute features are the attribute features output by the attribute prediction module, while generated simulation attribute features are attribute features obtained based on the attribute data of the sample objects. By considering two different methods to determine the target attribute features, and inputting the simulation splicing features from the second time step and the target attribute features into the reinforcement learning module, the adaptability of the reinforcement learning module to different attribute features can be improved, thus enhancing its generalization ability.
[0145] Step S604: After controlling the robot's movement based on the simulation output command of the second time step, obtain the motion reward.
[0146] In this process, after the reinforcement learning module outputs the simulation output command for the second time step, the robot can execute the corresponding action according to the simulation output command. After the robot moves, it will receive a motion reward. The motion reward is used to evaluate the quality of the robot's action.
[0147] In some embodiments, motion rewards include task-aware rewards and motion-aware rewards. Task-aware rewards are used to encourage the robot to perform tasks in the correct order, while motion-aware rewards are used to encourage the robot to generate smooth motion while performing tasks.
[0148] Step S605: Train the attribute prediction module based on the predicted simulation attribute features and the generated simulation attribute features, and train the reinforcement learning module based on the motion reward using reinforcement learning.
[0149] Among them, the attribute prediction module is based on the predicted simulation attribute features. The difference between the generated simulation attribute features z and the training loss is used. For example, a supervised regularization loss is formulated. based on The parameters of the attribute prediction module are adjusted, where sg[.] represents the stopping gradient operator and λ is a preset parameter. The goal of training is to minimize the loss.
[0150] In some embodiments, each time step corresponds to a simulation output command. After the robot executes the corresponding action, it obtains a motion reward, and the environmental state changes. The environmental state, simulation output command, corresponding motion reward, and new environmental state after the action execution at each time step are saved. Based on this data, the policy of the reinforcement learning module is updated at each time step using a reinforcement learning algorithm. For example, the policy of the reinforcement learning module is updated using a proximal policy optimization algorithm. Thus, at each time step, the policy is updated using a reinforcement learning algorithm based on the motion reward. The training process of the reinforcement learning module stops when the iteration stopping condition is met or the policy reaches the optimization objective.
[0151] In this embodiment, based on the simulation output command of the first time step and the observation data of the second time step, a simulation stitching feature for the second time step is generated. The simulation stitching feature of the second time step and the simulation stitching feature of at least one historical time step are input into the attribute prediction module to obtain predicted simulation attribute features. The simulation output command of the second time step and the target attribute features are input into the reinforcement learning module to obtain the simulation output command of the second time step. After controlling the robot's movement based on the simulation output command of the second time step, a motion reward is obtained. The attribute prediction module is trained based on the predicted simulation attribute features and the generated simulation attribute features, and the reinforcement learning module is trained using reinforcement learning based on the motion reward. This allows for effective training of the attribute prediction module and the reinforcement learning module, resulting in accurate attribute prediction and reinforcement learning modules. When outputting commands, the reinforcement learning module can consider the attribute data, interaction pose, and historical output commands of the object to be interacted with, reducing reliance on vision, better understanding the intrinsic characteristics of objects, exhibiting high generalization ability when handling unseen objects, demonstrating a high success rate in the real world, and reducing the gap between simulation and reality.
[0152] Figure 7 This is a schematic diagram of the implementation flow of a control model training method provided in an embodiment of this application. Figure 2 This method can be executed by the processor of a computer device. Based on Figure 6 The control model further includes an encoding module, and the training method further includes steps S701 and S702.
[0153] Step S701: Obtain the attribute data of the sample object.
[0154] In some embodiments, the attribute data includes at least one of the following: pivot center, pivot radius, object stiffness, object mass, object joint position, and handle grip signal. The pivot center refers to the center point or axis of rotation of an object or system. The pivot radius is the distance from the pivot center to a point on the rotating object. Object stiffness refers to the object's ability to resist deformation when subjected to external forces. Object mass is a physical quantity describing the magnitude of an object's inertia; it is a measure of the amount of matter contained in the object. Object joint position is the current orientation of an object's joints. The handle grip signal indicates whether the handle is currently gripped.
[0155] In some embodiments, the attribute data of the object to be interacted with is difficult to obtain in the real world, but can be directly obtained through simulation software during the simulation process. The attribute data of the sample object is the attribute data at the second time step.
[0156] Step S702: Input the attribute data of the sample object into the encoding module to generate the simulation attribute features.
[0157] The encoding module is used to generate simulation attribute features from the attribute data of the sample objects.
[0158] In some embodiments, the attribute data of the sample object is input into the encoding module, which generates simulated attribute features. The encoding module can be a multilayer perceptron.
[0159] In this embodiment, attribute data of the sample object is acquired; the attribute data of the sample object is input into the encoding module to generate simulation attribute features. In this way, accurate simulation attribute features can be extracted from the attribute data of the sample object.
[0160] In some embodiments, the training method further includes training the encoding module based on the predicted simulation attribute features and the generated simulation attribute features.
[0161] Among them, the encoding module is based on predicted simulation attribute features. The model is trained based on the difference between the generated simulation attribute features z and the actual simulation attribute features z. At each time step, the predicted simulation attribute features are output through the attribute prediction module. The generated simulation attribute features z are obtained through the encoding module, and the predicted simulation attribute features are then used as the basis for further analysis. The difference between the generated simulation attribute feature z and the actual value z is determined. For example, a supervised regularization loss is formulated. based on The parameters of the encoding module are adjusted, where sg[.] represents the stopping gradient operator and λ is a preset parameter. The goal of training is to minimize the loss.
[0162] In this embodiment, the encoding module is trained based on the predicted simulation attribute features and the generated simulation attribute features. This allows for accurate training of the encoding module.
[0163] In some embodiments, the motion reward includes at least one of the following: task-aware reward and motion-aware reward; the task-aware reward is used to encourage the robot to perform a desired motion sequence; the motion-aware reward is used to encourage the robot to perform smooth motion.
[0164] In this system, after the robot performs an action at each time step, the system rewards the action based on a reward function and the current environmental state. Task-aware rewards are used to encourage the robot to execute the desired motion sequence. For example, the focus of task-aware rewards is on executing the correct motion sequence, adhering to the order of A then B, rather than cheating to obtain an immediate reward. For instance, at time step t, the state is that the door is open and the doorknob is firmly gripped by the handle. The reward received was significantly greater than the state of not grabbing the doorknob. Motion-aware rewards are used to encourage robots to perform smooth movements. Motion-aware rewards encourage robots to generate smooth movements while maintaining a high success rate.
[0165] In this embodiment of the application, by setting task-aware rewards and motion-aware rewards, it is possible to ensure the correct order of task execution and generate smooth movements.
[0166] In some embodiments, the task-aware reward includes at least one of the following: a success reward to encourage the robot to complete the interactive task; a distance reward to encourage the robot to reduce the distance between its interactive device and the object to be interacted with; an object state reward to encourage the robot to increase the amount of joint position change of the object to be interacted with; and an interaction reward to encourage the robot to complete the target action.
[0167] Specifically, for the success reward, after the robot performs the interactive action at the current time step, the distance δ between the robot's interactive device and the object to be interacted with is obtained, and a first distance index is determined based on the distance δ. and the first interaction indicator and time indicator 1 s The formula for success reward is: When δ≤0.05, It should be 0.05, otherwise =0; when δ≤0.015∧1 contact In this case, It should be 0.5, otherwise 0; 1 s Consider it as 1 second. contact This represents the contact condition and can be either 1 or 0. For example, for a door opening task, δ is the distance between the robot's gripper and the door handle.
[0168] Specifically, for distance reward, the distance δ between the robot's interaction device and the object to be interacted with is obtained, and a second distance index is determined based on the distance δ. Second interaction index The formula for distance reward is: When δ≤0.015∧1 contact In this case, It should be 0.8, otherwise It is 0.
[0169] Specifically, for object state rewards, the distance δ between the robot's interaction device and the object to be interacted with is obtained, and the change in the joint position q of the object to be interacted with is also obtained. obj The first interaction index is determined based on the distance δ. and the third distance index and round length w len The formula for object state reward is: When δ≤0.05, It should be 0.5, otherwise The value is 0. For example, for a door-opening task, the change in the joint position q of the object to be interacted with is 0. obj The angle at which the door opens.
[0170] For interactions, the formula for interaction rewards is 0.2 * 1. g When δ≤0.015∧1 contact In the case of 1 g =1, otherwise 1 g The value is 0. For example, for a door-opening task, 1... g A value of 1 indicates that the gripper has caught the door handle. g A value of 0 indicates that the gripper did not grab the door handle.
[0171] Each reward corresponds to a weight. Each reward is multiplied by its corresponding weight, and the results are summed to obtain the task perception reward. For example, the weight of the success reward is 40, the weight of the distance reward is 0.6, the weight of the object state reward is 1, and the weight of the interaction reward is 0.05.
[0172] In this embodiment of the application, by dividing the task perception reward into success reward, distance reward, object state reward and interaction reward, and taking into account the task completion status and the order of action execution, the task perception reward can be accurately divided.
[0173] In some embodiments, motion perception rewards include at least one of the following: energy rewards to encourage the robot to reduce energy consumption; tracking position rewards to encourage the robot to reduce the difference between the actual position and the desired position; tracking rotation rewards to encourage the robot to reduce the difference between the actual angle and the desired angle; smoothness rewards to encourage the robot to perform smooth movements; and orientation rewards to encourage the robot to reduce movement deviations in the said orientation.
[0174] The directional rewards include: a first directional reward, which encourages the robot to reduce motion deviation in the first direction; and a second directional reward, which encourages the robot to reduce motion deviation in the second direction.
[0175] Specifically, for energy rewards, the joint torque τ and joint velocity of the robot's interactive devices are obtained. The formula for energy reward is: The weight is -0.05. For example, for a door opening task, the smaller the robot's joint torque and joint speed, the better, when the gripper grasps the door handle.
[0176] Among them, for the tracking position reward, the desired position c of the robot is obtained. pos and real location ee pos The formula for the location reward is exp(-4(c pos -ee pos ))*1 d The weight is 0.025. When δ ≤ 0.05, 1 d =1, otherwise 1 d The value is 0. The robot's desired position c is... pos and real location ee pos The smaller the difference between them, the higher the tracking location reward.
[0177] Specifically, for tracking rotation rewards, the desired rotation angle c of the robot is obtained. ori and the actual rotation angle ee ori The formula for tracking the rotation reward is exp(-4Δ(c)). ori -ee ori ))*1 d The weight is 0.004. When δ ≤ 0.05, 1 d =1, otherwise 1 d The value is 0. The robot's desired rotation angle c is... ori and the actual rotation angle ee ori The smaller the difference between them, the higher the tracking spin reward.
[0178] For the smoothness reward, the action 'a' at the current time step is obtained. tThe action a at the previous time step t-1 The formula for smoothness reward is: The weight is -0.001. sgn(a t ) is a t symbols, In sgn(a t )≠sgn(a t-1 The value is 1 if the condition is met, and 0 otherwise. The action 'a' at the current time step. t The action a at the previous time step t-1 The smaller the difference, the smoother the robot's movements.
[0179] Specifically, for the reward in the first direction, the motion component a in the first direction is obtained. t [y], the formula for the first direction reward is 1. dy *(a t [y]*15) 2 The weight is -0.005. For example, the first direction is along the y-axis, and the motion component a in the first direction... t The smaller [y] is, the better. When 0.02 ≤ δ ≤ 0.08, 1 dy The value is 1, otherwise 1. dy The value is 0.
[0180] Specifically, for the reward in the second direction, the motion component a in the second direction is obtained. t [z], the formula for the second-direction reward is 1. g *(a t [z]*15) 2 The weight is -0.07. For example, the second direction is along the z-axis, and the motion component a in the second direction... t [z] The smaller the better.
[0181] Each reward corresponds to a weight. Each reward is multiplied by its corresponding weight, and the results are summed to obtain the task perception reward.
[0182] In this embodiment of the application, motion perception reward can be accurately divided into energy reward, tracking position reward, tracking rotation reward, smoothness reward, first direction reward and second direction reward.
[0183] In some embodiments, controlling the robot's motion based on the simulation output command of the second time step includes: scaling the simulation output command to obtain a simulation control command; and controlling the robot's motion based on the simulation control command.
[0184] The simulation output commands obtained by the reinforcement learning module are not standard motion commands. The robot cannot directly execute actions based on these commands; therefore, they need to be scaled to obtain simulation control commands. The robot then executes actions based on these simulation control commands. Adjusting or scaling the simulation output commands ensures they can be executed safely and effectively by the robot, which may involve adjusting parameters such as speed, force, and range of the action. The scaled actions become "robot commands" that the robot can understand and execute. These commands are then sent to the robot's control system to drive the robot to perform the corresponding actions.
[0185] In some embodiments, the simulation output command is scaled based on a preset formula to obtain the simulation control command. For example, the simulation control command c... t The calculation formula is formula (4):
[0186] c t =clip(a t ,-1,1)*40+100 (4);
[0187] Among them, a t The clip function is a range limiting function that restricts the simulation output command to the range [-1, 1].
[0188] In some embodiments, the simulation output command is scaled by a pre-trained scaling module, the simulation output command is input into the scaling module, and the simulation control command is output.
[0189] The simulation control instructions include data related to the robot's interactive actions. Based on these instructions, the robot can be controlled to perform interactive actions at the current time step. The simulation control instructions are action commands that the robot can understand and effectively execute. For example, the simulation control instructions include target displacement, target posture, gripper action, and impedance control parameters.
[0190] In this embodiment, the simulation output command is scaled to obtain simulation control commands, and the robot's movement is controlled based on these commands. Thus, by scaling the simulation output command, simulation control commands that the robot can accurately recognize can be obtained. The robot can then accurately execute interactive actions according to the simulation control commands.
[0191] In some embodiments, the method further includes: generating a stiffness matrix of the robot based on the simulation output command; constructing a dynamic model for impedance control based on the robot's mass inertia matrix, damping matrix, and stiffness matrix; the dynamic model is used to control the balance between the robot's motion and contact force.
[0192] Stiffness represents a robot's ability to resist external deformation; by adjusting stiffness, the degree of deformation of the robot when subjected to external forces can be controlled. Damping represents the friction or energy dissipation capacity between the robot and its environment; by adjusting damping, the velocity change of the robot when subjected to external forces can be controlled. Inertia represents the robot's response speed to external forces; by adjusting inertia, the acceleration change of the robot when subjected to external forces can be controlled.
[0193] The dynamic model of impedance control is a mathematical model constructed based on the stiffness matrix, mass inertia matrix, and damping matrix. This model describes the dynamic response characteristics of the robot when subjected to external forces, that is, the balance between the robot's motion state (position, velocity, acceleration) and the contact force. By adjusting the parameters in these matrices, the robot's behavior in the contact environment can be controlled, such as maintaining a stable contact force and avoiding excessive deformation.
[0194] In some embodiments, the simulation output command of the reinforcement learning module includes stiffness. The stiffness is scaled to obtain the target stiffness, and the target stiffness is expanded to obtain the stiffness matrix. A dynamic model is constructed based on the adjusted stiffness matrix, the unadjusted mass-inertia matrix, and the damping matrix.
[0195] In this embodiment, the robot's stiffness matrix is generated based on simulation output commands; an impedance-controlled dynamic model is constructed based on the robot's mass inertia matrix, damping matrix, and stiffness matrix. Thus, by outputting stiffness parameters through a reinforcement learning module, impedance control is applied to the robot's motion, ensuring that the robot can perform smooth movements.
[0196] The following describes the application of the control method and control model training method provided in the embodiments of this application in real-world scenarios.
[0197] This application utilizes observation history as input for reinforcement learning, replacing point cloud features, to reduce the gap between simulation and reality and better understand the intrinsic properties of objects. Impedance control is employed to achieve a seamless transition from simulation to reality, resulting in smooth and compliant motion.
[0198] In some embodiments, articulated object manipulation presents a unique challenge compared to rigid object manipulation, as the object itself represents a dynamic environment. A novel reinforcement learning (RL)-based pipeline, equipped with variable impedance control and motion adaptation, is proposed for general articulated object manipulation leveraging observation history, focusing on smooth and dexterous motion during zero-shot simulation-to-real-world transfers. To reduce the simulation-to-real-world gap, the pipeline does not directly utilize visual data features (image RGBD / point cloud) as policy input, but instead first extracts useful low-dimensional data through off-the-shelf modules, thus reducing reliance on vision. Furthermore, the motion and intrinsic properties of the object are inferred from the observation history, and impedance control is applied in both simulated and real-world environments, further reducing the simulation-to-real-world gap. Additionally, a carefully designed training environment with highly stochastic and specialized reward systems (task awareness and motion awareness) is developed, enabling multi-stage, end-to-end manipulation without heuristic motion planning. Through extensive experimentation with various objects, the trained strategy achieved an 84% success rate in the real world.
[0199] Figure 8 Illustration of reinforcement learning strategy application provided in embodiments of this application Figure 1 .like Figure 8 As shown, an RL policy for opening doors and drawers was trained in the simulation. This policy adjusts its output action based on the motion of the object by utilizing historical observation data. This policy was directly applied to closed-loop variable impedance control in the real world to achieve 80% joint limit and a success rate of 84%, using only a single first-frame RGBD image.
[0200] In the diverse training settings, the observation history 801 is input into the adaptation module 802, the adaptation module 802 outputs action understanding 803, and then obtains the strategy action 804; the strategy action 804 is input into the impedance control module 805, the impedance control module 805 outputs robot command 806; the trained strategy is applied to the real world, and the trained strategy is used to operate on diverse objects 807 using only one RGBD image.
[0201] General-purpose robots represent a significant milestone in the field of robotics learning, with the potential to revolutionize everyday life. With articulated objects ubiquitous in home and industrial environments, learning how to effectively manipulate them is one of the major challenges in achieving this goal. Despite significant progress in embodied artificial intelligence, general-purpose articulated object manipulation remains an unsolved problem for various reasons. One major challenge is that true articulated characteristics (such as pivot center, friction, and stiffness) can only be identified after physical contact. For example, two objects may appear identical, yet their physical properties can be vastly different. Therefore, to achieve a general-purpose articulated object manipulation pipeline capable of seamless interaction with unseen objects, it is necessary to establish a closed-loop pipeline that adaptively infers these characteristics during the manipulation phase. Another difficulty lies in the joint constraints of the object, requiring the applied actions to conform to the actual joint motions of the object. If the robot's actions cannot tolerate joint motion and prioritize the completion of given instructions, excessive forces may be generated, potentially damaging both the object and the robot.
[0202] In related technologies, articulated object manipulation typically relies on visual information as the primary input to its pipeline. The visual input of the first frame (in the form of a point cloud or RGB image) is used to predict the operable parts and the sequence of operations or waypoint trajectories. This sequence or waypoint trajectory is then executed directly in an open-loop manner, regardless of any potential physical interactions with the object. While this approach is naturally intuitive, ignoring the object's intrinsic properties can lead to unsafe behavior. Other related technologies utilize an RL backbone, outputting actions in a closed-loop manner based on visual feedback. However, because this pipeline heavily relies on visual feedback at each iteration, a significant gap arises between the visual simulation and reality, inherited from the visual module. Furthermore, during the operation phase, this approach may output suboptimal operations due to occlusion of the operable parts. Related technologies attempt to utilize impedance control as an off-the-shelf low-level controller, adaptively adjusting predicted waypoints based on some sample-based heuristics. However, this method only affects the local trajectory between two predefined set points, resulting in unsmooth motion.
[0203] In this embodiment, closed-loop RL is combined with learnable impedance control. First, observation history is used to control the object in a closed loop, replacing visual input. Taking how a human opens a door in the dark as an example, intuition is demonstrated: given information about the doorknob's position and whether the door has a left or right hinge, the circular motion of the door is estimated based on the actions performed and the actual movement of the door. Then, even without direct visual input, the next action is progressively adjusted based on this feedback information to complete the task. Based on this intuition, utilizing observation history and reducing reliance on vision has two advantages: 1) using only vision as auxiliary information reduces the gap between visual simulation and reality; 2) utilizing observation and action history allows for implicit learning of the object's motion based on the positional error after each execution, thereby achieving a generalizable closed-loop pipeline.
[0204] Secondly, variable impedance control was introduced into the pipeline to address the importance of compliant motion in articulated object manipulation. Impedance control is suitable for tasks with high tolerance requirements for balance setpoint tracking and object joint motion, fundamentally distinguishing articulated object manipulation from rigid object manipulation. In the simulation, a high-frequency variable impedance controller was implemented while simultaneously learning its parameters using an RL policy. The inclusion of impedance control in a carefully designed training setup enabled the policy to learn smooth, continuous motion that conforms to the object's joint movements. Learning motion, rather than single actions or discrete waypoints, yields a higher success rate in real-world scenarios.
[0205] The contributions of this application are summarized as follows: A novel RL-based articulated object manipulation pipeline is proposed, using observation and motion history as primary inputs, with vision serving only as auxiliary information. A training environment is designed where each component simulates reality, and a reward function system is designed to achieve smooth, multi-stage, end-to-end manipulation without any heuristic motion planning. A variable impedance controller is introduced into RL to improve tolerance to object motion, thereby facilitating direct transfer from simulation to reality. Through extensive experiments on four tasks and 500 generalizations in the real world, zero-point inference achieves success rates of 96% in simulation and 84% in the real world, and demonstrates high versatility for a wide range of objects.
[0206] Manipulating articulated objects is extremely challenging due to the wide variety of their geometry and physical properties. Existing benchmarks in related technologies still have limited coverage. Research on articulated object manipulation can be broadly categorized into capability-based methods and RL-based methods. Capability-based methods rely on visual capability heatmaps, where each point corresponds to a success rate of manipulation, to select contact points and predict actions. However, these heatmaps are often ambiguous and difficult to annotate accurately. They also do not generalize well due to the gap between visual simulation and reality. On the other hand, RL-based methods have shown better generalization capabilities. However, utilizing point cloud features as policy inputs expands the exploration space and complicates the task. Furthermore, training with unrealistic entities, such as flying grippers or mobile Franka robots, to relax restrictions on inverse kinematics solvers (IK solvers) or collision avoidance conditions simplifies real-world tasks. The embodiments of this application utilize only low-dimensional visual information captured in the first frame and incorporate historical observations during the manipulation phase, thereby leveraging RL to better understand object motion.
[0207] In some embodiments, impedance control belongs to the position-force control family, in which position and force are not decoupled but processed simultaneously, thereby improving tolerance to feedback forces while maintaining good tracking. Many contact-rich robotic tasks, such as object placement or tool assembly, have successfully demonstrated the compatibility of such controllers with tasks that simultaneously consider setpoint tracking and object-robot force constraints. For learning-based methods, related techniques utilize impedance control as a readily available low-level controller for executing downstream instructions under policy guidance. Impedance control parameters are directly incorporated as learnable variables into RL, inverse RL, or analytical optimization methods. Variable impedance control is more suitable for different task settings and is less labor-intensive than manually adjusting impedance control. In the embodiments of this application, the application of impedance control is extended to the manipulation of articulated objects by learning impedance control parameters during training and applying them directly to the real world.
[0208] In some embodiments, given a hinged object O (corresponding to the interactive object in the above embodiments) and an operation task θ, a policy π is trained to output a dexterous action in a closed-loop manner to complete the task. The task is a more challenging and realistic adjustment based on related technologies. For pulling tasks (opening doors, drawers), the policy is required to reach, grasp the operable part, and then open until the joint position of the object reaches at least 80% of its joint limit, rather than approximately halfway. Since the robot needs to track the rigid body motion (including translation and rotation) of the object in actual three-dimensional space, this standard, especially when applied to rotary joints, requires the robot to perform a large number of dexterous and long-duration movements. Furthermore, in the setup, only robots (such as the Franka robot with a fixed base) are allowed to adopt realistically feasible inverse kinematics (IK) configurations, and unlike other pathpoint prediction pipelines using flying grippers or suction cups, the absolute feasibility of predicted motion is not assumed.
[0209] In some embodiments, the framework is designed to predict one dexterous action at a time, rather than a series of short, primitive actions. At time step t, the reinforcement learning policy (corresponding to the reinforcement learning module in the above embodiments) outputs action a. t ∈R 11 (The output command corresponding to the above embodiment) includes the target displacement. Target attitude R t ∈R 6 Gripper action G t ∈R 1 and impedance control parameters Subsequently, these actions a t The motion scaler will be used to convert the commands into robot commands. t ∈R 9 (Corresponding to the target control command in the above embodiment).
[0210] In some embodiments, at time step t, the observation metric o is acquired. t (Corresponding to the observation data in the above embodiments), including the target grasping pose g t ∈R 7 (corresponding to the first interactive pose in the above embodiment), robot joint configuration q t ∈R 7 The relative distance δ between the robot and the object t ∈R 1 End effector pose ee t ∈R 9 Including three-dimensional position and six-dimensional rotation, graspability signals
[0211] In the simulation, the target grasping pose is inferred directly from the handle bounding box, while in reality, it is obtained using an off-the-shelf grasping prediction module. In the simulation, the robot joint configuration can be obtained from the simulation software, while in reality, it is obtained from the robot's control system. The relative distance between the robot and the object can be obtained from the simulation software in the simulation, while in reality, it is obtained through sensors. The graspability signal is based on distance and contact perception conditions, rather than direct commands controlling the gripper's opening and closing. For the graspability signal, it can be obtained from the simulation software in the simulation, while in reality, it is obtained through sensors.
[0212] Among them, the observation indicator o t It also includes task-related observations, such as the pivot center with added noise in the door-opening task. Pivot radius and right hinge Boolean value Information such as this is used to guide smoother motion execution. Therefore, the observation index o t For formula (5):
[0213]
[0214] In some embodiments, privileged observation is used only in the simulation in order to better understand the environment. (Corresponding to the attribute data in the above embodiments), these are values that are difficult to track in the real world. [Privileged Observation] Including the pivot center Pivot radius object stiffness object mass Joint position of object Handle grab signal Therefore, privileged observation For formula (6):
[0215]
[0216] In some embodiments, manipulating articulated objects presents a unique challenge compared to manipulating rigid objects because the object itself is a dynamic environment. The object's motion can only be observed through physical interaction, or its actual ground position is hidden within the object, similar to motion tasks where environmental parameters (such as terrain friction and slope) are difficult to predict. To address this, an online strategy extraction pipeline, widely used in motion tasks, is employed, learning two independent modules: an adaptation module σ (corresponding to the attribute prediction module in the above embodiments) and a privileged observation encoder module φ (corresponding to the encoding module in the above embodiments). Privileged observations (e.g., pivot center, object stiffness, object mass, etc.) are used in the simulation. These features are extracted by the privileged observation encoder module, and the adaptation module is trained to infer this information from historical observations and action pairs.
[0217] The privileged observation encoder module φ is a shallow MLP used during training to learn the latent representation z of privileged observations. t (Corresponding to the simulation attribute features generated in the above embodiments). This 20-dimensional vector is paired with the (observation, action) at the current time step. Connect to form action input. Design the adaptation module σ as a temporal architecture, starting from H = 10 p... t Extract potential information about the environment (corresponding to the predicted simulation attribute features in the above embodiments). Only retain a portion of the action history as input to σ: target displacement. Gripping action G t and impedance control parameters
[0218] In this approach, because the traditional two-stage teacher-student pipeline may lead to feasibility gaps and simulation-reality gaps, the adaptation module and the privileged observation encoder module are trained simultaneously in a single training session. Specifically, when jointly training the adaptation module with the reinforcement learning (RL) backbone, it also learns to extract similar privileged information from the historical buffer. The method involves defining a supervised regularization loss on top of the PPO objective. (sg[.] represents the stopping gradient operator). Linear scheduling is applied to λ to prevent the policy from taking conservative actions in the initial phase.
[0219] In some embodiments, while the proposed framework has been widely used for motion tasks, applying this pipeline to fine manipulation tasks (such as articulated object manipulation) remains a challenge. To facilitate efficient execution of multi-stage motion using a single end-to-end policy, stage-conditional rewards are introduced, including task-aware rewards and motion-aware rewards. After the robot performs an action at each time step, the system rewards the action based on the reward function and the current environmental state. A task- and action-related reward system has been developed to ensure the correct sequence of task execution and encourage the policy to generate smooth actions.
[0220] The focus of task-aware reward is on executing the correct sequence of movements, adhering to the order of A before B, rather than cheating for immediate success. For example, at time step t, the state where the door is open and the handle is firmly gripped by the hand. The reward received was significantly greater than the state of not grabbing the doorknob.
[0221] Among these, motion-aware reward encourages the strategy to generate smooth motion while maintaining a high success rate. These clauses are typically activated after the policy training has completed its main tasks, thus acting as a fine-tuning mechanism to incentivize the policy to execute more smoothly. Incorporating these regularization clauses is crucial, helping to bridge the gap between simulation and reality by preventing unwanted motion or unattainable target poses.
[0222] In some embodiments, Table 1 represents the reward function. As shown in Table 1, it includes: 1 d The corresponding condition is δ≤0.05, 1 dy The corresponding condition is 0.02≤δ≤0.08, 1 g The corresponding condition is δ≤0.015∧1 contact τ is the joint torque. For joint velocity, w len As the round length weight, a t [y] represents the motion on the y-axis, a t [z] represents the action on the z-axis. For the reward of task perception, the formula for "success" is: The weight is 40.0; the formula for "distance" is... The weight is 0.6; the formula corresponding to "object state" is... The weight is 1.0; the formula for "scraping" is 0.2 * 1. g The weight is 0.05. For the reward of motion perception, the corresponding formula for "energy" is... The weight is -0.05; the formula for "tracking position" is exp(-4(c pos -ee pos ))*1 d The weight is 0.025; the formula for "tracking rotation" is exp(-4Δ(c ori -ee ori ))*1 d The weight is 0.004; the formula for "smoothness" is... The weight is -0.001; "y regularization" corresponds to formula 1. dy *(a t [y]*15) 2 The weight is -0.005; "z-regularization" corresponds to formula 1. g *(a t [z]*15) 2 The weight is -0.07.
[0223] Among them, for When δ≤0.05 When δ > 0.05, =0; for When δ≤0.015∧1contact hour It is 0.5, under other conditions. 0; 1 s This represents 1 second. For 1... d 1 dy 1 g The value is 1 if the corresponding condition is met, and 0 if the corresponding condition is not met. pos ee is the desired position of the strategy output. pos c is the actual location; ori ee is the desired rotation angle output by the strategy. ori This represents the actual rotation angle.
[0224] Specifically, for task perception rewards, the formula corresponding to each reward indicator is multiplied by its corresponding weight and then summed to obtain the task perception reward; for action perception rewards, the formula corresponding to each reward indicator is multiplied by its corresponding weight and then summed to obtain the action perception reward.
[0225] Table 1
[0226]
[0227]
[0228] In some embodiments, utilizing domain randomization training strategies may be beneficial for the transition from simulation to reality. The primary focus is on bridging the physical gap, requiring the strategy to understand the motion of objects through their interaction with the robot, where the inherent properties of the objects are noisy. In the simulation, training is performed by randomizing parameters such as friction, stiffness, and mass, and the predicted motion is executed using a high-frequency variable impedance controller. During training, object position and yaw rotation are randomized to cover a reasonable workspace in a real-world setting. In terms of physical intrinsics, joint friction, stiffness, and mass are altered to achieve a more robust simulation-to-realistic transition. For the desired grasping posture, random noise is introduced along the y and z axes after the posture is inferred from the part bounding box, and a randomly rotating target is introduced from a predefined spherical cone.
[0229] In some embodiments, the goal of impedance control is to take into account the external force F generated by the interaction between the robot and the environment. ext Under the circumstances, according to the expected trajectory x d Operation. The impedance control design employs a mass-spring-damped system, which dynamically adjusts the target setpoint based on feedback force and environmental stiffness. The dynamic model of the impedance control is given by formula (7):
[0230]
[0231] Where M is the robot's mass inertia matrix, D is the damping matrix, and K is the stiffness matrix. This is the impedance trajectory output. Indicates acceleration. x represents velocity. c Indicates location.
[0232] In some embodiments, within the pipeline, the stiffness coefficient k of the Cartesian impedance controller is learned and predicted. p And extend it into a six-dimensional diagonal matrix K. Assume M, K, and D are positive definite diagonal matrices to ensure system stability. The policy prediction k is obtained through formula (8). p Scaling:
[0233]
[0234] In both simulations and actual experiments, this numerical range consistently produces reasonable motion. Based on the stiffness matrix K, a critical damping condition can be deduced. The damping matrix.
[0235] In some embodiments, in policy output Then, the result is calculated using formula (8). Will Expanding this gives us K, and since M and D are known, we can obtain F. ext Based on F ext Impedance control can be applied to the robot's movements.
[0236] Figure 9 Illustration of reinforcement learning strategy application provided in embodiments of this application Figure 2 .like Figure 9 As shown, during the simulation, the privileged observation encoder module φ is trained to extract the latent representation z of the privileged observation. t Simultaneously, the adaptive module σ is trained from H = 10 previous (o) t a t-1 Similar privileged information can be inferred from this. Then, the latent representation z t p with the current time step t These are connected to form the policy input. In the real world, the trained policy is executed end-to-end based on the adaptation module σ, performing reaching, grasping, and manipulation operations. The required grasping pose is extracted using an RGBD image captured in the first frame through an off-the-shelf vision module. The reinforcement learning policy is trained in simulation and directly transferred to the real-world environment.
[0237] During the simulation training process, at time step t, the observation index o of the current time step is... t (901) and the action a at the previous time step t-1t-1 (902) constitutes an observation-action pair p t (903), p t (903) Stored in the history buffer 904. Adaptation module 905, based on the H historical observation-action pairs (p... t-1 to p t-H Extracting similar privileged information (906). Privileged observation encoder module 907 according to privileged observation (908) Latent representation of privileged observations z t (909). Randomize p t and Combining, or putting p t With z t The combination serves as input to the reinforcement learning module 910. This allows the downstream policy to adapt to z. t and At each time step t, the reinforcement learning module outputs action a. t (911), for a t Scaling is performed to obtain robot command c. t (912), c t The input impedance control system 913 applies impedance control during the execution of actions. The adaptation module and privileged observation encoder module are based on z... t and The reinforcement learning module is trained based on the differences between actions, and is trained using the PPO algorithm based on the rewards corresponding to the actions.
[0238] In real-world applications, the robot's camera acquires an RGB image 914 and a corresponding depth image 915. Based on the depth image, a point cloud 916 corresponding to the environment is obtained. The RGB image is input into the first module (SAM) 917, which outputs a mask image of the image to be grasped. The point cloud is input into the second module (GSNet) 918, which outputs multiple grasping poses. The target grasping pose 919 is determined from the multiple grasping poses using the mask image of the image to be grasped. The observation metric o containing the target grasping pose at the current time step is then used. t (920) and the output action a of the previous time step t-1 (921) Forming an observation-action pair p t (922), p t Stored in the history buffer 923. The adaptation module 924 adapts the H historical observation-action pairs (p...) t-1 to p t-H Extracting similar privileged information (925). [The following is a list of items / items] and p tThe combined input is fed into reinforcement learning module 926, which outputs the action 'a' at the current time step. t (927), for a t Scaling is performed to obtain robot command c. t (928), c t The input impedance control system 929 applies impedance control during the execution of actions.
[0239] In some embodiments, extensive evaluations were conducted in simulated and real-world environments to validate the effectiveness of the proposed method. For data and task settings, in simulations, experiments were conducted on the IsaacGym simulator and the large-scale PartNet-Mobility dataset, following the settings of relevant techniques. Simulations were performed using a Franka robot with a fixed base and a total of 346 articulated 3D objects, including doors and drawers (modified furniture). In real-world environments, experiments were conducted using the Franka Emika robotic arm equipped with a handheld RealSense D415 camera to capture RGBD images of various household items. The first frame of the RGBD image was used for operable part point cloud extraction and grasping prediction using the off-the-shelf modules Segment Anything (SAM) and GSNet.
[0240] The proposed pipeline was evaluated using the following two tasks: OpenDoor / OpenDoor+ and OpenDrawer / OpenDrawer+. OpenDoor / OpenDoor+: A door is initially closed, and the robot needs to open it to a position greater than 15% or 80% of its maximum opening width. A key requirement of this task setting is that the gripper must firmly grasp the door handle when opening the door, and cheating by opening the door from the side or using the robot's body is not allowed. OpenDrawer / OpenDrawer+: A drawer is initially closed, and the robot needs to open it to a position greater than 20% or 80% of its maximum opening length. Similar to the door opening task, the gripper must firmly grasp the handle when opening the drawer. Success rate (SR) was used as the primary evaluation metric in both the simulation and real-world settings.
[0241] In some embodiments, the proposed methods are compared with articulated object manipulation pipelines employing a simulation-to-real RL paradigm. Comparison Method 1: Directly uses a proximal policy optimization algorithm to learn a state-based policy for each task, with detailed PPO parameters and training strategies similar to those in the embodiments of this application. Comparison Method 2: This is an affordability learning framework that utilizes partial point cloud predictions to assess the affordability of visual operability, incorporating partial masks as an additional dimension to the task while keeping other aspects unchanged. Comparison Method 3: This is a vision-based policy learning method that first trains a state-based expert using part-based normalization and part-aware rewards, then refines the knowledge into a vision-based student policy. Comparison Method 4: This is a pure image learning method that utilizes a hand-eye monocular camera to actively perceive connected objects from multiple angles to improve the accuracy of 6D pose estimation. Comparison Method 5: A vision-based method that first performs cross-category part segmentation and pose estimation, then uses the predicted part poses for heuristic manipulation.
[0242] To highlight the contribution and effectiveness of each module, four comprehensive ablation studies were conducted: strategy-free distillation: using only the current time step. t The policy is trained using observed data, omitting the adaptation module and privileged observation encoder module. No variable impedance control: Cartesian position control is used as the low-level controller for the policy. No regularization: Motion-aware rewards are excluded from the reward function. No randomization: All forms of randomization are excluded, including object pose, desired grasping pose, friction, stiffness, mass, and intrinsic noise factors.
[0243] In some embodiments, the results of the simulation experiments are shown in Table 2. It can be seen that while most baseline methods perform reasonably well on the training set, their performance on the test set shows a significant downward trend. In contrast, the method of this embodiment consistently maintains strong performance on the evaluation set without a sharp decline, highlighting the excellent generalization ability of the method of this embodiment. Even without any direct gain reward, the controller can learn to adapt to different operational phases.
[0244] Table 2
[0245]
[0246]
[0247] Figure 10 This is a schematic diagram of the experimental setup for an embodiment of this application. The strategy was extensively evaluated in the real world, with test subjects including a large number of unseen items that varied in appearance 1001, size 1002, hinge direction 1003, hinge stiffness 1004, and 6D posture 1005. Figure 10 As shown, Figure 10 The different images in the document demonstrate that different objects vary in appearance (1001), size (1002), hinge direction (1003), hinge stiffness (1004), and 6D posture (1005). Our performance was showcased in a reasonable workspace, with test objects facing forward or slightly tilted around the z-axis.
[0248] Figure 11 This is a schematic diagram of the experimental results of an embodiment of this application. Figure 11 The horizontal axis represents the number of time frames; the larger the number of frames, the further back in time the data is processed. The vertical axis represents the stiffness parameter k. p The value of k. It can be seen that in the original k... p When it is greater than 1, for the original k p Scaling is performed, scaling k p It is 140; in the original k p When less than -1, for the original k p Scaling is performed, scaling k p The value is 60. At the initial time, the robot's stiffness parameter k... p The stiffness parameter k of the robot is relatively large, and as time progresses, it increases during the process of approaching the object. p It will get smaller, and the robotic arm will become softer.
[0249] In some embodiments, from Figure 11 It can be seen that the learned impedance control parameters can actively adapt to the operational phase, even without direct gain rewards: impedance control hardens as it approaches the target and softens as it approaches the target. From Figure 11 It can be seen that when the robotic arm moves away from the object, by setting the impedance control parameter to a higher k... p The robotic arm becomes stiffer when the distance is shortened to minimize collision penalty; conversely, it softens when the distance is reduced. p The smaller.
[0250] In some embodiments, Table 3 lists the policy output performance in the real world. Fifty experiments were conducted on different objects for the pipeline and each ablation model (500 runs in total).
[0251] Table 3
[0252]
[0253] In some embodiments, based on Figure 10The success rate was further investigated by distinguishing between cases where failure occurred during the grasping phase due to grasping pose estimation failure and cases where failure occurred during the opening phase due to pipeline failure. For OpenDoor+, it was found that 6 / 50 of the inferences failed during the grasping phase, while only 4 / 50 failed during the opening phase. This suggests that if a stable grasping pose is initiated, the strategy may achieve a success rate of 40 / 44 = 0.90. For OpenDrawer+, 7 / 8 of the failure cases were due to unsuccessful grasping.
[0254] In some embodiments, the ablation study results shown in Table 3 indicate that, in addition to the SR decline in both simulation and the real world, several other issues were identified, highlighting the non-smooth motion execution in the real world. For control without variable impedance, the primary cause of the failure cases (a 40% decline) was the low flexibility of position control, requiring precise execution of each predicted action. This generates significant joint torque to overcome the object's feedback force, causing the robotic arm to be triggered to stop. In the simulation, this behavior did not appear to significantly impact performance, as evidenced by a success rate greater than 0.8.
[0255] However, in the real world, high torque is extremely dangerous and can trigger emergency stops, highlighting the necessity of impedance control. Even with manual adjustment of the impedance control baseline, "no-strategy distillation" and "no randomization" often fail halfway through. This behavior is attributed to the gap between physical simulation and reality caused by non-variable training settings and short-term observations. In the "no regularization" case, the reaching and opening movements are very stiff, which is highly undesirable and can lead to grasping failures and loss of contact during execution.
[0256] In some embodiments, the simulation was evaluated using a modified subset of StorageFurniture from PartNet-Mobility, containing 346 objects, achieving a success rate of 93% for opening doors and 96% for opening drawers. In the real world, validation was performed using 10 different articulated objects, varying in appearance, size, hinge direction, hinge stiffness, and 6D pose. The policy achieved an 80% success rate for opening doors and an 84% success rate for opening drawers. The policy demonstrated good generalization ability on unseen objects, with a success rate higher than all previous work. Furthermore, it generated smooth and continuous motions not seen in previous work.
[0257] This application introduces a reliable RL strategy and seamlessly deploys it to various real-world environments. Experiments in simulated and real-world scenarios demonstrate that, in simulations, the operational phase should be learned as a smooth, continuous motion, rather than a discrete waypoint. Combined with the tolerance of impedance control, closed-loop real-world transmission becomes more efficient even with slightly less-than-ideal motion prediction.
[0258] This application introduces a novel RL framework equipped with variable impedance control for end-to-end articulated object manipulation. This framework adaptively learns object motion through observation and motion history, rather than naively executing a predicted trajectory before robot contact with the object. It demonstrates strong simulation-to-reality conversion capabilities on various real-world test objects, achieving success rates of 80% and 84% in door-opening and drawer-opening tasks, respectively, outperforming related technologies. While obtaining quantifiable results, the policy generates smooth and dexterous movements due to carefully designed training settings and reward functions. This application provides an alternative method for utilizing visual information and other potential patterns (such as tactile grasping signals), thereby better bridging the gap between simulation and reality for future RL-based manipulation.
[0259] In some embodiments, this solution has been used in the application of opening and closing doors and drawers in home settings. Implementation: Through reinforcement learning strategies in both simulation and the real world, the robot is trained to recognize the 6D poses of doors and drawers, and high-frequency variable impedance control is used to ensure smooth movement. Grasping posture input provided by vision modules (such as SAM and GSNet) ensures the robot can successfully grasp and open / close household items. Results: The robot can perform opening and closing operations in simulation with a high success rate (93% for opening doors, 96% for opening drawers), and achieves a success rate of 80% or higher in reality. This method improves the automation capabilities of home service robots in handling everyday items and ensures smooth and continuous movements.
[0260] In some embodiments, this solution may be used in assembly line robots in automobile manufacturing. Implementation: In automobile production, robots can utilize this technology to automatically assemble various articulated components (such as doors, hoods, etc.). The robot can automatically adjust its operating strategy based on the posture and 6D pose information of different assembled parts, achieving precise motion control. Effects: In automobile manufacturing, this technology can improve assembly accuracy and efficiency, especially when handling complex mechanical parts, enabling smooth assembly movements, reducing equipment wear, and improving the overall automation level of the production line.
[0261] Based on the foregoing embodiments, this application provides a control component, which includes the included units and the modules included in each unit, which can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0262] Figure 12 This is a schematic diagram of the composition structure of a control component provided in an embodiment of this application, as shown below. Figure 12 As shown, the control component 1200 includes: a determining unit 1210, a generating unit 1220, a first input unit 1230, and a second input unit 1240, wherein:
[0263] The determining unit 1210 is used to determine the first interaction pose for the object to be interacted with based on the environmental image of the robot.
[0264] The generation unit 1220 is used to generate the splicing features of the current time step based on the output command of the previous time step and the first interactive pose.
[0265] The first input unit 1230 is used to input the concatenated features of the current time step and the concatenated features of at least one historical time step into the trained attribute prediction module to obtain attribute features; the attribute features are used to characterize the attribute data of the object to be interacted with that are related to the interaction.
[0266] The second input unit 1240 is used to input the attribute features and the concatenated features of the current time step into the trained reinforcement learning module to obtain the output instruction of the current time step.
[0267] In some embodiments, the trained attribute prediction module and the trained reinforcement learning module are obtained after training using attribute data related to the interaction of the sample objects to be interacted with in the simulation environment; wherein, at any time step in the training process, the attribute prediction module adjusts its parameters based on the difference between the simulation attribute features generated at that time step and the simulation attribute features predicted by the attribute prediction module; the generated simulation attribute features are obtained based on the attribute data of the sample objects; the reinforcement learning module adjusts its parameters based on motion rewards, which are obtained after the robot responds to the simulation output command to move, and the simulation output command is output by the reinforcement learning module at that time step based on the generated simulation attribute features or the predicted simulation attribute features.
[0268] In some embodiments, the first input unit 1230 is further configured to store the splicing features of the current time step in a history cache; and to obtain the splicing features of the at least one historical time step from the history cache.
[0269] In some embodiments, the second input unit 1240 is further configured to scale the output command to obtain a target control command.
[0270] In some embodiments, the second input unit 1240 is further configured to control the interactive device based on the target control command to perform an interactive action at the current time step.
[0271] In some embodiments, the environmental image includes a depth image and a color image; the determining unit 1210 is further configured to perform object recognition on the color image to obtain a mask image of the object to be interacted with; convert the depth image to obtain environmental point cloud data; use the environmental point cloud data to determine a plurality of second interaction poses; and use the mask image to determine the first interaction pose among the plurality of second interaction poses.
[0272] In some embodiments, the training method of the attribute prediction module includes: generating simulation splicing features for a second time step based on simulation output instructions at a first time step and observation data at a second time step; the first time step being the previous time step of the second time step; inputting the simulation splicing features of the second time step and simulation splicing features from at least one historical time step into the attribute prediction module to obtain predicted simulation attribute features; the simulation attribute features are used to characterize attribute data of a sample object to be interacted with in the simulation environment that are related to the interaction; inputting the simulation splicing features of the second time step and target attribute features into a reinforcement learning module to obtain simulation output instructions for the second time step; the target attribute features being the predicted simulation attribute features or simulation attribute features generated based on the attribute data of the sample object; and training the attribute prediction module based on the predicted simulation attribute features and the generated simulation attribute features.
[0273] In some embodiments, the training method of the reinforcement learning module includes: after controlling the robot to move based on the simulation output command of the second time step, obtaining a motion reward; and training the reinforcement learning module through reinforcement learning based on the motion reward.
[0274] This application provides a robot, which includes the control components described above.
[0275] Figure 13 This is a schematic diagram of the composition structure of a training device provided in an embodiment of this application, as shown below. Figure 13 As shown, the training device 1300 includes: a simulation generation module 1310, a first input module 1320, a second input module 1330, an acquisition module 1340, and a training module 1350, wherein:
[0276] The simulation generation module 1310 is used to generate simulation splicing features of the second time step based on the simulation output command of the first time step and the observation data of the second time step; the first time step is the previous time step of the second time step.
[0277] The first input module 1320 is used to input the simulation splicing features of the second time step and the simulation splicing features of at least one historical time step into the attribute prediction module to obtain the predicted simulation attribute features; the simulation attribute features are used to characterize the attribute data of the sample objects to be interacted with in the simulation environment that are related to the interaction.
[0278] The second input module 1330 is used to input the simulation output command and target attribute features of the second time step into the reinforcement learning module to obtain the simulation output command of the second time step; the target attribute features are the predicted simulation attribute features or simulation attribute features generated based on the attribute data of the sample object.
[0279] The acquisition module 1340 is used to acquire motion rewards after controlling the robot's movement with simulation output instructions based on the second time step;
[0280] Training module 1350 is used to train the attribute prediction module based on the predicted simulation attribute features and the generated simulation attribute features, and to train the reinforcement learning module based on the motion reward through reinforcement learning.
[0281] In some embodiments, the control model further includes an encoding module; the simulation generation module 1310 is further configured to acquire the attribute data of the sample object; input the attribute data of the sample object into the encoding module to generate the simulation attribute features.
[0282] In some embodiments, the training module 1350 is further configured to train the encoding module based on the predicted simulation attribute features and the generated simulation attribute features.
[0283] In some embodiments, the motion reward includes at least one of the following: task-aware reward and motion-aware reward; the task-aware reward is used to encourage the robot to perform a desired motion sequence; the motion-aware reward is used to encourage the robot to perform smooth motion.
[0284] In some embodiments, the task-aware reward includes at least one of the following: a success reward to encourage the robot to complete the interactive task; a distance reward to encourage the robot to reduce the distance between its interactive device and the object to be interacted with; an object state reward to encourage the robot to increase the amount of joint position change of the object to be interacted with; and an interaction reward to encourage the robot to complete the target action.
[0285] In some embodiments, the motion perception reward includes at least one of the following: an energy reward to encourage the robot to reduce energy consumption; a tracking position reward to encourage the robot to reduce the difference between the actual position and the desired position; a tracking rotation reward to encourage the robot to reduce the difference between the actual angle and the desired angle; a smoothness reward to encourage the robot to perform smooth movements; and a direction reward to encourage the robot to reduce movement deviation in the said direction.
[0286] In some embodiments, the acquisition module 1340 is further configured to scale the simulation output command to obtain a simulation control command; and control the robot's movement based on the simulation control command.
[0287] In some embodiments, the acquisition module 1340 is further configured to generate the stiffness matrix of the robot based on the simulation output command; construct an impedance control dynamic model based on the robot's mass inertia matrix, damping matrix and stiffness matrix; the dynamic model is used to control the balance between the robot's motion and contact force.
[0288] The descriptions of the apparatus embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. In some embodiments, the functions or modules included in the apparatus provided in this application can be used to perform the methods described in the method embodiments above. For technical details not disclosed in the apparatus embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0289] It should be noted that, in the embodiments of this application, if the above-described control method and control model training method are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the related technology, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0290] This application provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.
[0291] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium can be transient or non-transient.
[0292] This application provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.
[0293] This application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.
[0294] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referred to interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0295] Figure 14 This application provides a hardware entity diagram of a computer device as an embodiment of the present application, such as... Figure 14 As shown, the hardware entity of the computer device 1400 includes a processor 1401 and a memory 1402, wherein the memory 1402 stores a computer program that can run on the processor 1401, and the processor 1401 executes the program to implement the steps in the method of any of the above embodiments.
[0296] The memory 1402 stores computer programs that can run on the processor. The memory 1402 is configured to store instructions and applications that can be executed by the processor 1401. It can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 1401 and various modules in the computer device 1400. It can be implemented by flash memory or random access memory (RAM).
[0297] The processor 1401 executes the program to implement the steps of any of the above-mentioned control methods or control model training methods. The processor 1401 typically controls the overall operation of the computer device 1400.
[0298] This application provides a computer storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the control method or control model training method as described in any of the above embodiments.
[0299] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0300] The aforementioned processor can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), Controller, Microcontroller, and Microprocessor. It is understood that other electronic devices can also implement the functions of the aforementioned processor, and this application does not specifically limit the specific implementation.
[0301] The aforementioned computer storage media / memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various terminals that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0302] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A control method, characterized in that, Applied to robots, the method includes: Determine the first interaction pose for the object to be interacted with based on the image of the robot's environment. Based on the output command of the previous time step and the first interactive pose, generate the splicing features of the current time step; The concatenated features of the current time step and the concatenated features of at least one historical time step are input into the trained attribute prediction module to obtain attribute features; the attribute features are used to characterize the attribute data of the object to be interacted with that are related to the interaction. The attribute features and the concatenated features of the current time step are input into the trained reinforcement learning module to obtain the output instructions for the current time step.
2. The method according to claim 1, characterized in that, The trained attribute prediction module and the trained reinforcement learning module are obtained by training on the attribute data of the sample objects to be interacted with in the simulation environment that are related to the interaction. In this process, at any time step, the attribute prediction module adjusts its parameters based on the difference between the simulated attribute features generated at that time step and the simulated attribute features predicted by the attribute prediction module; the generated simulated attribute features are obtained based on the attribute data of the sample object. The reinforcement learning module adjusts parameters based on motion rewards, which are obtained after the robot responds to the simulation output command. The simulation output command is output by the reinforcement learning module at any time step based on the generated simulation attribute features or the predicted simulation attribute features.
3. The method according to claim 1 or 2, characterized in that, The method further includes: Store the splicing features of the current time step in the history cache; retrieve the splicing features of at least one historical time step from the history cache; And / or, scale the output command to obtain a target control command; control the interactive device based on the target control command to execute the interactive action at the current time step; And / or, the environmental image includes a depth image and a color image; determining the first interaction pose for the object to be interacted with based on the environmental image of the robot includes: Object recognition is performed on the color image to obtain a mask image of the object to be interacted with; The depth image is converted to obtain environmental point cloud data; Using the environmental point cloud data, multiple second interaction poses are determined; The first interaction pose is determined using the mask image among the plurality of second interaction poses.
4. The method according to claim 1 or 2, characterized in that, The training method for the attribute prediction module includes: Based on the simulation output command of the first time step and the observation data of the second time step, the simulation stitching feature of the second time step is generated; the first time step is the previous time step of the second time step. The simulation splicing features of the second time step and the simulation splicing features of at least one historical time step are input into the attribute prediction module to obtain the predicted simulation attribute features; the simulation attribute features are used to characterize the attribute data of the sample objects to be interacted with in the simulation environment that are related to the interaction. The simulation stitching features and target attribute features of the second time step are input into the reinforcement learning module to obtain the simulation output command of the second time step; the target attribute features are the predicted simulation attribute features or simulation attribute features generated based on the attribute data of the sample object. The attribute prediction module is trained based on the predicted simulation attribute features and the generated simulation attribute features; The training method for the reinforcement learning module includes: After controlling the robot's movement based on the simulation output command at the second time step, a motion reward is obtained; The reinforcement learning module is trained using the aforementioned exercise rewards through a reinforcement learning approach.
5. A method for training a control model, characterized in that, The control model includes an attribute prediction module and a reinforcement learning module; the training method includes: Based on the simulation output command of the first time step and the observation data of the second time step, the simulation stitching feature of the second time step is generated; the first time step is the previous time step of the second time step. The simulation splicing features of the second time step and the simulation splicing features of at least one historical time step are input into the attribute prediction module to obtain the predicted simulation attribute features; the simulation attribute features are used to characterize the attribute data of the sample objects to be interacted with in the simulation environment that are related to the interaction. The simulation output command and target attribute features of the second time step are input into the reinforcement learning module to obtain the simulation output command of the second time step; the target attribute features are the predicted simulation attribute features or simulation attribute features generated based on the attribute data of the sample object. After controlling the robot's movement based on the simulation output command at the second time step, a motion reward is obtained; The attribute prediction module is trained based on the predicted simulation attribute features and the generated simulation attribute features, and the reinforcement learning module is trained based on the motion reward through reinforcement learning.
6. The method according to claim 5, characterized in that, The control model further includes an encoding module; the training method further includes: Obtain the attribute data of the sample object; The attribute data of the sample object is input into the encoding module to generate the simulation attribute features; The encoding module is trained based on the predicted simulation attribute features and the generated simulation attribute features.
7. The method according to claim 5 or 6, characterized in that, The motion reward includes at least one of the following: task-aware reward and motion-aware reward; the task-aware reward is used to encourage the robot to perform a desired motion sequence; the motion-aware reward is used to encourage the robot to perform smooth motion; The task-aware reward includes at least one of the following: Success rewards are used to encourage robots to complete interactive tasks; Distance rewards are used to encourage reducing the distance between the robot's interaction device and the object to be interacted with; Object state reward is used to encourage the robot to increase the amount of joint position change of the object to be interacted with; Interactive rewards are used to encourage robots to perform actions that complete the target task. And / or, the motion-aware reward includes at least one of the following: Energy rewards are used to encourage robots to reduce energy consumption; Tracking location rewards are used to encourage the robot to reduce the difference between its actual location and its desired location; Tracking rotation rewards are used to encourage the robot to reduce the difference between the actual angle and the desired angle; Smoothness reward is used to encourage robots to perform smooth movements; Directional rewards are used to encourage the robot to reduce motion deviations in the stated direction.
8. The method according to claim 5 or 6, characterized in that, The simulation output command based on the second time step for controlling robot movement includes: scaling the simulation output command to obtain a simulation control command; and controlling the robot movement based on the simulation control command. And / or, the method further includes: The robot's stiffness matrix is generated based on the simulation output instructions; An impedance control dynamic model is constructed based on the robot's mass inertia matrix, damping matrix, and stiffness matrix; the dynamic model is used to control the balance between the robot's motion and contact force.
9. A control component, characterized in that, The control components, installed in the robot, include: The determining unit is used to determine the first interaction pose for the object to be interacted with based on the environmental image of the robot. The generation unit is used to generate the splicing features of the current time step based on the output instructions of the previous time step and the first interactive pose. The first input unit is used to input the concatenated features of the current time step and the concatenated features of at least one historical time step into the trained attribute prediction module to obtain attribute features; the attribute features are used to characterize the attribute data of the object to be interacted with that are related to the interaction. The second input unit is used to input the attribute features and the concatenated features of the current time step into the trained reinforcement learning module to obtain the output instruction of the current time step.
10. A robot, characterized in that, The robot includes the control components as described in claim 9.
Citation Information
Patent Citations
Model training method and device, robot control method and device, terminal and storage medium
CN118196586A
Training autoencoders for generating latent representations
US20240096077A1