A Method for Transferring Simulation of Robot Contact-Intensive Operations to Reality Based on Contact Force Direction Prediction

By constructing expert control strategies and adaptive compliant controllers in a simulation environment, the problems of high cost of real data and difficulty in migrating simulation data in robot contact-intensive operations are solved, enabling stable and safe operation of robots in complex environments.

CN122490868APending Publication Date: 2026-07-31ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2026-07-02
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies for robot contact-intensive operations suffer from problems such as high cost of real data, difficulty in transferring force prediction strategies trained on simulation data to the real world, and lack of contact intent expression in pure position strategies, leading to operational instability and safety risks.

Method used

An expert control strategy is constructed in a simulation environment. Demonstration data containing the contact point normal is collected, the strategy network is trained to predict the contact state and the desired normal contact force direction, and an adaptive compliant controller is deployed. Based on the prediction results, damping control is applied in the normal direction, high stiffness control is applied in the tangential direction, and low stiffness tracking is maintained in other directions, realizing the direct transfer from pure simulation training to the real environment.

Benefits of technology

It improves the success rate and stability of contact-intensive operations, reduces training costs, enhances robustness to external disturbances, and enables robots to operate safely and reliably in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490868A_ABST
    Figure CN122490868A_ABST
Patent Text Reader

Abstract

This invention discloses a method for transferring simulation-to-real-world contact-intensive operations in robots based on contact force direction prediction, belonging to the field of robot control technology. The method includes: constructing an expert control strategy in a simulation environment to collect demonstration data; training the strategy network using imitation learning based on the demonstration data; deploying the trained strategy network on a real robot; when no contact is predicted, generating omnidirectional low-stiffness admittance control signals through an adaptive compliant controller to track the predicted pose; when contact is predicted, generating a damping control signal in the desired normal contact force direction to maintain the set contact force, generating a high-stiffness admittance control signal in the tangential direction where the pose difference is determined to overcome resistance, and generating low-stiffness admittance control signals in other directions through the adaptive compliant controller; and executing operations according to the admittance control signals. This invention enables direct transfer from pure simulation training to the real environment, thereby stably and safely completing contact-intensive operation tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot control technology, specifically relating to a method for transferring simulation of robot contact-intensive operations to reality based on contact force direction prediction. Background Technology

[0002] Contact-intensive operations are widely used in practical applications such as workpiece assembly, door and cabinet opening and closing, and surface wiping, and are a crucial capability for robots to complete complex interactive tasks. In contact-intensive operations, robots not only need to reach the appropriate pose but also need to apply appropriate contact forces when interacting with the environment to ensure stability and safety during the operation. How to effectively learn control strategies incorporating force information to enable robots to reliably and safely complete contact-intensive operations is a significant challenge in the field of robot manipulation.

[0003] Existing technical solutions can be divided into the following two main categories: The first category of methods involves policy learning directly in the real world, including imitation learning-based and reinforcement learning-based methods. Imitation learning-based methods use manually collected demonstration data from the real world to train policies that predict desired contact forces or controller stiffness parameters. Reinforcement learning-based methods optimize control policies through repeated trial and error during real robot-environment interactions. These methods can obtain accurate force data in real-world environments, resulting in excellent performance of trained policies in data-covered environments. However, collecting real-world data is costly, inefficient, and poses safety risks, making it difficult to scale to diverse task scenarios.

[0004] The second type of approach first uses simulation data for policy learning, and then transfers it to the real world for deployment. Simulation environments can collect large-scale and diverse demonstration data at low cost and extremely high efficiency. However, the contact dynamics in simulation environments differ significantly from those in the real world, resulting in inaccurate forces generated when the robot comes into contact with the environment in simulation. Using such force data with obvious differences between simulation and reality to train policies will lead to a rapid decline in performance after being transferred to the real world.

[0005] To achieve the transfer of simulation to reality for robot contact-intensive operations, existing technologies include: Dynamic parameter calibration methods: By identifying the system or manually adjusting parameters, the simulator parameters are matched to the actual contact properties, thereby reducing the difference between simulated and real force data. However, each task requires individual calibration, and calibration usually requires real robot-environment interaction data, making it difficult to model complex contact conditions.

[0006] Fine-tuning with a small amount of real data: This method first trains the policy in simulation, then fine-tunes it using a small amount of real data to improve its performance in real-world environments. However, this method still requires the collection of real data, and is therefore limited by the cost, security, and scalability issues associated with real data.

[0007] Dynamic randomized reinforcement learning involves randomizing dynamic parameters such as friction, stiffness, and mass over a wide range, and then training a control strategy adaptable to various dynamics through reinforcement learning. However, the dynamic differences between the simulation environment and the real world mainly stem from the contact dynamics modeling errors of the simulator. Randomizing the parameters of the contact dynamics model may not necessarily cover the contact dynamics in the real world. Moreover, this method essentially requires implicitly training a dynamic parameter identifier, which rapidly increases the training cost and difficulty.

[0008] Combining a pure positional strategy with a passive compliant controller: A pure positional strategy is trained using simulation data, and then combined with a fixed passive compliant controller in the real world to ensure safe execution. However, the passive compliant controller has uniform stiffness in all directions and lacks task-specific active adaptation, often requiring a trade-off between completing the task and ensuring safety.

[0009] In summary, the core challenges currently faced by the technology are: the high cost and difficulty in scaling up real-world data; the difficulty in transferring force prediction strategies trained on simulation data to the real world due to the dynamic differences between simulation and reality; and the lack of contact intent expression in pure position strategies trained on simulation data, thus only achieving passive compliance. Therefore, there is an urgent need for a simulation-to-real-world transfer learning method that can effectively overcome the differences in contact dynamics and achieve reliable transfer of contact force information. Summary of the Invention

[0010] In view of the above, the purpose of this invention is to provide a method for transferring robot contact-intensive operations from simulation to reality based on contact force direction prediction. By constructing an expert control strategy in the simulation environment, collecting demonstration data containing the true values ​​of the contact point normal, and training the strategy network, it is possible to learn transferable contact force direction prediction capabilities under pure simulation data conditions. When deployed in the real environment, an adaptive compliant controller that can actively and adaptively adjust the contact force according to the contact force direction is designed. Based on the prediction results, damping control is achieved in the normal direction, high stiffness control in the tangential direction, and low stiffness tracking is maintained in other directions. This enables direct transfer from pure simulation training to the real environment, thereby stably and safely completing various contact-intensive operation tasks such as workpiece assembly, door and cabinet opening and closing, and surface wiping.

[0011] To achieve the above-mentioned objectives, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for transferring simulation to reality of robot contact-intensive operations based on contact force direction prediction, comprising the following steps: In a simulation environment, an expert control strategy is constructed, and the operation task is divided into a free motion stage and a contact interaction stage. Pose trajectories are generated and executed to collect demonstration data. In the contact interaction stage, the surface normal of the object at the contact point is recorded as the true value of the direction of the expected normal contact force. The policy network is trained by imitation learning based on the demonstration data, and the end pose, expected normal contact force direction and contact state of the next several steps are output. The trained policy network is deployed on a real robot. When no contact is predicted, an adaptive compliant controller generates low-stiffness admittance control signals in all directions to track the predicted pose. When contact is predicted, the adaptive compliant controller generates damping control signals in the direction of the desired normal contact force to maintain the set contact force, generates high-stiffness admittance control signals in the tangential direction determined by the difference between the current pose and the predicted pose to overcome resistance, and generates low-stiffness admittance control signals in the other directions. The robot performs operations according to the admittance control signals.

[0012] Preferably, the expert control strategy generates a pose trajectory from the current end-effector pose to the key pose centered on the object through linear interpolation during the free motion phase and executes it sequentially. During the contact interaction phase, it generates an end-effector pose trajectory that satisfies the geometric and mechanical constraints of the task and executes it sequentially according to the rules of each operation task.

[0013] Preferably, during the demonstration data collection process, the lighting, background, object pose, initial end-effector pose of the robotic arm, camera intrinsic and extrinsic parameters are randomized each time the simulation environment is reset, and the task completion status is automatically determined according to the task success indicators provided by the simulator, retaining only the successful demonstration data.

[0014] Preferably, the input to the policy network includes the observed image at the current moment, the pose of the robotic arm end effector, and the task language description, and the output includes the robotic arm end effector pose, the desired normal contact force direction, and the action sequence of binary contact state for multiple future time steps.

[0015] Preferably, the policy network includes a visual encoding module, a language encoding module, a conditional fusion module, and a diffusion policy module; the visual encoding module is used to encode the observed image into visual features, the language encoding module is used to encode the task language description into language features, the conditional fusion module is used to fuse visual features and language features into conditional features through a cross-attention mechanism, and the diffusion policy module is used to encode the current robotic arm end-effector pose and the noisy future action sequence, perform cross-attention calculation with the conditional features, and combine the denoised time step to perform noise prediction to generate the future action sequence.

[0016] Preferably, the dynamic equation of the adaptive compliant controller is: , in, and These are the admittance reference velocity and acceleration, respectively. For admittance reference pose, The robot arm's end-effector pose predicted by the policy network. For external contact force, In order to expect external forces, , , These are mass, damping, and stiffness coefficients, respectively.

[0017] Preferably, when the policy network predicts that the robot will not make contact with the object, the following settings are made: Furthermore, with all directions set to the default low stiffness coefficient, the calculated admittance reference pose is... This constitutes a low-stiffness admittance control signal in all directions; When the policy network predicts that the robot will make contact with an object, the contact force will be in the direction of the expected normal direction. Top settings ,in For the desired normal contact force amplitude, the calculated admittance reference pose constitutes the damping control signal in that direction. In the tangential direction determined by the difference between the current pose and the predicted pose, the stiffness coefficient is increased to several times the default low stiffness coefficient to overcome the drag. The calculated admittance reference pose constitutes the high stiffness admittance control signal in that direction. In other directions, the default low stiffness coefficient is maintained, and the calculated admittance reference pose constitutes the low stiffness admittance control signal in these directions.

[0018] Secondly, the present invention provides a robot contact-intensive operation simulation-to-reality transfer system based on contact force direction prediction, implemented using the aforementioned robot contact-intensive operation simulation-to-reality transfer method based on contact force direction prediction, including: Simulation data construction module: Constructs expert control strategies in the simulation environment, divides the operation task into free motion stage and contact interaction stage, generates and executes pose trajectories to collect demonstration data, and records the surface normal of the object at the contact point as the true value of the expected normal contact force direction in the contact interaction stage. Policy network training module: Based on the demonstration data, the policy network is trained by imitation learning, and the end pose, expected normal contact force direction and contact state of the next several steps are output. Real-world deployment module: The trained policy network is deployed on a real robot. When no contact is predicted, an adaptive compliant controller generates low-stiffness admittance control signals in all directions to track the predicted pose. When contact is predicted, the adaptive compliant controller generates damping control signals in the desired normal contact force direction to maintain the set contact force, and generates high-stiffness admittance control signals in the tangential direction determined by the difference between the current pose and the predicted pose to overcome resistance. Low-stiffness admittance control signals are generated in other directions. The robot performs operations according to the admittance control signals.

[0019] Thirdly, an electronic device provided by an embodiment of the present invention includes a memory and one or more processors. The memory is used to store a computer program, and the processor is used to implement the above-mentioned method for transferring simulation of robot contact-intensive operations based on contact force direction prediction when executing the computer program.

[0020] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a computer, implements the aforementioned method for migrating robot contact-intensive operations from simulation to reality based on contact force direction prediction.

[0021] Compared with the prior art, the beneficial effects of the present invention include at least the following: (1) Improved success rate of contact-intensive operation tasks: The present invention uses a pure simulation data training strategy network to predict the contact state and the expected normal contact force direction, and combines it with an adaptive compliant controller that actively adapts to the contact force adjustment, so that the robot can adjust the admittance control parameters in different directions according to the contact state and task requirements during operation, avoiding operation failure due to insufficient force or damage to objects or robots due to excessive force, thereby completing contact-intensive operation tasks stably and safely.

[0022] (2) Enhanced robustness to external disturbances during operation: The present invention independently adjusts stiffness and damping characteristics in different directions through an adaptive compliant controller. It achieves damping control in the normal direction during contact to maintain the set contact force, achieves high stiffness control in the tangential direction to overcome resistance, and combines low stiffness control in other directions. Thus, while ensuring the completion of the task, it significantly improves the adaptability to position errors and external environmental disturbances, avoids the safety emergency stop triggered by disturbances due to high stiffness in all directions, and improves the stability and reliability of task execution.

[0023] (3) Reduced training costs and deployment threshold for contact-intensive tasks: This invention uses simulation data entirely and automatically collects large-scale and diverse successful demonstration data using expert control strategies, eliminating the need for any real-world manual demonstrations or real robot trial and error, and avoiding tedious steps such as dynamic parameter calibration or real data fine-tuning. Compared with real data acquisition methods, it has outstanding advantages such as low labor costs, high collection efficiency, high data diversity, and high security. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart illustrating the method for transferring simulation of robot contact-intensive operations to reality based on contact force direction prediction, as provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the framework of the robot contact-intensive operation simulation to reality transfer method based on contact force direction prediction provided in the embodiments of the present invention; Figure 3 This is a schematic diagram of the structure of the robot contact-intensive operation simulation to reality transfer system based on contact force direction prediction provided in the embodiment of the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not limit the scope of protection of this invention.

[0027] The inventive concept of this invention is as follows: Addressing the problem that the significant differences between contact dynamics in simulated environments and real-world environments make it difficult to transfer force-related control strategies trained based on simulation data to real-world applications, this invention provides a method for migrating robot contact-intensive operations from simulation to reality based on contact force direction prediction. By collecting demonstration data containing the surface normal of the contact point in a simulated environment using an expert control strategy, a strategy network is trained to predict the contact state and the desired normal contact force direction. This allows the subsequent adaptive compliant controller to distinguish between the contact force direction and the contact intention based on the prediction results. In real-world deployment, the controller applies damping control in the normal direction to maintain the set contact force, applies high stiffness control in the tangential direction to overcome resistance, and maintains low stiffness compliant tracking in other directions. This allows the strategy network trained purely on simulation to stably perform contact tasks in real-world environments without the need for fine-tuning with real data or calibration of dynamic parameters.

[0028] like Figure 1 As shown in the embodiment, this method provides a simulation-to-real-world transfer method for robot contact-intensive operations based on contact force direction prediction, including the following steps: S1. In the simulation environment, an expert control strategy is constructed, and the operation task is divided into a free motion stage and a contact interaction stage. The pose trajectory is generated and executed to collect demonstration data. In the contact interaction stage, the surface normal of the object at the contact point is recorded as the true value of the direction of the expected normal contact force.

[0029] S1.1, Simulation environment construction.

[0030] In this embodiment, the Isaac Lab simulation platform is used to construct a training environment for contact-intensive tasks.

[0031] (1) Construction of Robot and Object Simulation Assets: Robot simulation assets including robotic arms and grippers were provided by the Isaac Lab platform. The target manipulated object was obtained as a textured model file through 3D reconstruction. For objects with movable structures, the model was divided into reasonable blocks using Blender software, and joint constraints were added between the components of the object in the Isaac Sim platform to form object simulation assets with realistic mechanical structures. For objects without joints or block requirements, the whole model could be used directly. The simulation assets were then exported to an Isaac Lab callable format.

[0032] (2) Environment setup: The simulation environment includes a robotic arm and gripper, target object, wrist camera, external fixed camera, background model and lighting system.

[0033] S1.2, Expert control strategy design.

[0034] In this embodiment, privileged information such as the end-effector pose and object pose provided by the simulator (Isaac Lab platform) is used to construct an expert control strategy to control the robot arm to complete contact-intensive tasks in a simulation environment. This expert control strategy divides the contact-intensive task into two phases: a free movement phase when the robot is not in contact with the object, and a contact interaction phase after contact occurs.

[0035] During the free motion phase, a pre-contact key pose (i.e., a reference pose where the robotic arm's end effector maintains a preset safe distance from the object and has not yet made contact) and a contact key pose (i.e., a reference pose where the robotic arm's end effector is just making contact with the object's surface or is about to begin interaction) centered on the object are first defined. Using the object pose information provided by the simulator, these key poses are transformed into the robotic arm's base coordinate system. The robotic arm performs linear interpolation from the current end effector pose to the pre-contact pose, issues end effector pose commands in a time sequence, and then interpolates from the pre-contact pose to the contact pose. For tasks requiring grasping, a gripper closing command is issued after reaching the contact pose.

[0036] During the contact interaction phase, the end-effector pose trajectory is generated based on the task geometry and mechanical constraints, as shown in the following example: (1) Turn on the microwave oven: With the door hinge of the microwave oven as the center of rotation, generate a circular trajectory of the contact point around the hinge; (2) Insertion hole: Generate a straight insertion trajectory along the hole axis to ensure that the insertion axis is collinear with the hole axis; (3) Erasing the whiteboard: The eraser plane is kept parallel to the whiteboard plane, and a horizontal covering trajectory is generated in the whiteboard plane; (4) Opening the door: In the first stage, a circular trajectory is generated around the door handle pivot to complete the pressing action; in the second stage, a circular trajectory is generated around the door pivot to complete the opening action.

[0037] S1.3, Simulation demonstration data acquisition.

[0038] In this embodiment, the task completion status is automatically determined based on the privileged status and task success indicators provided by the simulator, and only the successful trajectory data is saved, thereby collecting large-scale and diverse demonstration data at low cost to support subsequent policy training.

[0039] Task success metrics include information such as the maximum completion time and the geometric or physical conditions required for task completion, as shown in the example below: (1) Microwave oven: The door opens to an angle of more than 50° within 120 seconds; (2) Insertion hole: Insertion depth exceeds 10 mm within 60 seconds; (3) Erasure of whiteboard: All writing is removed within 120 seconds; (4) Opening the door: The opening angle exceeds 30° within 120 seconds.

[0040] Randomize the following parameters each time the environment is reset: object pose, initial pose of the robotic arm end effector, intrinsic and extrinsic parameters of the wrist camera and external camera, and lighting and background parameters.

[0041] Each demonstration data record captures the following at each moment: wrist camera RGB image, external camera RGB image, robotic arm end-effector pose, binary gripper state, binary gripper control commands, and binary contact state (indicating whether the object grasped by the robotic arm end-effector or gripper is in contact with the target object; the criterion for contact is whether the distance is less than a threshold). During the contact interaction phase, the surface normal of the object at the point of contact between the robotic arm end-effector and the object is recorded as the true value of the expected normal contact force direction: when no contact occurs, it is recorded as a 3D zero vector; when contact occurs, it is recorded as the surface normal of the object at the contact point, which is a unit 3D vector. For example, the expected normal contact force direction for a socket is upward along the vertical direction of the socket, and the expected normal contact force direction for wiping a whiteboard is upward perpendicular to the whiteboard.

[0042] S2, based on the demonstration data, performs imitation learning training on the policy network, and outputs the end pose, expected normal contact force direction and contact state for the next several steps.

[0043] S2.1, Policy Network Construction.

[0044] In this embodiment, the policy network includes a visual encoding module, a language encoding module, a conditional fusion module, and a diffusion policy module.

[0045] (1) Visual encoding module: The wrist RGB image is input into the DINOv2 model and the SigLIP model respectively to obtain feature vectors, and then the feature vectors are added element by element after being mapped by a multilayer perceptron to obtain the wrist visual features. The same operation is performed on the external RGB image to obtain the external visual features. The two sets of visual features are concatenated and self-attention is calculated to obtain the visual features.

[0046] (2) Language encoding module: Input the task language description into the SigLIP model, extract semantic features and then map them through a multilayer perceptron to obtain language features.

[0047] (3) Conditional Fusion Module: Using visual features as queries and linguistic features as keys and values, cross-attention calculation is performed. The calculation result is processed by a feedforward network to obtain new visual features. The linguistic features are concatenated with a set of learnable features and then self-attention calculation is performed. The part of the self-attention output corresponding to the learnable features is used as the query, and the aforementioned new visual features are used as keys and values. Cross-attention calculation is performed again, and the calculation result is processed by a feedforward network. At the same time, the part of the self-attention output corresponding to the linguistic features is sent to another feedforward network. The outputs of the two feedforward networks are concatenated to form conditional features.

[0048] (4) Diffusion Strategy Module: A Transformer-based diffusion strategy structure is adopted. The Denoising Diffusion Probability Model (DDPM) is used in the training phase, and the Denoising Diffusion Implicit Model (DDIM) is used in the inference phase. Specifically, the current end-effector pose and the noisy future action sequence are respectively input into the multilayer perceptron for encoding. The two sets of encoded features are concatenated and self-attention calculation is performed. The denoising time step is injected through the adaptive layer normalization zero initialization (AdaLN-Zero). The self-attention output is used as the query, and the conditional features are used as the key and value to perform cross-attention calculation. The cross-attention result is input into the feedforward network and injected into the denoising time step again through the AdaLN-Zero method. The output of the feedforward network is mapped through the multilayer perceptron to obtain the noise prediction result.

[0049] S2.2, Policy Network Training.

[0050] (1) Divide the demonstration data into (observation, action sequence) data pairs. Observations include wrist camera RGB images, external camera RGB images, robotic arm end-effector pose, and task description. Each action in the action sequence includes: desired robotic arm end-effector pose, binary gripper control command, binary contact state, and desired normal contact force direction; (2) Randomly sample batch data from the demonstration data, uniformly sample time steps, and sample Gaussian noise; (3) Add noise to the action sequence according to the forward diffusion process; (4) Input the observation, time step and noisy action sequence into the policy network, and output the predicted noise after forward propagation of the network; (5) Calculate the mean squared error loss between the predicted noise and the actual added noise, and update the policy network parameters through gradient backpropagation so that the policy network learns the conditional distribution from the observation to the action sequence.

[0051] S3 deploys the trained policy network onto the real robot. When no contact is predicted, the adaptive compliant controller generates low-stiffness admittance control signals in all directions to track the predicted pose. When contact is predicted, the adaptive compliant controller generates damping control signals in the direction of the desired normal contact force to maintain the set contact force, generates high-stiffness admittance control signals in the tangential direction determined by the difference between the current pose and the predicted pose to overcome resistance, and generates low-stiffness admittance control signals in the other directions. The robot performs operations according to the admittance control signals.

[0052] S3.1, Adaptive Compliant Controller Design.

[0053] In this embodiment, the dynamic equation of the adaptive compliant controller is as follows: , in, , and These are the admittance reference pose, velocity, and acceleration, respectively (which are the unknowns in the dynamic equations). (These are the instructions ultimately sent to the underlying controller). The pose of the robotic arm's end effector predicted by the policy network. This is an expression representing the external contact force acting on the end effector of the robotic arm in the robotic arm's base coordinate system, estimated based on readings from the joint torque sensors. The expected external force (which will be adjusted in real time based on the strategy network prediction and rules). , , The mass, damping, and stiffness coefficients of the admittance controller are set to default values ​​manually and adjusted in real time based on strategy network predictions and rules.

[0054] (1) When the binary contact state predicted by the policy network is 0 (no contact), set The adaptive compliant controller exhibits low stiffness admittance control, with a default low stiffness coefficient set in all directions (preferably in the range of 50~100, 50 is used in the embodiment) to achieve low stiffness admittance control in all directions, so as to compliantly and safely track the desired end-effector pose predicted by the strategy network.

[0055] (2) When the binary contact state predicted by the policy network is 1 (contact): i, in the direction of the expected normal contact force predicted by the policy network. Up, settings The expected normal contact force amplitude The controller exhibits damping control in this direction: , in, The magnitude of the desired normal contact force is set artificially. For external forces In direction The amount on, for In direction scalar value on and In direction The admittance reference velocity and acceleration are measured. The controller will drive the robotic arm to make contact with the object and ensure that the contact force is at the set value. This ensures stable and safe contact between the robotic arm and the object in that direction.

[0056] ii. Simultaneously, calculate the desired tangential contact force direction. : , in, superscript For transpose, It is the identity matrix. In the direction of the desired tangential contact force... The stiffness becomes several times the default value (preferably 4 to 8 times, 4 times in the embodiment), and the damping is calculated according to the fixed damping ratio formula, so that it exhibits high stiffness admittance control in this direction to overcome the drag.

[0057] iii. In the other directions, low stiffness admittance control is maintained to ultimately achieve robustness against errors and disturbances.

[0058] S3.2, Reasoning in a Real-World Context.

[0059] (1) Set the parameter of the desired normal contact force. ; (2) Collect the RGB image of the wrist camera, the RGB image of the external camera, the pose of the robotic arm end effector and the task language description at the current moment, as the observation input of the policy network; (3) Randomly sample Gaussian noise and construct a pure noise sequence with the same dimension as the target action sequence as the initial action sequence; (4) Input the observed input, the current denoised time step and the noisy action sequence into the trained policy network. The network forward propagates and outputs the predicted noise. According to the DDIM sampling formula, the predicted noise is used to denoise and update the current action sequence to obtain the action sequence of the next time step. (5) Repeat step (4) until all denoising time steps are completed to obtain the final denoising action sequence, which includes the expected end pose of the robotic arm, the binary gripper control command, the binary contact state and the expected normal contact force direction for the next 32 time steps.

[0060] (6) The generated action sequence is sent to the adaptive compliant controller at a fixed frequency to drive the robot to perform the operation.

[0061] To further verify the effectiveness of the present invention, the following comparative experiments were conducted.

[0062] The baseline methods compared to this method are the visual-language-action (VLA) models Pi0 and E2VLA trained with the same data, and the six baseline models constructed using omnidirectional low / medium / high stiffness admittance controllers (L / M / HA): Pi0+LA, Pi0+MA, Pi0+HA, E2VLA+LA, E2VLA+MA, and E2VLA+HA.

[0063] Physical experiments were conducted on four contact-intensive tasks: opening a microwave oven, plugging in an electrical outlet, wiping a whiteboard, and opening a door. Each model underwent 20 tests without disturbance and 5 tests with disturbance. Each test randomly assigned the initial pose of the robotic arm's end effector and the pose of the target object. The disturbance applied in the experiments referred to as: (1) Opening the microwave oven: After the robotic arm grabs the microwave oven door handle, manually shake the microwave oven; (2) Insertion hole: When the robotic arm has inserted the column part into the hole, the position of the hole is manually moved; (3) Erasing the whiteboard: After the robotic arm holding the eraser has made contact with the whiteboard, manually lower or raise the height of the whiteboard, or change the tilt angle of the whiteboard. (4) Opening the door: After the robotic arm grabs the door handle, the door is manually shaken.

[0064] Table 1 shows the success rate and failure modes under undisturbed conditions. The results show that, compared with the best-performing baseline method, the present invention improves the success rate by 19% in four contact-intensive operation tasks: opening a microwave oven, plugging in a socket, wiping a whiteboard, and opening a door. At the same time, it reduces the number of failures caused by insufficient or excessive contact force to 0, achieving stable and safe contact operation.

[0065] As shown in Table 2, the success rate and failure modes under disturbance conditions are statistically analyzed. The results show that the present invention improves the disturbance robustness success rate by 20% compared with the best existing baseline method in four contact-intensive operation tasks: opening a microwave oven, plugging in a socket, wiping a whiteboard, and opening a door.

[0066] Table 1. Statistics on success rate and failure modes under undisturbed conditions

[0067] Table 2. Statistics on success rate and failure modes under perturbation.

[0068] In summary, the robot contact-intensive operation simulation-to-real-world transfer method based on contact force direction prediction provided by this invention uses pure simulation data to train the policy network, requiring no input or output dynamics-related variables. Therefore, it does not exhibit significant degradation when transferred to a real-world environment and can make reliable predictions based on real-world observations. Furthermore, this invention utilizes the contact state and normal direction predicted by the policy, combined with a manually set desired normal contact force amplitude, to achieve adaptive adjustment of the adaptive compliant controller parameters. This allows the controller to exhibit differentiated stiffness and damping characteristics in different directions, maintaining stable contact between the robotic arm and the object while providing the driving force needed to overcome resistance. It also exhibits robustness to robotic arm end-effector pose prediction errors and external disturbances, thereby supporting stable, safe, and efficient control of real robots in contact-intensive operation tasks.

[0069] Based on the same inventive concept, such as Figure 3 As shown, this embodiment of the invention also provides a robot contact-intensive operation simulation-to-reality transfer system 300 based on contact force direction prediction, including: a simulation data construction module 310, a policy network training module 320, and a reality transfer deployment module 330.

[0070] The simulation data construction module 310 is used to construct expert control strategies in a simulation environment, divide the operation task into a free motion stage and a contact interaction stage, generate and execute pose trajectories to collect demonstration data, and record the surface normal of the object at the contact point as the true value of the expected normal contact force direction in the contact interaction stage. The policy network training module 320 is used to train the policy network by imitation learning based on demonstration data, and outputs the end pose, expected normal contact force direction and contact state for future multiple steps. The real-world transfer deployment module 330 is used to deploy the trained policy network onto the real robot. When no contact is predicted, the adaptive compliant controller generates low-stiffness admittance control signals in all directions to track the predicted pose. When contact is predicted, the adaptive compliant controller generates damping control signals in the desired normal contact force direction to maintain the set contact force, generates high-stiffness admittance control signals in the tangential direction determined by the difference between the current pose and the predicted pose to overcome resistance, and generates low-stiffness admittance control signals in the other directions. The robot performs operations according to the admittance control signals.

[0071] Based on the same inventive concept, embodiments of the present invention also provide an electronic device, including a memory and one or more processors, wherein the memory is used to store a computer program, and the processor is used to implement the above-described method for transferring simulation of robot contact-intensive operations based on contact force direction prediction to reality when the computer program is executed.

[0072] Based on the same inventive concept, embodiments of the present invention also provide a computer-readable storage medium storing a computer program. When the computer program is executed by a computer, the above-described method for transferring simulation of robot contact-intensive operations based on contact force direction prediction to reality is realized.

[0073] It should be noted that the robot contact-intensive operation simulation to reality transfer system, electronic device and computer-readable storage medium based on contact force direction prediction provided in the above embodiments all belong to the same inventive concept as the robot contact-intensive operation simulation to reality transfer method based on contact force direction prediction. For details of the specific implementation process, please refer to the embodiments of the robot contact-intensive operation simulation to reality transfer method based on contact force direction prediction, which will not be repeated here.

[0074] The specific embodiments described above illustrate the technical solution and beneficial effects of the present invention in detail. It should be understood that the above description is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A robot contact-intensive operation simulation-to-reality migration method based on contact force direction prediction, characterized in that, Includes the following steps: In a simulation environment, an expert control strategy is constructed, and the operation task is divided into a free motion stage and a contact interaction stage. Pose trajectories are generated and executed to collect demonstration data. In the contact interaction stage, the surface normal of the object at the contact point is recorded as the true value of the direction of the expected normal contact force. The policy network is trained by imitation learning based on the demonstration data, and the end pose, expected normal contact force direction and contact state of the next several steps are output. The trained policy network is deployed on a real robot. When no contact is predicted, an adaptive compliant controller generates low-stiffness admittance control signals in all directions to track the predicted pose. When contact is predicted, the adaptive compliant controller generates damping control signals in the direction of the desired normal contact force to maintain the set contact force, generates high-stiffness admittance control signals in the tangential direction determined by the difference between the current pose and the predicted pose to overcome resistance, and generates low-stiffness admittance control signals in the other directions. The robot performs operations according to the admittance control signals.

2. The robot contact-intensive operation simulation-to-reality migration method based on contact force direction prediction according to claim 1, characterized in that, The expert control strategy generates a pose trajectory from the current end-effector pose to the key pose centered on the object through linear interpolation during the free motion phase and executes it sequentially. During the contact interaction phase, it generates an end-effector pose trajectory that satisfies the geometric and mechanical constraints of the task and executes it sequentially according to the rules of each operation task. 3.The robot contact-intensive operation simulation-to-reality migration method based on contact force direction prediction of claim 1, wherein, During the demonstration data collection process, the lighting, background, object pose, initial end-effector pose of the robotic arm, camera intrinsic and extrinsic parameters are randomized each time the simulation environment is reset. The task completion status is automatically determined based on the task success indicators provided by the simulator, and only successful demonstration data is retained.

4. The method for transferring simulation to reality of robot contact-intensive operations based on contact force direction prediction according to claim 1, characterized in that, The input to the policy network includes the observed image at the current moment, the pose of the robotic arm end effector, and the task language description. The output includes the robotic arm end effector pose, the desired normal contact force direction, and the action sequence of binary contact state for multiple future time steps.

5. The method for transferring simulation to reality of robot contact-intensive operations based on contact force direction prediction according to claim 4, characterized in that, The policy network includes a visual encoding module, a language encoding module, a conditional fusion module, and a diffusion strategy module. The visual encoding module encodes the observed image into visual features, the language encoding module encodes the task language description into language features, the conditional fusion module fuses the visual and language features into conditional features through a cross-attention mechanism, and the diffusion strategy module encodes the current robotic arm end-effector pose and the noisy future action sequence, performs cross-attention calculation with the conditional features, and combines the denoised time step to perform noise prediction to generate the future action sequence.

6. The method for transferring simulation to reality of robot contact-intensive operations based on contact force direction prediction according to claim 1, characterized in that, The dynamic equation of the adaptive compliant controller is: , in, and These are the admittance reference velocity and acceleration, respectively. For admittance reference pose, The robot arm's end-effector pose predicted by the policy network. For external contact force, In order to expect external forces, , , These are mass, damping, and stiffness coefficients, respectively.

7. The method for transferring simulation to reality of robot contact-intensive operations based on contact force direction prediction according to claim 6, characterized in that, When the policy network predicts that the robot will not make contact with the object, set... Furthermore, with all directions set to the default low stiffness coefficient, the calculated admittance reference pose is... This constitutes a low-stiffness admittance control signal in all directions; When the policy network predicts that the robot will make contact with an object, the contact force will be in the direction of the expected normal direction. Top settings ,in For the desired normal contact force amplitude, the calculated admittance reference pose constitutes the damping control signal in that direction. In the tangential direction determined by the difference between the current pose and the predicted pose, the stiffness coefficient is increased to several times the default low stiffness coefficient to overcome the drag. The calculated admittance reference pose constitutes the high stiffness admittance control signal in that direction. In other directions, the default low stiffness coefficient is maintained, and the calculated admittance reference pose constitutes the low stiffness admittance control signal in these directions.

8. A robot contact-intensive operation simulation-to-real-world transfer system based on contact force direction prediction, implemented using the method described in any one of claims 1 to 7, characterized in that, include: Simulation data construction module: Constructs expert control strategies in the simulation environment, divides the operation task into free motion stage and contact interaction stage, generates and executes pose trajectories to collect demonstration data, and records the surface normal of the object at the contact point as the true value of the expected normal contact force direction in the contact interaction stage. Policy network training module: Based on demonstration data, the policy network is trained by imitation learning, and the end pose, expected normal contact force direction and contact state of the next several steps are output. Real-world deployment module: The trained policy network is deployed on a real robot. When no contact is predicted, an adaptive compliant controller generates low-stiffness admittance control signals in all directions to track the predicted pose. When contact is predicted, the adaptive compliant controller generates damping control signals in the desired normal contact force direction to maintain the set contact force, and generates high-stiffness admittance control signals in the tangential direction determined by the difference between the current pose and the predicted pose to overcome resistance. Low-stiffness admittance control signals are generated in other directions. The robot performs operations according to the admittance control signals.

9. An electronic device comprising a memory and one or more processors, the memory for storing a computer program, characterized in that, The processor is used to implement, when executing a computer program, the simulation-to-reality transfer method for robot contact-intensive operations based on contact force direction prediction as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by a computer, it implements the method for transferring simulation of robot contact-intensive operations to reality based on contact force direction prediction as described in any one of claims 1 to 7.