Visual-haptic control policy decoupling method and system based on physical alignment reward
By constructing a multimodal feature representation based on 3D point clouds and a high-frequency tactile residual strategy, combined with physical alignment reward function optimization, the inefficiency and hesitant behavior of the visual-tactile strategy in fine-grained physical interaction are solved, enabling the robot to operate efficiently and safely in contact-intensive tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN INST OF TECH
- Filing Date
- 2026-06-11
- Publication Date
- 2026-07-14
AI Technical Summary
Existing diffusion-based visual-tactile strategies exhibit inefficiencies at the control level and hesitant behavior in risk-avoidance during fine-grained physical interactions. Sparse rewards lack progressive feedback on contact stability or force modulation, leading to longer execution times and increased failure risks for robots in contact-intensive tasks.
By constructing a unified multimodal feature representation based on 3D point clouds, long-view action blocks are generated and closed-loop dynamic correction is performed using a high-frequency tactile residual strategy. A dense reward function based on physical alignment is designed for reinforcement learning optimization. The visual-tactile control strategy is decoupled by combining low-frequency basic strategy and high-frequency residual strategy.
It significantly improves the efficiency and safety of robots in contact-intensive operations, enhances dynamic response capabilities, reduces execution time and failure risk, and strengthens robustness against positional uncertainties.
Smart Images

Figure CN122378745A_ABST
Abstract
Description
Technical Field
[0001] This invention pertains to robot control and artificial intelligence technologies, specifically relating to a decoupling method and system for visual-tactile control strategies based on physical alignment rewards. Background Technology
[0002] Robotic manipulation in physical environments requires a close integration of perception and control, especially in contact-intensive tasks. While existing diffusion-based visual-tactile strategies have achieved significant success in long-field-of-view tasks, they often exhibit inefficiencies at the control level and hesitant behavior in risk-avoidance during fine-grained physical interactions.
[0003] Existing methods model a series of control actions as a single generated object, resulting in open-loop control within the block execution window. Consequently, high-frequency tactile feedback arriving within the execution window cannot be incorporated step-wise to update control commands, limiting the robot's responsiveness to rapid contact dynamics.
[0004] Existing reinforcement learning training typically utilizes sparse rewards that evaluate the performance of binary tasks. This sparse reward lacks stepwise feedback on contact stability or force modulation, causing the policy to converge to a suboptimal, risk-averse, conservative strategy, which not only prolongs execution time but also increases the risk of contact failure. Summary of the Invention
[0005] This invention provides a decoupling method and system for visual-tactile control strategies based on physical alignment rewards, which is used to solve the problem of how to efficiently and safely combine visual and high-frequency tactile feedback for fine physical interaction in robots during contact-intensive operations (such as high-precision dual-arm precision assembly, plug-in and unplug tasks).
[0006] This invention is achieved through the following technical solution: A decoupling method for visual-haptic control strategies based on physical alignment rewards, the decoupling method comprising the following steps: Step 1: Construct a unified multimodal feature representation based on 3D point clouds; Step 2: Based on the multimodal feature table from Step 1, generate long-view action blocks using a low-frequency basic strategy; Step 3: Based on the long-field action block from Step 2, use a high-frequency tactile residual strategy for closed-loop dynamic correction; Step 4: Based on the closed-loop dynamic correction in Step 3, perform reinforcement learning optimization based on physical alignment dense rewards; Step 5: Based on the optimized reinforcement learning from Step 4, perform two-stage frequency perception training to achieve decoupling of visual-tactile control strategies.
[0007] Furthermore, step 1 specifically involves unifying the multimodal sensor signals from different dimensions into a shared 3D coordinate system, and using the visual geometric point cloud acquired by the RGB-D camera. High-resolution tactile point clouds, i.e., high-frequency tactile cues, are acquired by fingertip tactile sensors and mapped into 3D space. And the position of the arm joints encoding the kinematic state, i.e., the proprioceptive state. By extracting a unified feature representation: This is for subsequent network calls to strategies.
[0008] Furthermore, step 2 specifically involves constructing a low-frequency basic strategy parameterized as a conditional diffusion model in order to establish a global geometric prior and maintain the consistency of long-term planning. ; Step 2-1: The basic strategy receives low-frequency global basic observation information; Step 2-2: Learn the conditional distribution by minimizing the noise prediction error to generate a reference action block containing multiple consecutive steps. This stage provides a globally stable motion trajectory reference throughout the entire execution window of the action block.
[0009] Furthermore, step 3 specifically involves constructing a lightweight multilayer perceptron as a high-frequency residual strategy to compensate for the open-loop execution defects of action blocks. ; Step 3-1: Bypassing high-latency visual processing, this residual strategy focuses only on transient, high-frequency tactile cues. and proprioceptive state ; Step 3-2: Residual strategy freezes the basic action blocks in the background. and in-block time embedding Given the condition, output progressively corrective actions. ; Step 3-3: The final control command is formed by superimposing the predicted action of the basic policy and the response correction of the residual policy:
[0010] in It serves as a scaling factor, thereby actively adjusting the contact dynamics without compromising the global semantic intent.
[0011] Furthermore, step 4 specifically involves designing a physically aligned dense reward function to replace sparse rewards in reinforcement learning, thereby enabling the model to internalize stable control dynamics; the reward function The formula is:
[0012] in, For proportional weighting, For integral weights, For differential weights, To track rewards, As a penalty for oscillation damping, As a penalty for steady-state error, A reward will be given for success.
[0013] Further, step 4-1: Force tracking reward Measuring the tactile ability of a robot's fingertips With reference to the target The alignment between them is calculated using the Gaussian kernel function; Step 4-2: Steady-state error penalty The time integral of the penalty sliding window internal force deviation and task incomplete bias forces the strategy to actively address persistent contact errors and prevent the robot from stagnating and hesitating near the target. Step 4-3: Oscillation Damping Penalty At the moment of contact, the hinge loss function is used to penalize excessive speed of the end effector. It acts as an active damper to suppress oscillations caused by unstable contact.
[0014] Furthermore, step 5 specifically involves: Basic policy training: First, the basic policy is initialized on real-world demonstration data through behavior cloning, and then transferred to the simulation environment to fine-tune it using the diffusion policy optimization algorithm to adapt it to the dynamics of the environment; Residual policy training: Freeze the parameters of the basic policy, use the PPO algorithm and combine it with the aforementioned "physical alignment-based dense reward" to train the high-frequency residual policy, so that it focuses on learning the stepwise corrective actions of stable contact interactions.
[0015] A decoupling system for visual-haptic control strategies based on physical alignment rewards, the decoupling system employing the aforementioned decoupling method for visual-haptic control strategies based on physical alignment rewards, the decoupling system comprising... Modal feature construction module: Constructs a unified multimodal feature representation based on 3D point clouds; Long-field action block generation module: Based on the multimodal feature table constructed by the modal feature construction module, long-field action blocks are generated using a low-frequency basic strategy; Closed-loop dynamic correction module: Based on the long field of view action block generated by the long field of view action block generation module, a high-frequency tactile residual strategy is used for closed-loop dynamic correction; Reinforcement learning optimization module: Based on the closed-loop dynamic correction module, reinforcement learning optimization based on physical alignment dense rewards is performed. Perception Training Module: Based on the optimized reinforcement learning optimization module, two-stage frequency perception training is carried out to achieve decoupling of visual and tactile control strategies.
[0016] An application of the above-mentioned decoupling method for visual-tactile control strategy based on physical alignment reward is characterized in that it enables robots in physical environments to efficiently and safely combine visual and high-frequency tactile feedback for fine physical interaction in contact-intensive operations.
[0017] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method described above.
[0018] The beneficial effects of this invention are: This invention significantly improves execution efficiency.
[0019] This invention can significantly improve the security and compliance of interactions.
[0020] This invention has extremely strong robustness against positional uncertainty. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the structure of the present invention.
[0022] Figure 2 This is a schematic diagram of the collaborative control architecture of the low-frequency basic strategy and the high-frequency residual strategy of the present invention.
[0023] Figure 3 This is a closed-loop dynamic correction physical characteristic analysis diagram of the present invention, wherein (a) is the internal residual modulation, (b) is the end effector speed curve, and (c) is the contact force dynamics. Detailed Implementation
[0024] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary detail.
[0025] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0026] It should also be understood that the terminology used in this application specification is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this application specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0027] The following is in conjunction with the appendix to this application specification. Figure 1-3 The technical solutions in the embodiments of this application are clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0028] Many specific details are set forth in the following description in order to provide a full understanding of this application. However, this application may also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.
[0029] Implementation Method 1 This embodiment provides a decoupling method for visual-haptic control strategies based on physical alignment rewards. The decoupling method includes the following steps: Step 1: Construct a unified multimodal feature representation based on 3D point clouds; Step 2: Based on the multimodal feature table from Step 1, generate long-view action blocks using the low-frequency base policy; Step 3: Based on the long-field action block from Step 2, use a high-frequency tactile residual policy for closed-loop dynamic correction; Step 4: Based on the closed-loop dynamic correction in Step 3, perform reinforcement learning optimization based on Physics-Aligned Dense Reward; Step 5: Based on the optimized reinforcement learning in Step 4, perform two-stage frequency-aware training to decouple the visual-tactile control strategy.
[0030] Furthermore, step 1 specifically involves unifying the multimodal sensor signals from different dimensions into a shared 3D coordinate system, and using the visual geometric point cloud acquired by the RGB-D camera. High-resolution tactile point clouds, i.e., high-frequency tactile cues, are acquired by fingertip tactile sensors and mapped into 3D space. (Including 3D position and force magnitude), and the position of the arm joints encoding the kinematic state, i.e., the proprioceptive state. By extracting a unified feature representation: This is for subsequent network calls to strategies.
[0031] Furthermore, step 2 specifically involves constructing a low-frequency basic policy parameterized as a Conditional Diffusion Model (DDPM) in order to establish a global geometric prior and maintain the consistency of long-term planning. ; Step 2-1: The basic strategy receives low-frequency global basic observation information, including visual point clouds, tactile point clouds, and proprioception; Step 2-2: Learn the conditional distribution by minimizing the noise prediction error to generate a reference action block containing multiple consecutive steps. This stage provides a globally stable motion trajectory reference throughout the entire execution window of the action block.
[0032] Furthermore, step 3 specifically involves constructing a lightweight multilayer perceptron (MLP) as a high-frequency residual strategy to compensate for the open-loop execution defects of action blocks. ; Step 3-1: Bypassing high-latency visual processing, this residual strategy focuses only on transient, high-frequency tactile cues. and proprioceptive state ; Step 3-2: Residual strategy freezes the basic action blocks in the background. and in-block time embedding Given the condition, output progressively corrective actions. ; Step 3-3: Action Synthesis: The final control command is formed by superimposing the predicted action of the basic policy and the response correction of the residual policy.
[0033] in It serves as a scaling factor, thereby actively adjusting the contact dynamics without compromising the global semantic intent.
[0034] Furthermore, step 4 specifically involves designing a physically aligned dense reward function to replace sparse rewards in reinforcement learning, thereby enabling the model to internalize stable control dynamics; the reward function The formula is:
[0035] in, For proportional weighting, For integral weights, For differential weights, To track rewards, As a penalty for oscillation damping, As a penalty for steady-state error, A reward will be given for success.
[0036] This is used to amplify the impact of force tracking rewards, prompting the model to actively follow the target in a dynamic process; This is used to amplify the effect of steady-state error penalty, prompting the model to eliminate persistent small biases in long-term operation; This is used to amplify the effect of oscillation damping penalty, prompting the model to suppress drastic fluctuations in actions or states.
[0037] Further, step 4-1: Force tracking reward Measuring the tactile ability of a robot's fingertips With reference to the target The alignment between them is calculated using a Gaussian kernel function, which encourages the network to learn compliant contact behavior similar to stiffness; Step 4-2: Steady-state error penalty The time integral of the penalty sliding window internal force deviation and task incomplete bias forces the strategy to actively address persistent contact errors and prevent the robot from stagnating and hesitating near the target. Step 4-3: Oscillation Damping Penalty At the moment of contact, the hinge loss function is used to penalize excessive velocity of the end effector. It acts as an active damper to suppress oscillations caused by unstable contact.
[0038] Furthermore, step 5 specifically involves: Basic policy training: First, the basic policy is initialized on real-world demonstration data using behavior cloning (BC), and then transferred to the simulation environment for fine-tuning using the diffusion policy optimization (DPPO) algorithm to adapt it to the dynamics of the environment; Residual policy training: Freeze the parameters of the basic policy, use the PPO algorithm and combine it with the aforementioned "physical alignment-based dense reward" to specifically train the high-frequency residual policy, so that it focuses on learning the stepwise corrective actions of stable contact interactions.
[0039] In a precision plug-in assembly task (such as the five high-precision component assembly tasks in the AutoMate dataset) equipped with a dual-arm ALOHA system and fingertip TacSL tactile sensors.
[0040] When the parameters are at their optimal values (e.g., the action block size is set to 16 steps, the haptic update interval frequency is set to the highest frequency of k=1, and the P, I, and D terms in the reward function are activated simultaneously): The robot utilizes visual priors for rapid global approach. When it reaches a dense contact area a few centimeters from the target, the high-frequency tactile residual strategy performs subtle calibrations on the gripper position at each step. If angular misalignment or jamming occurs during insertion into the hole, conventional algorithms would continuously apply downward force, causing jamming. However, the algorithm of this invention, by detecting abnormal high-frequency tactile force, immediately triggers an active damping mechanism (D), rapidly slowing down and smoothly adjusting the end effector's posture, stabilizing the operational contact force within an extremely low amplitude range. Ultimately, this not only achieves an average assembly success rate of 89%, but also makes the entire assembly process approximately 30% faster than the current best technology.
[0041] Table 1 shows the results of the bi-arm operation.
[0042] Compared to the existing advanced visual-tactile baseline method VT-Refine, this invention eliminates the robot's "start-stop" hesitation behavior during the contact phase, reducing the task execution time by 12% while maintaining a high success rate, and lowering the average number of execution steps to about 151 steps.
[0043] Thanks to the differential term damping correction capability of the high-frequency residual strategy, this invention can actively dissipate contact impact energy, reducing the peak interaction force by 70%, thus meeting the stringent safety requirements of precision manufacturing for non-destructive interactions.
[0044] When there is a random spatial disturbance of up to 1.4 cm at the initial position, the success rate of traditional open-loop block control drops significantly to 58.2%, while the present invention, with its real-time pose correction capability of high-frequency proportional terms, still maintains a robust success rate of 69.8%.
[0045] Implementation Method 2 This embodiment provides a visual-haptic control strategy decoupling system based on physical alignment rewards. The decoupling system uses a visual-haptic control strategy decoupling method based on physical alignment rewards as described in Embodiment 1. The decoupling system includes: Modal feature construction module: Constructs a unified multimodal feature representation based on 3D point clouds; Long-field action block generation module: Based on the multimodal feature table constructed by the modal feature construction module, long-field action blocks are generated using the low-frequency base policy; Closed-loop dynamic correction module: Based on the long-field action block generated by the long-field action block generation module, closed-loop dynamic correction is performed using a high-frequency tactile residual policy. Reinforcement learning optimization module: Based on the closed-loop dynamic correction module, reinforcement learning optimization based on Physics-Aligned Dense Reward is performed. Perception Training Module: Based on the optimized reinforcement learning optimization module, two-stage frequency-aware training is performed to achieve decoupling of visual and tactile control strategies.
[0046] Implementation Method 3 This embodiment provides an application of a visual-tactile control strategy decoupling method based on physical alignment reward as described in Embodiment 1. It is applied to robots in physical environments to efficiently and safely combine visual and high-frequency tactile feedback for fine physical interaction in contact-intensive operations.
[0047] Implementation Method 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method described in Embodiment 1.
Claims
1. A method for decoupling visual-haptic control strategies based on physical alignment rewards, characterized in that, The decoupling method includes the following steps: Step 1: Construct a unified multimodal feature representation based on 3D point clouds; Step 2: Based on the multimodal feature table from Step 1, generate long-view action blocks using a low-frequency basic strategy; Specifically, step 2 involves constructing a low-frequency basic strategy parameterized as a conditional diffusion model in order to establish a global geometric prior and maintain the consistency of long-term planning. ; Step 2-1: The basic strategy receives low-frequency global basic observation information; Step 2-2: Learn the conditional distribution by minimizing the noise prediction error to generate a reference action block containing multiple consecutive steps. This phase provides a globally stable motion trajectory reference throughout the entire execution window of the action block. Step 3: Based on the long-field action block from Step 2, use a high-frequency tactile residual strategy for closed-loop dynamic correction; Step 3 specifically involves constructing a lightweight multilayer perceptron as a high-frequency residual strategy to compensate for the open-loop execution defects of action blocks. ; Step 3-1: Bypassing high-latency visual processing, this residual strategy focuses only on transient, high-frequency tactile cues. and proprioceptive state ; Step 3-2: Residual strategy freezes the basic action blocks in the background. and in-block time embedding Given the condition, output progressively corrective actions. ; Step 3-3: The final control command is formed by superimposing the predicted action of the basic policy and the response correction of the residual policy: in As a scaling factor, it actively adjusts the contact dynamics without compromising the global semantic intent; Step 4: Based on the closed-loop dynamic correction in Step 3, perform reinforcement learning optimization based on physical alignment dense rewards; Step 5: Based on the optimized reinforcement learning from Step 4, perform two-stage frequency perception training to achieve decoupling of visual-tactile control strategies.
2. The decoupling method according to claim 1, characterized in that, Step 1 specifically involves unifying the multimodal sensor signals from different dimensions into a shared 3D coordinate system, and using the visual geometric point cloud acquired by the RGB-D camera. High-resolution tactile point clouds, i.e., high-frequency tactile cues, are acquired by fingertip tactile sensors and mapped into 3D space. And the position of the arm joints encoding the kinematic state, i.e., the proprioceptive state. By extracting a unified feature representation: This is for subsequent network calls to strategies.
3. The decoupling method according to claim 1, characterized in that, Step 4 specifically involves designing a physically aligned dense reward function to replace sparse rewards in reinforcement learning, thereby enabling the model to internalize stable control dynamics; the reward function... The formula is: in, For proportional weighting, For integral weights, For differential weights, To track rewards, As a penalty for oscillation damping, As a penalty for steady-state error, A reward will be given for success.
4. The decoupling method according to claim 3, characterized in that, Step 4-1: Force Tracking Reward Measuring the tactile ability of a robot's fingertips With reference to the target The alignment between them is calculated using the Gaussian kernel function; Step 4-2: Steady-state error penalty The time integral of the penalty sliding window internal force deviation and task incomplete bias forces the strategy to actively address persistent contact errors and prevent the robot from stagnating and hesitating near the target. Step 4-3: Oscillation Damping Penalty At the moment of contact, the hinge loss function is used to penalize excessive speed of the end effector. It acts as an active damper to suppress oscillations caused by unstable contact.
5. The decoupling method according to claim 1, characterized in that, Specifically, step 5 is as follows: Basic policy training: First, the basic policy is initialized on real-world demonstration data through behavior cloning, and then transferred to the simulation environment to fine-tune it using the diffusion policy optimization algorithm to adapt it to the dynamics of the environment; Residual policy training: Freeze the parameters of the basic policy, use the PPO algorithm and combine it with the aforementioned "physical alignment-based dense reward" to train the high-frequency residual policy, so that it focuses on learning the stepwise corrective actions of stable contact interactions.
6. A decoupling system for visual-haptic control strategies based on physical alignment rewards, characterized in that, The decoupling system uses a visual-haptic control strategy decoupling method based on physical alignment rewards as described in any one of claims 1-5, and the decoupling system includes: Modal feature construction module: Constructs a unified multimodal feature representation based on 3D point clouds; Long-field action block generation module: Based on the multimodal feature table constructed by the modal feature construction module, long-field action blocks are generated using a low-frequency basic strategy; Closed-loop dynamic correction module: Based on the long field of view action block generated by the long field of view action block generation module, a high-frequency tactile residual strategy is used for closed-loop dynamic correction; Reinforcement learning optimization module: Based on the closed-loop dynamic correction module, reinforcement learning optimization based on physical alignment dense rewards is performed. Perception Training Module: Based on the optimized reinforcement learning optimization module, two-stage frequency perception training is carried out to achieve decoupling of visual and tactile control strategies.
7. An application of a method for decoupling visual-haptic control strategies based on physical alignment rewards as described in any one of claims 1-5, characterized in that, Robots applied in physical environments can efficiently and safely combine visual and high-frequency tactile feedback for fine physical interaction in contact-intensive operations.
8. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements the method as described in any one of claims 1-5.