Joint denoising method for robot visual motion prediction

By combining visual, tactile, and robot motion data, a joint denoising method solves the problem of insufficient information in robot vision sensors under occlusion conditions, achieves efficient fusion and collaborative generation of multimodal information, and improves the accuracy and robustness of robot operations. It is suitable for scenarios such as home services, industrial assembly, and medical assistance.

CN120707419APending Publication Date: 2025-09-26ROBOTICS RESEARCH CENTER OF YUYAO CITY +1
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510781781.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In existing technologies, robot vision sensors have difficulty accurately judging the state of objects under occlusion. Tactile sensors provide local contact information, but insufficient coordination between modalities leads to complex fine force control and in-hand operations. Existing methods find it difficult to effectively integrate multimodal information, affecting the accuracy and reliability of predictions.

Method used

By fusing the images and depth maps captured by the depth camera, the robot motion data collected by the Piper robotic arm CAN line communication, and the tactile images collected by the Gelsight Mini visual and tactile sensor, a unified generation model is constructed. Variational autoencoders and convolutional neural networks are used for encoding, and the masked self-attention mechanism and Transformer architecture are combined for joint denoising and generation, dynamically adjusting the modal contribution.

Benefits of technology

It achieves efficient fusion of multimodal information, improves the robot's perception ability and prediction accuracy in complex environments, ensures operational stability and precision in scenarios with visual occlusion or fine force control, and is suitable for a variety of scenarios such as home services, industrial assembly, and medical assistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707419A_ABST
    Figure CN120707419A_ABST
Patent Text Reader

Abstract

The invention discloses a joint denoising method for robot visual motion prediction, and the method comprises the steps: constructing a unified generative model through fusing an image and a depth map collected by a depth camera, motion data collected by CAN line communication of a Piper mechanical arm, and a tactile image collected by a Gelsight Mini tactile sensor; the method comprises two steps of data acquisition and input coding, and joint denoising and generation: firstly, multi-modal data are coded into low-dimensional potential representation, and then future images, depth maps, tactile data and robot actions are cooperatively predicted through a joint denoising framework based on Transform. A mask self-attention mechanism is innovatively introduced, information interaction between modes is dynamically adjusted, action generation is guided through tactile feedback, and the force control precision is improved. The model adopts a de-noising diffusion probability loss function to jointly optimize multi-modal prediction, so that the output consistency is ensured. According to the method, the robustness and the accuracy of flexible operation of the robot are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of visual information processing, and in particular relates to a joint denoising method for robot visual action prediction. Background Art

[0002] Currently, the dexterous manipulation of robots in the field of embodied intelligence is limited and affected by:

[0003] 1. Limitations of visual information (especially occlusion). Problem description: Although visual sensors can provide global information about the environment, they often encounter occlusion problems in contact-intensive tasks (such as grasping objects or in-hand manipulation), resulting in the robot being unable to accurately judge the state or contact of objects. For example, when grasping eggs or taking grapes out of a bag, visual information may be blocked and unable to provide sufficient details. The role of touch: Tactile sensors can provide local contact information (such as whether contact exists, the amount of force applied, and the contact pattern), compensating for the shortcomings of vision in occlusion situations, thereby enhancing the robot's ability to perceive the state of objects.

[0004] 2. The need for precise force control. Problem description: When handling fragile objects (such as eggs or grapes) or tasks requiring precise force control, it is difficult to determine the appropriate amount of force applied based on vision alone, which can lead to damage or grasping failure. The role of touch: Through dense tactile sensors (such as the 16×16 array mentioned in the article), the robot can sense the applied pressure in real time and adjust the grasping force based on the feedback, ensuring safe and stable operation.

[0005] 3. Complexity of in-hand manipulation. Problem description: In-hand manipulation (such as adjusting the position of a hex wrench or scooping food with a spoon) requires the robot to perceive the state and position changes of the object in the hand. Visual sensors have difficulty capturing these subtle changes, especially when the object is obscured by the hand. The role of tactile sensing: Tactile sensors provide continuous contact information and local geometric information, helping the robot track the posture of the object in the hand (such as 6-DoF pose estimation), thereby enabling more flexible adjustments and manipulations.

[0006] 4. A more realistic engineering problem is that factories are currently highly automated, but the plugging and unplugging of cables is still not automated and is done by workers. This actually involves the problem of visual information occlusion during the dexterous operation of the robotic arm, the need for fine force control, and the complexity of in-hand operation.

[0007] Even though some studies have attempted to introduce multimodal information, the following problems still exist:

[0008] 1. Insufficient intermodal collaboration: The intermodal collaboration mechanism is not perfect, and it cannot fully utilize the complementarity between different modalities. For example, vision can provide global information, while touch can provide local contact details, but existing methods have difficulty in effectively integrating this information.

[0009] 2. Independence of the Generation Process: Existing methods typically treat the generation processes of different modalities as independent tasks and lack a unified generation framework. This makes it difficult to ensure consistency between different modalities when predicting future states, affecting the accuracy and reliability of the predictions.

[0010] 3. Lack of fine force control capabilities: When handling fragile objects or tasks that require precise force control, existing diffusion strategies have difficulty accurately predicting and controlling the magnitude and direction of contact forces, which can easily lead to operational failure or object damage.

[0011] These limitations severely restrict the application of diffusion strategies in complex robotic manipulation tasks, especially in scenarios requiring fine manipulation and force control. Therefore, a new diffusion strategy framework that can integrate multimodal information and achieve collaborative generation is urgently needed. Summary of the Invention

[0012] The purpose of the present invention is to provide a joint denoising method for robot visual action prediction to solve the above technical problems.

[0013] To solve the above technical problems, the specific technical solution of the joint denoising method for robot visual action prediction of the present invention is as follows:

[0014] A joint denoising method for robot visual motion prediction is proposed. This method fuses images and depth maps acquired by a depth camera, robot motion data collected via CAN communication with a Piper robotic arm, and tactile images collected by a Gelsight Mini visual-tactile sensor to build a unified generative model. The method includes the following steps:

[0015] Step 1: Data collection and input coding;

[0016] Step 2: Joint denoising and generation process.

[0017] Furthermore, the step 1 includes the following steps:

[0018] Step 1.1: Data sources: images and depth maps;

[0019] Step 1.2: Encoding process: image and depth map encoding.

[0020] Furthermore, the step 1.1 includes the following steps:

[0021] The depth camera collects RGB images and corresponding depth information in real time. Robot motion: High-dimensional posture data, including end-effector position, rotation angle, and gripper status, is collected through the CAN line communication protocol of the Piper robotic arm. Tactile image: High-resolution tactile data is collected using the Gelsight Mini visual tactile sensor. Gelsight Mini captures the morphology and force distribution of the contact surface through its gel surface and internal camera to generate high-resolution tactile images.

[0022] Furthermore, the step 1.2 includes the following steps:

[0023] The RGB image and depth map are mapped to the latent space through the pre-trained variational autoencoder VAE to generate a low-dimensional latent representation; the latent representation is then converted into a tag sequence through a patchify process; robot motion encoding: the posture data collected by the CAN line is encoded into a latent representation through a multi-layer perceptron MLP, and linearly projected into a single tag; tactile image encoding: preprocessing: the tactile images collected by Gelsight Mini are first preprocessed, including denoising and normalization, to eliminate possible noise on the gel surface; feature extraction: the features of the tactile image are extracted through a specially designed convolutional neural network CNN to capture the texture, edge and force distribution information of the contact surface; latent representation: the extracted tactile features are mapped to the latent space through a VAE encoder to generate a low-dimensional latent representation; then it is converted into a tag sequence through a patching process; force distribution embedding: an additional force distribution embedding layer is introduced to encode the contact force information implicit in the tactile image into a low-dimensional vector, and splice it with the tactile latent representation to further enrich the information of the tactile modality.

[0024] Furthermore, the step 2 includes the following steps:

[0025] Step 2.1: Joint denoising The current observation is mapped to the noise latent space through each encoder, which is initialized to white noise;

[0026] Step 2.2: Dynamic masking with masked self-attention mechanism: Introducing masked self-attention mechanism to allow the model to flexibly handle missing modal data;

[0027] Step 2.3: Haptic Guidance: The tactilely generated latent representation provides detailed physical feedback to action generation via a masked self-attention mechanism;

[0028] Step 2.4: Output generation: The model simultaneously generates k-step future images, depth maps, tactile images, and robot actions;

[0029] Step 2.5: Model architecture;

[0030] Step 2.6: Training target.

[0031] Furthermore, the step 2 includes the following steps:

[0032] The noisy latent representation of tactile data consists of two parts: the latent representation of the tactile image and the force distribution embedding. These two parts are spliced ​​with the conditional latent representation in the channel dimension to form a multimodal conditional noisy latent representation. The transformer block jointly denoises the spliced ​​label sequence through a multi-layer self-attention mechanism to predict the multimodal output in the next k steps.

[0033] Furthermore, the masked self-attention mechanism in step 2.3 dynamically adjusts the contribution of the tactile modality to action generation in a weighted manner.

[0034] Furthermore, step 2.4 includes tactile output decoding: the tactile image is reconstructed into a high-resolution image through a decoder, reflecting the morphology of the future contact surface; the contact force distribution is reconstructed into a force vector through an MLP decoder, representing the force change in the next k steps.

[0035] Furthermore, the model architecture of step 2.4 is as follows: the transformer backbone adopts a multi-layer transformer as the core architecture, supports the splicing and processing of multimodal tag sequences, and the attention layer dynamically adapts to the input length and missing conditions of different modalities through a mask mechanism. An enhanced attention sublayer is specially designed for the tactile modality to specifically handle the interaction between tactile morphology and force distribution, ensuring the efficient use of tactile information.

[0036] Furthermore, step 2.6 uses the denoising diffusion probability model DDPM loss function to jointly optimize multimodal prediction:

[0037]

[0038] λ I ,λ D ,λ T ,λ F ,λ A are the balance hyperparameters for image, depth, tactile topography, tactile force distribution and motion loss, respectively, F It can be dynamically adjusted to balance the contribution of tactile morphology and force distribution. The joint denoising method for robot visual action prediction of the present invention has the following advantages: 1. Efficient fusion of multimodal information

[0039] This paper constructs a unified generative model by fusing visual data (RGB images and depth maps), robot motion data (collected via CAN communication), and tactile information (collected via the Gelsight Mini sensor). This multimodal fusion mechanism leverages the complementarity between different modalities, significantly improving the robot's perception of complex environments, especially in scenarios with visual occlusion or requiring precise force control.

[0040] 2. Joint Denoising and Collaborative Generation

[0041] A Transformer-based joint denoising framework enables the collaborative generation of multimodal data. This unified generation process ensures consistency among images, depth maps, tactile data, and robot motion predictions, avoiding the inconsistencies inherent in independent generation in traditional methods and improving prediction accuracy and reliability.

[0042] 3. Dynamic Masked Self-Attention Mechanism

[0043] The introduction of a masked self-attention mechanism enables the model to flexibly handle missing modal data and dynamically adjust the contribution of different modalities to the prediction results. This mechanism enhances the robustness of the system, maintaining high prediction performance even when some modal data is missing.

[0044] 4. Tactile-guided fine motor control

[0045] The tactile modality senses the topography and force distribution of the contact surface, providing detailed physical feedback for robot motion generation. The tactile-enhanced attention sublayer specifically processes the interaction between tactile topography and force distribution, ensuring efficient utilization of tactile information. This enables the robot to dynamically adjust gripping force and contact position when manipulating fragile objects or tasks requiring precise force control, avoiding object damage or operational failure.

[0046] 5. Optimized encoding and decoding architecture

[0047] Using pre-trained variational autoencoders (VAE) and specially designed convolutional neural networks (CNN)

[0048] The system efficiently encodes data from different modalities to generate low-dimensional latent representations. The decoder is then able to reconstruct high-resolution tactile images and accurate force distribution vectors, further improving the generation quality and practicality of the model.

[0049] 6. Joint Optimization Loss Function

[0050] Using the denoising diffusion probability model (DDPM) loss function, the weights of different modal predictions (such as tactile topography and force distribution) are dynamically balanced to ensure the coordination and consistency of multimodal predictions. This joint optimization mechanism significantly improves the overall performance of the model.

[0051] 7. Real-time and system integration

[0052] Hardware trigger signals enable efficient synchronization of multimodal data (with an error of less than 10ms), and combined with an optimized network architecture, a real-time inference frequency of 30Hz is achieved. The complete system integration solution (including hardware configuration, software architecture, and interface definition) facilitates practical applications.

[0053] 8. Wide range of application scenarios

[0054] This method is applicable to various scenarios such as home services (such as grasping fragile items), industrial assembly (such as precision connector plugging), and medical assistance (such as sorting medical devices), and can significantly improve the accuracy and robustness of robots in complex operation tasks.

[0055] In summary, the present invention effectively solves the problems of insufficient modal collaboration, strong independence of the generation process, and lack of fine force control capabilities in the existing technology through a multimodal joint denoising generation framework, a dynamic mask self-attention mechanism, and tactile-guided motion generation, providing an efficient, robust, and practical solution for robot visual motion prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 This is a diagram of the multimodal joint denoising network architecture of the present invention;

[0057] Figure 2 Schematic diagram of the Diffusion Transformer architecture of the present invention. DETAILED DESCRIPTION

[0058] In order to better understand the purpose, structure and function of the present invention, the following further describes in detail a joint denoising method for robot visual action prediction of the present invention in conjunction with the accompanying drawings.

[0059] like Figure 1As shown in the figure, the present invention proposes a joint denoising method for robot visual motion prediction, which builds a unified generative model by fusing images and depth maps collected from depth cameras, robot motion data collected by Piper robot arm CAN line communication, and tactile images collected by Gelsight Mini visual tactile sensor. The framework uses the transformer (Diffusion Transformer) architecture to simultaneously predict future images, robot motion, depth maps and tactile data through a joint denoising process. In particular, a clever masked self-attention mechanism is adopted to enable image generation, depth generation and tactile generation to collaboratively guide motion generation. Tactile generation provides fine physical feedback for motion generation by sensing the morphology and force distribution of the contact surface, thereby significantly improving the accuracy and robustness of robot control.

[0060] like Figure 1 As shown, the specific steps include:

[0061] Step 1: Data collection and input coding:

[0062] Step 1.1: Data sources: Image and depth map: RGB images and corresponding depth information are collected in real time through a depth camera (such as Intel RealSense or similar devices). Robot motion: High-dimensional posture data, including end effector position, rotation angle, and gripper state (7-dimensional vector), are collected through the CAN line communication protocol of the Piper robotic arm. Tactile image: High-resolution tactile data is collected using the Gelsight Mini visual tactile sensor. Gelsight Mini captures the topography and force distribution of the contact surface through its gel surface and internal camera to generate high-resolution tactile images (e.g., 256x256 pixels). These images not only reflect the texture and geometry of the contact area, but also indirectly characterize the magnitude and direction of the contact force through gel deformation.

[0063] Step 1.2: Encoding Process: Image and Depth Map Encoding: The RGB image and depth map are mapped to a latent space using a pretrained variational autoencoder (VAE) to generate a low-dimensional latent representation. The latent representation is then converted into a sequence of tokens (e.g., 256 tokens for a 256x256 image) through a patchifying process. Robot Action Encoding: The posture data collected by the CAN line is encoded into a latent representation using a multi-layer perceptron (MLP) and linearly projected into a single token. Tactile Image Encoding: Preprocessing: Tactile images collected by the Gelsight Mini are first preprocessed, including denoising and normalization, to remove potential noise on the gel surface (such as artifacts caused by illumination changes or gel aging). Feature Extraction: Features of the tactile images are extracted using a specially designed convolutional neural network (CNN), capturing information about the texture, edges, and force distribution of the contact surface. For example, the CNN can extract a gradient map of gel deformation as an approximate representation of the force distribution. Latent Representation: The extracted tactile features are mapped to a latent space using a VAE encoder to generate a low-dimensional latent representation (e.g., 32x32x4). It is then converted into a sequence of markers (e.g., 16 markers) through a block-based process. Force distribution embedding: An additional force distribution embedding layer is introduced to encode the contact force information implicit in the tactile image (inferred by gel deformation) into a low-dimensional vector (e.g., 8-dimensional), which is then concatenated with the tactile latent representation to further enrich the information of the tactile modality.

[0064] Step 2: Joint denoising and generation process:

[0065] Step 2.1: Joint Denoising The current observation (image, depth map, robot state, tactile data) is mapped to a noisy latent space through its respective encoders, initialized to white noise. The noisy latent representation of the tactile data consists of two parts: a latent representation of the tactile image (reflecting topography and texture) and a force distribution embedding (reflecting contact forces). These two parts are concatenated with the conditional latent representation (the encoding of the current observation) in the channel dimension to form a multimodal conditional noisy latent representation. The transformer block jointly denoises the concatenated labeled sequence through a multi-layer self-attention mechanism to predict the multimodal output (image, depth map, tactile image, contact forces, and robot motion) for the next k steps.

[0066] Step 2.2: Dynamic Masking: A masked self-attention mechanism is introduced to allow the model to flexibly handle missing modal data. For example, if tactile data is unavailable, the mask will suppress the relevant markers and only retain the contribution of the valid modality.

[0067] Step 2.3: Tactile guidance: The latent representation generated by tactile generation (including topography and force distribution) provides detailed physical feedback to action generation through a masked self-attention mechanism. For example, in a grasping task, the contact force distribution predicted by tactile generation can guide the force adjustment of the gripper action to avoid slipping or damaging the object. The topographic information generated by tactile generation works in conjunction with image and depth generation to enhance the understanding of the object surface. For example, tactile perception can perceive tiny textures that are invisible in the image, helping the model predict more accurate contact points. Cross-modal interaction: The masked self-attention mechanism dynamically adjusts the contribution of the tactile modality to action generation in a weighted manner. For example, when the tactile signal detects a high-friction surface, the action generation will prioritize adjusting the gripper's grasping angle.

[0068] Step 2.4: Output Generation: The model simultaneously generates k-step-ahead images, depth maps, tactile images (including topography and contact force distribution), and robot motions. Tactile Output Decoding: The tactile image is reconstructed into a high-resolution image by a decoder (the VAE decoder corresponding to the encoder) that reflects the future contact surface topography. The contact force distribution is reconstructed into a force vector by an MLP decoder that represents the force variation k steps into the future.

[0069] Step 2.5: Model architecture: Figure 2 As shown, the Transformer backbone uses a multi-layered Transformer (Diffusion Transformer) as its core architecture, supporting the concatenation and processing of multimodal labeled sequences. The attention layer dynamically adapts to the input length and missingness of different modalities through a masking mechanism. An enhanced attention sublayer is specifically designed for the tactile modality to specifically process the interaction between tactile topography and force distribution, ensuring efficient utilization of tactile information.

[0070] Step 2.6: Training objective: Jointly optimize multimodal prediction using the denoising diffusion probability model (DDPM) loss function:

[0071]

[0072] λ I ,λ D ,λ T ,λ F ,λ A are the balancing hyperparameters for image, depth, tactile topography, tactile force distribution, and motion loss, respectively. In particular, λ F Dynamically adjustable to balance the contributions of tactile topography and force distribution.

[0073] This paper proposes a multimodal joint denoising and generation framework that integrates data from a depth camera, a Piper robotic arm, and a Gelsight Mini tactile sensor to construct a unified generation model. The following is a specific implementation:

[0074] 1. System hardware configuration

[0075] Depth camera: Intel RealSense D435i, resolution 1280×720, 30fps, mounted on the robot head or arm, captures RGB images and depth maps.

[0076] Robotic arm: Piper robotic arm, CAN bus communication, sampling frequency 100 Hz, 7 degrees of freedom, two-finger gripper at the end, maximum load 5 kg.

[0077] Tactile sensor: Gelsight Mini visual tactile sensor, with a resolution of 640×480 and 30fps, is installed on the inside of the end effector to capture the contact surface topography and force distribution.

[0078] Computing platform: NVIDIA RTX 3090 GPU (24GB video memory), Intel Core i9-11900K processor, 64GB memory for model training and real-time inference.

[0079] 2 Data collection and preprocessing process

[0080] 2.1 Data Collection

[0081] Visual data: The depth camera collects RGB images and depth maps at 30Hz with a resolution of 640×480, ensuring stable lighting.

[0082] Robot motion: The end effector position (x, y, z), posture (quaternion) and gripper state are collected at 100 Hz via the CAN bus and stored as a 7-dimensional vector.

[0083] Tactile data: Gelsight Mini collects tactile images at 30 Hz with a resolution of 640 × 480, records gel deformation, and calculates contact force distribution.

[0084] Synchronization mechanism: Hardware trigger signals ensure data synchronization with an error of less than 10ms, and timestamps align data streams of different frequencies.

[0085] 2.2 Data Preprocessing

[0086] Visual data: RGB images are normalized to [0, 1], depth maps are padded and filtered, cropped to 256×256, and data augmented (flipped, rotated, color jittered).

[0087] Robot motion: Position and gripper state are normalized to [-1, 1] and [0, 1], rotation is represented by quaternion and downsampled to 30Hz.

[0088] Tactile data: tactile image denoising, calculation of gel deformation gradient map to extract force distribution, cropping to 256×256, and data augmentation.

[0089] 3 Network Architecture and Implementation Details

[0090] 3.1 Encoder Design

[0091] Visual Encoder: ResNet-50 backbone combined with VAE to encode RGB images and depth maps into 32×32×8 latent representations, partitioned into 256 tokens.

[0092] Action encoder: 3-layer MLP (hidden layers 128, 256, 128), encoding 7-dimensional action vectors into 64-dimensional tokens.

[0093] Tactile encoder: A 5-layer CNN (channels 32, 64, 128, 256, 512) extracts topography and force features, which are encoded into a 16×16×4 latent representation with blocks of 64 tokens; a 2-layer MLP encodes the contact force into an 8-dimensional vector.

[0094] 3.2 Diffusion Transformer Architecture

[0095] Overall architecture: 12-layer Transformer, 8-head self-attention, hidden layer dimension 768, feedforward dimension 3072, with LayerNorm and residual connections.

[0096] Masked self-attention: The mask matrix supports information exchange between modalities, action markers prioritize tactile markers, and the weights are learnable.

[0097] Tactile enhancement sublayer: Fusion of tactile topography and force distribution, and dynamic adjustment of the gating mechanism to influence action generation.

[0098] 3.3 Denoising Process

[0099] Initialization: The current observation is mapped to the latent space, and the next k-step targets are initialized to white noise (1000 noise levels).

[0100] Conditional splicing: The observation latent representation is spliced ​​with the noisy target, and the tactile topography and force distribution are spliced ​​separately.

[0101] Iterative denoising: 20 steps of denoising, each step reduces the noise level by 50, Transformer predicts the noise, and updates the target.

[0102] Output generation: Decoding generates k-step images, depth maps, tactile data, and actions, with a typical k value of 10 (approximately 0.33 seconds).

[0103] 4 Training method and parameter setting

[0104] Dataset: 50 objects (rigid and flexible), 5 operations (grasping, lifting, placing, pushing, rotating), each repeated 10 times, a total of 2500 sequences, each sequence is 3 seconds (90 frames, 30Hz).

[0105] Training parameters: batch size 32, learning rate 3e-4 (cosine annealing), AdamW optimizer, weight decay 1e-4, 200 epochs, loss weights λI = 1.0, λD = 0.8, λT = 1.2, λF = 1.5, λA = 2.0, training time approximately 72 hours.

[0106] Data augmentation: random cropping, flipping, rotation, color jittering, and Gaussian noise (σ=0.02) added to tactile images.

[0107] Validation and early stopping: 20% validation set, evaluation every 5 rounds, early stopping if there is no improvement in validation loss after 10 rounds.

[0108] 5 Experimental Verification and Performance Evaluation

[0109] Task settings:

[0110] Grab and lift different objects.

[0111] Remove items from container.

[0112] Precise placement of objects (accuracy ±5mm).

[0113] Handling fragile items (such as eggs).

[0114] Evaluation indicators: action prediction accuracy, task success rate, operation stability, force control accuracy, and inference speed.

[0115] 6 Application Scenario Examples

[0116] Home service robot: Take eggs out of the refrigerator and put them into a bowl, using tactile sensing to control the force to avoid damage.

[0117] Industrial assembly: Plugging precision connectors, tactile guidance for precise insertion.

[0118] Medical assistance: Sorting medical devices, using tactile control to ensure damage-free operation.

[0119] It will be understood that the present invention is described by way of some embodiments, and it will be appreciated by those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are intended to be protected by the present invention.

Claims

1. A joint denoising method for robot visual motion prediction, which builds a unified generative model by fusing images and depth maps collected by a depth camera, robot motion data collected by the Piper manipulator CAN line communication, and tactile images collected by the Gelsight Mini visual tactile sensor. The method is characterized by: The steps include: Step 1: Data collection and input coding; Step 2: Joint denoising and generation process.

2. The joint denoising method for robot visual action prediction according to claim 1, characterized in that: The step 1 comprises the following steps: Step 1.1: Data sources: images and depth maps; Step 1.2: Encoding process: image and depth map encoding.

3. The joint denoising method for robot visual action prediction according to claim 2, characterized in that: The step 1.1 includes the following steps: The depth camera collects RGB images and corresponding depth information in real time. Robot motion: High-dimensional posture data, including end-effector position, rotation angle, and gripper status, is collected through the CAN line communication protocol of the Piper robotic arm. Tactile image: High-resolution tactile data is collected using the Gelsight Mini visual tactile sensor. Gelsight Mini captures the morphology and force distribution of the contact surface through its gel surface and internal camera to generate high-resolution tactile images.

4. The joint denoising method for robot visual action prediction according to claim 2, characterized in that: The step 1.2 includes the following steps: The RGB image and depth map are mapped to the latent space through a pre-trained variational autoencoder (VAE) to generate a low-dimensional latent representation. The latent representation is then converted into a labeled sequence through a patchify process. Robot action encoding: The posture data collected by the CAN line is encoded into a latent representation through a multi-layer perceptron (MLP) and linearly projected into a single label. Tactile image encoding: Preprocessing: The tactile images collected by GelsightMini are first preprocessed, including denoising and normalization, to eliminate the noise that may exist on the gel surface; Features Extraction: A specially designed convolutional neural network (CNN) is used to extract features from tactile images, capturing the texture, edges, and force distribution information of the contact surface. Latent representation: The extracted tactile features are mapped to the latent space through a VAE encoder to generate a low-dimensional latent representation, which is then converted into a labeled sequence through a block process. Force distribution embedding: An additional force distribution embedding layer is introduced to encode the contact force information implicit in the tactile image into a low-dimensional vector, which is then concatenated with the tactile latent representation to further enrich the information of the tactile modality.

5. The joint denoising method for robot visual action prediction according to claim 1, characterized in that: The step 2 comprises the following steps: Step 2.1: Joint denoising The current observation is mapped to the noise latent space through each encoder, which is initialized to white noise; Step 2.2: Dynamic masking with masked self-attention mechanism: Introducing masked self-attention mechanism to allow the model to flexibly handle missing modal data; Step 2.3: Haptic Guidance: The tactilely generated latent representation provides detailed physical feedback to action generation via a masked self-attention mechanism; Step 2.4: Output generation: The model simultaneously generates k-step future images, depth maps, tactile images, and robot actions; Step 2.5: Model architecture; Step 2.6: Training target.

6. The joint denoising method for robot visual action prediction according to claim 5, characterized in that: The step 2 comprises the following steps: The noisy latent representation of tactile data consists of two parts: the latent representation of the tactile image and the force distribution embedding. These two parts are spliced ​​with the conditional latent representation in the channel dimension to form a multimodal conditional noisy latent representation. The transformer block jointly denoises the spliced ​​label sequence through a multi-layer self-attention mechanism to predict the multimodal output in the next k steps.

7. The joint denoising method for robot visual action prediction according to claim 5, characterized in that: The masked self-attention mechanism in step 2.3 dynamically adjusts the contribution of the tactile modality to action generation in a weighted manner.

8. The joint denoising method for robot visual action prediction according to claim 5, characterized in that: The step 2.4 includes tactile output decoding: the tactile image is reconstructed into a high-resolution image through a decoder, reflecting the morphology of the future contact surface; the contact force distribution is reconstructed into a force vector through an MLP decoder, representing the force change in the next k steps.

9. The joint denoising method for robot visual motion prediction according to claim 5, characterized in that: The model architecture of step 2.4 is as follows: the transformer backbone adopts a multi-layer transformer as the core architecture, supports the splicing and processing of multimodal tag sequences, and the attention layer dynamically adapts to the input length and missing conditions of different modalities through a masking mechanism. An enhanced attention sublayer is specially designed for the tactile modality to specifically handle the interaction between tactile morphology and force distribution, ensuring the efficient use of tactile information.

10. The joint denoising method for robot visual action prediction according to claim 5, characterized in that: Step 2.6 uses the denoising diffusion probability model DDPM loss function to jointly optimize multimodal prediction: λ I ,λ D ,λ T ,λ F ,λ A are the balance hyperparameters for image, depth, tactile topography, tactile force distribution and motion loss, respectively, F Dynamically adjustable to balance the contributions of tactile topography and force distribution.

Citation Information

Cited By

  • Power transmission line deicing equipment clamping control method and device

    CN121035890A

  • Power transmission line de-icing device clamping control method and device

    CN121035890B

  • Robot action prediction method and system based on cross-modal feature enhancement

    CN121105044A

  • A spherical space text-driven haptic signal generation method, device, medium and product

    CN122363530A

  • A spherical space text-driven haptic signal generation method, device, medium and product

    CN122363530B