The invention discloses a joint denoising method for
robot visual motion prediction, and the method comprises the steps: constructing a unified
generative model through fusing an image and a
depth map collected by a depth camera, motion data collected by CAN line communication of a
Piper mechanical arm, and a tactile image collected by a Gelsight Mini
tactile sensor; the method comprises two steps of
data acquisition and input coding, and joint denoising and generation: firstly, multi-
modal data are coded into low-dimensional potential representation, and then future images, depth maps, tactile data and
robot actions are cooperatively predicted through a joint denoising framework based on Transform. A
mask self-attention mechanism is innovatively introduced, information interaction between
modes is dynamically adjusted, action generation is guided through tactile feedback, and the force control precision is improved. The model adopts a de-noising
diffusion probability
loss function to jointly optimize multi-
modal prediction, so that the output consistency is ensured. According to the method, the robustness and the accuracy of flexible operation of the
robot are remarkably improved.