A robot manipulation method, device, system, and medium based on a diffusion model.
By using diffusion model learning to generate robotic arm manipulation actions from human hand teaching data, the problem of high computational load and high teaching cost in the control of flexible linear objects in existing technologies is solved, realizing efficient and accurate flexible object manipulation, which is suitable for multiple application scenarios.
Patent Information
- Application Number
- CN202511333352.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-18
AI Technical Summary
Existing technologies struggle to control flexible linear objects accurately and in real time. They involve large computational loads, rely on complex physical modeling and annotation data, and have high teaching costs and poor versatility.
A diffusion model is adopted to construct training samples by collecting information on the position of the index finger tip and the state of the flexible linear object during the teaching process. The noise prediction network in the diffusion model is used to learn the mapping from the high-noise state to the clear target state, and generate the control action sequence of the robotic arm.
It improves the accuracy and real-time performance of manipulating flexible linear objects, reduces data acquisition costs, and enhances the scalability and versatility of tasks, making it suitable for scenarios such as intelligent manufacturing, service robots, and medical-assisted surgery.
Smart Images

Figure CN120828424B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a robot manipulation method, apparatus, system and medium based on a diffusion model. Background Technology
[0002] Flexible linear objects (such as cables, ropes, and conduits) are widely used in various fields, including industrial manufacturing, medical surgery, and home service robots. Compared to rigid objects, flexible linear objects possess characteristics such as infinite degrees of freedom, nonlinear deformation, and strong coupling, posing significant challenges to robot manipulation. How to enable robots to accurately perceive and manipulate flexible linear objects has become one of the cutting-edge issues in robot perception and control research in recent years.
[0003] To manipulate flexible objects, many related technologies rely on physical modeling and simulation, such as using finite element models, point mass spring models, or Kirchhoff models to simulate the deformation of the object. However, these methods depend on accurate modeling parameters, involve large computational loads, and are difficult to apply in real-time to practical task scenarios.
[0004] On the other hand, vision-based methods for perceiving flexible objects have gained attention in recent years. However, most methods still rely on large-scale labeled data or hand-designed intermediate representations and lack the ability to model complex deformations.
[0005] In trajectory generation, most methods employ imitation learning or reinforcement learning strategies for policy modeling. While combining imitation learning with graph neural networks improves accuracy, it still faces challenges in generalization and computational complexity.
[0006] In addition, in terms of teaching methods, robot learning usually relies on dedicated teach pendants or manually labeled data, which not only increases the cost of data collection, but also limits the scalability and versatility of tasks. Summary of the Invention
[0007] In view of this, the purpose of the embodiments of the present invention is to provide a robot manipulation method, apparatus, system and medium based on a diffusion model, so as to solve one or more technical problems existing in the prior art and provide at least one beneficial option or create conditions.
[0008] On one hand, embodiments of the present invention provide a robot manipulation method based on a diffusion model, the method comprising the following steps:
[0009] Collect information on the position of the index finger tip and the state of the flexible linear object during the hand demonstration process to construct training samples;
[0010] A diffusion model is used to train the training samples. The noise prediction network in the diffusion model learns the mapping from a high-noise state to a clear target state, and a trained diffusion model is obtained. The training process of the diffusion model includes forward noise addition and inverse noise removal.
[0011] The state of the flexible linear object and the position of the end effector of the robotic arm are collected to construct the current observation environment. The current observation environment is then input into the trained diffusion model to generate the target action sequence. Based on the target action sequence, the robotic arm is controlled to perform manipulation actions.
[0012] Optionally, the step of collecting the position information of the index finger tip and the state information of the flexible linear object during the hand-teaching process to construct training samples includes:
[0013] Acquire video images during the hand demonstration process, convert the flexible linear objects in the video images into an ordered sequence of points, and simultaneously acquire the position information of the index finger tip during the hand demonstration process;
[0014] The ordered point sequence of each flexible linear object, the position of the index finger tip, and the state nodes of the target are stitched together to form the observation environment;
[0015] The observation environment of the current frame, the observation environment of the previous frame, and the state node of the target are used as inputs, and the action sequence formed by the trajectory of the index finger at the corresponding moment is used as training labels to form training samples.
[0016] Optionally, the step of converting the flexible linear object in the video image into an ordered sequence of points and simultaneously acquiring the position information of the index finger tip during the hand demonstration process includes:
[0017] Based on visual algorithms, state modeling is performed on flexible linear objects in video images, and an ordered point sequence representing the curve shape of the flexible linear object is extracted; the ordered point sequence includes multiple state nodes uniformly distributed along the center line of the flexible linear object.
[0018] A hand tracking module was used to identify key points of the hand skeleton in the video image to obtain the position information of the index finger tip during the hand teaching process.
[0019] Optionally, the step of performing state modeling on the flexible linear object in the video image based on the visual algorithm and extracting an ordered sequence of points representing the curve shape of the flexible linear object includes:
[0020] The centerline of the flexible linear object is obtained by the skeleton extraction algorithm. Multiple sampling points are evenly selected on the centerline at preset intervals, and the sampling points are used as candidate nodes of the ordered point sequence.
[0021] A Gaussian mixture model is used to treat each candidate node as a centroid of a Gaussian distribution. The position of the candidate node is iteratively adjusted by the expectation-maximization algorithm so that the candidate node fits the actual shape of the flexible linear object.
[0022] Each candidate node is smoothed to obtain an ordered sequence of points that characterizes the curve shape of a flexible linear object.
[0023] Optionally, the step of training the training samples using a diffusion model, and learning the mapping from a high-noise state to a clear target state through the noise prediction network in the diffusion model, to obtain a trained diffusion model, includes:
[0024] In the forward noise addition stage, Gaussian noise is gradually added to the action sequence according to the preset noise scheduling strategy to generate noise action sequences at different noise levels.
[0025] In the reverse denoising stage, the noise action sequence and the observation environment are used as inputs. The loss value is determined based on the predicted noise output by the noise prediction network. The parameters of the diffusion model are updated to reduce the loss value to the loss threshold, and the trained diffusion model is obtained.
[0026] Optionally, the step of taking the noisy action sequence and the observation environment as input, determining the loss value based on the predicted noise output by the noise prediction network, and updating the parameters of the diffusion model to reduce the loss value to a loss threshold to obtain a trained diffusion model includes:
[0027] Randomly sample the noisy action sequence, select one time step from all noisy time steps each time, and extract the noisy action sequence, the original action sequence, and the real noise corresponding to that time step.
[0028] The noise action sequence, the embedding vector of the time step, and the observation environment are input into a noise prediction network based on a one-dimensional U-Net structure, and the predicted noise is output.
[0029] Calculate the mean squared error loss between the predicted noise and the actual noise, and update the parameters of the noise prediction network through the backpropagation algorithm to reduce the mean squared error loss value to the loss threshold, thus obtaining the trained noise prediction network.
[0030] Optionally, the step of acquiring the state of the flexible linear object and the position of the robotic arm's end effector to construct the current observation environment, inputting the current observation environment into a trained diffusion model to generate a target action sequence, and controlling the robotic arm to perform manipulation actions based on the target action sequence includes:
[0031] Real-time acquisition of current state images of flexible linear objects, conversion into ordered point sequences, and fusion with the real-time position coordinates of the robotic arm end effector to construct the current observation environment;
[0032] The current observation environment and target state node are input into the trained diffusion model, and the target action sequence of the end effector is generated through a multi-step inverse denoising process. This target action sequence contains the position and orientation information of the end effector in Cartesian space.
[0033] The robotic arm control system drives the joint motors to perform corresponding control actions based on the generated target motion sequence, thereby realizing the deformation and adjustment of the flexible linear object towards the target state.
[0034] On the other hand, embodiments of the present invention provide a robot manipulation device based on a diffusion model, comprising:
[0035] The data acquisition module is used to collect information on the position of the index finger tip and the state information of the flexible linear object during the hand demonstration process to build training samples;
[0036] The model training module is used to train the training samples using the diffusion model. By training the noise prediction network in the diffusion model, the model learns the mapping from a high-noise state to a clear target state, thus obtaining a trained diffusion model. The training process of the diffusion model includes forward noise addition and inverse noise removal.
[0037] The motion control module is used to collect the state of the flexible linear object and the position of the robotic arm end effector, construct the current observation environment, input the current observation environment into the trained diffusion model, generate the target action sequence, and control the robotic arm to perform manipulation actions based on the target action sequence.
[0038] On the other hand, embodiments of the present invention provide a robot control system based on a diffusion model, comprising:
[0039] At least one processor;
[0040] At least one memory for storing at least one program;
[0041] When the at least one program is executed by the at least one processor, the at least one processor performs the method described above.
[0042] On the other hand, embodiments of the present invention provide a computer-readable storage medium storing a processor-executable program, which, when executed by a processor, is used to perform the above-described method.
[0043] The embodiments of the present invention have the following beneficial effects:
[0044] This invention constructs training samples by collecting information on the position of the index finger tip and the state information of a flexible linear object during the teaching process. A diffusion model is then used to train the training samples, enabling the model to learn the mapping from a high-noise state to a clear target state. In actual operation, the model generates a target action sequence based on the current observation environment to control the robotic arm to perform manipulation actions. This effectively solves the problems of traditional methods, such as reliance on accurate physical modeling parameters, high computational load, insufficient generalization ability, and high teaching costs. It improves the accuracy and real-time performance of the robot's manipulation of flexible linear objects, reduces data acquisition costs, and enhances the scalability and versatility of the task. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a flowchart illustrating the steps of a robot manipulation method based on a diffusion model provided in an embodiment of the present invention.
[0047] Figure 2 This is a framework diagram of a robot manipulation method based on a diffusion model provided in an embodiment of the present invention;
[0048] Figure 3 This is a schematic diagram of the diffusion model generating action sequences in an embodiment of the present invention;
[0049] Figure 4 This is a schematic diagram of a one-dimensional U-Net action sequence noise prediction network in an embodiment of the present invention;
[0050] Figure 5 This is a structural block diagram of a robot control device based on a diffusion model provided in an embodiment of the present invention;
[0051] Figure 6 This is a structural block diagram of a robot control system based on a diffusion model provided in an embodiment of the present invention. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0053] It should be noted that although the device diagram shows a modular division and the flowchart illustrates a logical order, in some cases, the steps shown or described may be performed in a different order than the modular division in the device or the order shown in the flowchart. The terms "first," "second," etc., used in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0055] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0056] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0057] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0058] In related technologies, image-to-trajectory mapping methods based on deep neural networks extract the state of flexible objects from images and then predict the trajectory of the robot's end effector. Video prediction models are used to infer the future state of the flexible object, thereby generating a control strategy. While these methods are effective, their predictions are unstable, convergence is slow, and they are limited by the quality of the dataset.
[0059] The relevant technologies still have the following main problems in robot manipulation of flexible linear objects:
[0060] Limited shape representation methods: Related technologies typically use images, point clouds, or meshes to represent the shape of flexible objects, lacking effective modeling of the essential geometric features of the shape, making it difficult to use as a control target to accurately guide robots to complete manipulation tasks.
[0061] The teaching method is highly dependent: the teaching method of related technologies requires dedicated hardware or cumbersome annotation process, which not only increases the cost of data collection and processing, but also limits the scalability and practicality of the system.
[0062] This invention aims to provide a robot manipulation method, device, system, and medium based on a diffusion model. This invention utilizes a visual approach to reduce the shape of flexible objects to an ordered series of points, thus better representing their target form. It uses human hand manipulation videos as teaching data, equating fingertips with robot end effectors, significantly lowering the data acquisition threshold and improving training efficiency. It is suitable for robots to perform precise manipulation and shape adjustment when handling flexible linear objects such as cables, wires, ropes, and hoses. This method generates trajectories through visual perception combined with a diffusion model, possessing good versatility and robustness. It can be widely applied to flexible object manipulation tasks in intelligent manufacturing, service robots, and medical-assisted surgery, improving automation and operational accuracy.
[0063] like Figure 1 and Figure 2 As shown, Figure 1 A robot manipulation method based on a diffusion model is provided in this embodiment of the invention. The method includes the following steps:
[0064] S100 collects information on the position of the index finger tip and the state of the flexible linear object during the hand teaching process to construct training samples.
[0065] S200, a diffusion model is used to train the training samples. The noise prediction network in the diffusion model is trained to learn the mapping from a high-noise state to a clear target state, and a trained diffusion model is obtained. The training process of the diffusion model includes forward noise addition and reverse noise removal.
[0066] S300 collects the state of the flexible linear object and the position of the end effector of the robotic arm, constructs the current observation environment, inputs the current observation environment into the trained diffusion model, generates the target action sequence, and controls the robotic arm to perform manipulation actions based on the target action sequence.
[0067] This invention proposes a robot manipulation method, device, system, and medium based on a diffusion model. During the teaching process, existing vision solutions are used to detect the state representation of a flexible linear object and the position of the fingertip. When acquiring human hand teaching videos, a binocular vision system or depth camera is used to obtain three-dimensional coordinate information. The two-dimensional image coordinates are converted into three-dimensional spatial coordinates through camera calibration parameters, ensuring that the spatial position data of the index fingertip position and the state of the flexible linear object have a unified coordinate system, thereby improving the accuracy of model training and inference.
[0068] Specifically, the state of the flexible linear object is abstracted as a set of ordered nodes evenly distributed along the centerline of the flexible linear object and the position of the index fingertip. These two sets of information are merged and considered as the model's observation environment, which serves as the model's input during training. The final state of the flexible linear object is used as the training constraint, and the change in the fingertip's position is considered the model's action, serving as the training label during training. The diffusion model is then trained to possess action prediction capabilities. During the inference phase, the real-time state of the flexible linear object and the position of the robotic arm's end effector are acquired through a camera and a robot teach pendant to obtain the observation environment. The state of the target shape is then input into the trained diffusion model. The diffusion model generates a sequence of actions for the robotic arm's end effector, which the robot executes to achieve shape control of the deformable flexible linear object.
[0069] In some embodiments, the step of collecting the position information of the index finger tip and the state information of the flexible linear object during the hand teaching process to construct training samples includes:
[0070] S110, acquire video images during the hand teaching process, convert the flexible linear objects in the video images into an ordered point sequence, and simultaneously acquire the position information of the index finger tip during the hand teaching process;
[0071] S120 stitches together the ordered point sequence of each frame of the flexible linear object, the position of the index finger tip, and the state nodes of the target to form the observation environment;
[0072] S130 takes the observation environment of the current frame, the observation environment of the previous frame, and the state node of the target as input, and takes the action sequence formed by the trajectory of the index finger at the corresponding moment as training labels to form training samples.
[0073] In this embodiment, video images of the hand demonstration process are acquired, and the contour features of the flexible linear object are extracted using image recognition technology, converting them into an ordered sequence of points uniformly distributed along the center line. Simultaneously, a target detection algorithm is used to capture the real-time position coordinates of the index fingertip during the hand demonstration. The ordered point sequence, index fingertip position coordinates, and preset target state nodes of each frame are concatenated to construct the model's observation environment data structure. Multiple consecutive frames of observation environment data are arranged chronologically and combined with the action sequence formed by the index fingertip movement trajectory within the corresponding time period to form a training sample set for parameter learning of the diffusion model.
[0074] This embodiment effectively reduces the complexity of teaching data collection by transforming visual information during the human hand teaching process into structured training data, while preserving the key correlation features between the deformation of flexible linear objects and the manipulation actions, providing high-quality input samples for the subsequent training of the diffusion model.
[0075] The implementation process of the solution proposed in this invention is as follows: Figure 2 As shown, the specific implementation steps for each stage are as follows:
[0076] I. Teaching and Data Collection
[0077] Currently, most diffusion model strategies take video of the working environment as input and output an end-to-end motion control scheme for the robot. However, due to the high dimensionality and large data volume of video information, traditional models often struggle to achieve good training results with limited video data. Therefore, in this invention, the input of the diffusion model uses low-dimensional point sequence information after visual processing. Specifically, a visual tracking algorithm for deformable flexible linear objects is used, which represents the shape changes by an ordered sequence of points uniformly distributed along the object's axis.
[0078] Furthermore, traditional data acquisition methods using robotic arms for teaching are limited by time consumption, high labor intensity, and high cost. However, hand-based teaching data can significantly reduce data acquisition time and cost. While the differences between human hands and robotic arms need to be addressed for video-based end-to-end diffusion model strategies, the state representation-based input method of this invention only focuses on the position of the point of interaction with the cable. Therefore, the position of the fingertip can be directly captured through a visual solution and used as the interaction point input for training the diffusion model.
[0079] This approach enables the model to learn the motion patterns of the interaction points. During the inference phase, by using the hand-eye calibration of the camera and the robotic arm, combined with the inverse kinematics of the robotic arm, a mapping relationship between the position of the moving point and the position of the robotic arm's end effector can be established, thus enabling the robotic arm to precisely control the shape of deformable flexible linear objects.
[0080] By reducing the dimensionality of model inputs and simplifying the process of collecting teaching data, the complexity of data collection and training is reduced, while the robustness and practicality of the model in real-world applications are improved. This innovative approach effectively reduces reliance on complex manual teaching, lowers development costs, and adapts to a wider range of application scenarios.
[0081] In some embodiments, converting the flexible linear object in the video image into an ordered sequence of points and simultaneously acquiring the position information of the index finger tip during the hand teaching process includes:
[0082] Based on visual algorithms, state modeling is performed on flexible linear objects in video images, and an ordered point sequence representing the curve shape of the flexible linear object is extracted; the ordered point sequence includes multiple state nodes uniformly distributed along the center line of the flexible linear object.
[0083] A hand tracking module was used to identify key points of the hand skeleton in the video image to obtain the position information of the index finger tip during the hand teaching process.
[0084] This embodiment models the state of a flexible linear object using visual algorithms. Specifically, it employs deep learning-based image segmentation technology to accurately extract the contour region of the flexible linear object from the video image. Then, a curve fitting algorithm is used to process the contour to obtain its centerline. Subsequently, multiple points are uniformly selected along the centerline at preset sampling intervals to form an ordered point sequence. The coordinate information of each point collectively constitutes a quantitative description of the current shape of the flexible linear object. For the hand tracking module, an existing open-source hand pose estimation model can be integrated. This model can detect the hand in the video image in real time and output the three-dimensional coordinates of multiple skeletal key points, including the fingertip. In terms of synchronous acquisition, by aligning the frame number of the video image with the timestamp output by the hand tracking module, it is ensured that the ordered point sequence of the flexible linear object in each frame corresponds to the position information of the fingertip at the same moment, laying a precise data foundation for the subsequent construction of training samples.
[0085] Specifically, to achieve an abstract representation of the state of a flexible linear object, this invention employs a pure vision algorithm for real-time tracking of deformable linear objects in occluded scenarios. Its core objective is to abstract the state of the flexible linear object into an ordered sequence of points uniformly distributed along its centerline. ,in Indicates the number of nodes. Indicates the first The pixel coordinates of each node. The idea is to use the cable's status nodes... It is considered as the centroid of multiple Gaussian distributions in a Gaussian mixture model (GMM), i.e., the algorithm assumes that the point cloud originates from a flexible linear object. The points in the diagram are randomly sampled from these Gaussian distributions, that is... ,in, This represents the nth point in the point cloud at time t. This represents the coordinates of the nth state node at time t. Let V be the variance of the Gaussian distribution. The nodes are represented by a Gaussian distribution. The Expectation-Maximization (EM) algorithm is then used for iterative optimization to determine their final positions. These nodes strictly adhere to the topology of the flexible linear object, uniformly distributed along its centerline at fixed intervals, completely and continuously outlining the object's geometry. The connections between adjacent nodes form piecewise linear curves, accurately reflecting the bending, stretching, and other deformation states of the flexible linear object in space. The coordinates of each node not only identify its specific location in the image but also convey the object's overall posture information through the relative positions of the nodes; for example, changes in the direction of the node sequence can indicate the object's bending direction.
[0086] In some embodiments, the state modeling of the flexible linear object in the video image based on the visual algorithm, and the extraction of an ordered sequence of points representing the curve shape of the flexible linear object, includes:
[0087] The centerline of the flexible linear object is obtained by the skeleton extraction algorithm. Multiple sampling points are evenly selected on the centerline at preset intervals, and the sampling points are used as candidate nodes of the ordered point sequence.
[0088] A Gaussian mixture model is used to treat each candidate node as a centroid of a Gaussian distribution. The position of the candidate node is iteratively adjusted by the expectation-maximization algorithm so that the candidate node fits the actual shape of the flexible linear object.
[0089] Each candidate node is smoothed to obtain an ordered sequence of points that characterizes the curve shape of a flexible linear object.
[0090] This embodiment uses a skeleton extraction algorithm to separate the main structure of a flexible linear object from a preprocessed video image. Redundant pixels are removed through a thinning operation, retaining the centerline contour with a single-pixel width. Next, a sampling interval is set based on the object's length and shape complexity, and candidate nodes are selected at equal intervals along the centerline. This ensures that the node distribution fully covers the object's shape while avoiding data redundancy due to excessive density. Subsequently, the candidate nodes are used as the initial centroids of a Gaussian mixture model. The parameters of each Gaussian distribution are continuously optimized using an expectation-maximization algorithm, ensuring that the point cloud distribution generated by the model highly matches the actual contour features of the object in the image, thereby accurately adjusting the node positions. Finally, a moving average filter is used to smooth the node coordinates, eliminating positional fluctuations caused by noise interference, resulting in a continuous and stable ordered point sequence, providing accurate shape representation data for subsequent model training.
[0091] First, the input video image is preprocessed, including image denoising and contrast enhancement, to improve the accuracy of subsequent processing. Next, an edge detection algorithm is used to extract the contour information of the flexible linear object, initially determining its approximate region in the image. Then, a deep learning-based instance segmentation model is employed to separate the flexible linear object from the complex background, resulting in a binarized image containing only the target object. Afterward, a skeleton extraction algorithm is used to obtain the centerline of the flexible linear object, which accurately reflects the overall shape and orientation of the object. Based on this, multiple sampling points are uniformly selected along the centerline according to preset interval parameters; these sampling points serve as initial candidate nodes for an ordered point sequence. To further optimize the node positions, a Gaussian Mixture Model (GMM) is introduced, treating each candidate node as a centroid of a Gaussian distribution. The Expectation-Maximization (EM) algorithm iteratively adjusts the node positions, allowing the nodes to better conform to the actual shape of the flexible linear object. Simultaneously, dynamic programming is incorporated to smooth the node sequence, ensuring consistent spacing between adjacent nodes and avoiding local abrupt changes. Finally, an ordered point sequence that accurately represents the curve shape of the flexible linear object is obtained.
[0092] Then, the Media Pipe framework is used to identify key points on the operator's hand to obtain the position information of the fingertips. Media Pipe is an open-source, cross-platform computer vision framework developed by Google, widely used in tasks such as real-time pose estimation, face detection, and hand tracking, and has good computational efficiency and real-time performance. In this invention, its built-in Hand Tracking module is used to process the video images captured by the camera, which can output coordinate data containing 21 three-dimensional hand key points in real time, including the position of the index fingertip.
[0093] During data acquisition, the system continuously captures images of hand movements and uses the Media Pipe model to identify and locate key points of the hand in each frame, extracting the coordinates of the index finger's tip. This location is considered an expression of the human operator's control intention and serves as a representative point of the "robot's end-effector" in subsequent data annotation and model training, participating in the mapping and modeling of the entire manipulation action.
[0094] In some embodiments, training the training samples using a diffusion model, and learning the mapping from a high-noise state to a clear target state through the noise prediction network in the diffusion model to obtain a trained diffusion model, includes:
[0095] In the forward noise addition stage, Gaussian noise is gradually added to the action sequence according to the preset noise scheduling strategy to generate noise action sequences at different noise levels.
[0096] In the reverse denoising stage, the noise action sequence and the observation environment are used as inputs. The loss value is determined based on the predicted noise output by the noise prediction network. The parameters of the diffusion model are updated to reduce the loss value to the loss threshold, and the trained diffusion model is obtained.
[0097] Specifically, the noise prediction network adopts the U-Net architecture. Its input includes the current noise action sequence, observation environment data (such as the ordered point sequence of the flexible linear object, the position of the index finger tip, and the target state node), and the time-step embedding vector of the current noise level. The output is the predicted value of the original noise. During training, the network parameters are continuously optimized by minimizing the mean squared error loss function between the predicted noise and the actual added noise. After multiple rounds of iterative training, training stops when the loss function converges and the model's action prediction accuracy on the validation set reaches a preset threshold, resulting in a diffusion model with stable action prediction capabilities. Given the initial observation environment and target shape, this model can gradually generate near-optimal robotic arm end-effector action sequences through a reverse denoising process, achieving precise shape control of flexible linear objects.
[0098] In some embodiments, the step of taking a noisy action sequence and the observation environment as input, determining a loss value based on the predicted noise output by the noise prediction network, and updating the parameters of the diffusion model to reduce the loss value to a loss threshold to obtain a trained diffusion model includes:
[0099] Randomly sample the noisy action sequence, select one time step from all noisy time steps each time, and extract the noisy action sequence, the original action sequence, and the real noise corresponding to that time step.
[0100] The noise action sequence, the embedding vector of the time step, and the observation environment are input into a noise prediction network based on a one-dimensional U-Net structure, and the predicted noise is output.
[0101] Calculate the mean squared error loss between the predicted noise and the actual noise, and update the parameters of the noise prediction network through the backpropagation algorithm to reduce the mean squared error loss value to the loss threshold, thus obtaining the trained noise prediction network.
[0102] This embodiment constructs an observation environment by effectively fusing the state information of a flexible linear object with the position of the robotic arm's end effector. Combined with the powerful noise prediction and denoising capabilities of the diffusion model, it achieves end-to-end modeling from complex environment perception to precise action sequence generation. This diffusion model-based training method not only fully learns the relationship between action characteristics and environmental constraints during human teaching but also enhances the model's adaptability to the randomness and uncertainty of action sequences through a two-way process of forward noise addition and reverse denoising. This lays a solid model foundation for the real-time and precise manipulation of the flexible linear object by the robotic arm in the subsequent inference stage.
[0103] II. Model Training
[0104] The diffusion model used in this invention is a type of generative model, which aims to recover data from noise through a stepwise denoising process.
[0105] like Figure 3 As shown, the diffusion model consists of a forward noise addition process and a reverse noise reduction process, which correspond to the model's training and inference processes, respectively.
[0106] Specifically, the forward noise addition process involves passing the original data (action sequence) through multiple time steps. (like The process of gradually transforming into pure noise at each time step The process of adding noise to data can be expressed as formula (1).
[0107] (1);
[0108] in, yes The sequence of actions at each time step; yes The sequence of actions at each time step; It is control The hyperparameter of the time step noise level belongs to the pre-defined noise variation schedule. Indicates a Gaussian distribution; Represents the identity matrix; express The noise figure of a time step is used to scale the action sequence of the previous time step during the noise addition process in order to control the intensity of the noise addition. yes The covariance matrix of the time-step noise addition process.
[0109] Since forward noise addition to the data can be considered a Markov process, a given initial action sequence can be obtained through mathematical derivation. and after a specified noise-adding step Noise action sequence after If the relationship is made , , Let the noise figure at time step j be an example. Then we have:
[0110] (2);
[0111] in, Indicates the initial action sequence. This represents the noise figure at time step t. This represents the product of the noise figures from the first time step to the t-th time step. This represents the mean term after scaling the initial action sequence. Its value gradually decreases as the time step t increases, reflecting the continuous reduction of the influence weight of the original action sequence during the noise addition process. This represents the proportion of the original action sequence that is replaced by noise after t steps of noise addition. The larger the value, the stronger the impact of noise on the data. The covariance matrix represents the noise action sequence at time step t, reflecting the degree of dispersion of the data after adding noise.
[0112] In the actual training process, the sampling formula (3) is used to randomly sample the noise action sequence after adding noise. Each time, a time step t is selected from all the noise-added time steps, and the noise action sequence, the original action sequence and the real noise corresponding to that time step are extracted. The noise action sequence, the embedding vector at time step t, and the observation environment are input to the noise prediction network, which outputs a predicted value for the noise. .
[0113] (3);
[0114] in, Represents real noise. ;
[0115] Then, train a neural network. To predict noise, calculate the predicted noise. With real noise The mean squared error loss between the two is used to update the parameters of the noisy prediction network through the backpropagation algorithm. This minimizes the loss function value. Through iterative optimization using a large number of training samples, the noise prediction network gradually learns the denoising rules of action sequences under different noise levels, thereby constructing a diffusion model that can accurately predict noise and generate clear action sequences.
[0116] The goal of training is to minimize the following formula:
[0117] (4);
[0118] in, This represents the loss function, used to measure the difference between the predicted noise and the actual noise in the output of a neural network. Represents the learnable parameters of a neural network. It refers to the observation environment and the target attitude of flexible linear objects. It is used for outputting prediction noise. Neural networks; This represents the initial action sequence for all possible time steps t. and real noise The expected outcome is that by minimizing this loss function, the noise prediction network can accurately learn the mapping from the noisy state to the clear target state.
[0119] like Figure 4 As shown, during training, the noise prediction network of this invention employs a one-dimensional U-Net neural network. U-Net is a convolutional neural network structure for image segmentation tasks, featuring an encoder-decoder architecture. One-dimensional U-Net is well-suited for processing sequential data. The noise prediction network learns a mapping from a noisy state to a clear target state by progressively denoising the input data through perturbation.
[0120] In the process of constructing training samples, the present invention adopts the following strategy:
[0121] The existing vision system is used to perform node-based abstract representation of the state of flexible linear objects, and the position information of the index finger tip during the human hand teaching process is acquired simultaneously. The state nodes of the flexible linear object, the index finger tip position, and the target state nodes of each frame are stitched together to form the observation environment. The trajectory of the index finger tip in 16 consecutive frames is used as the original action sequence and as the training label for diffusion modeling. In each training sample, the input is the observation environment (current frame and previous frame) and the state node of the target in two frames, and the training label is the action sequence formed by the index finger trajectory in that time period.
[0122] In the reverse denoising inference process, from standard Gaussian noise Beginning, at each time step The neural network output is used to predict noise, and denoising sampling is performed.
[0123] (5);
[0124] in, Indicates variance scheduling. Indicates sampling randomness. Through continuous iteration, we eventually obtain... Noiseless action sequence To enable the robot to perform the task.
[0125] In some embodiments, the process of acquiring the state of the flexible linear object and the position of the robotic arm's end effector, constructing the current observation environment, inputting the current observation environment into a trained diffusion model, generating a target action sequence, and controlling the robotic arm to perform manipulation actions based on the target action sequence includes:
[0126] Real-time acquisition of current state images of flexible linear objects, conversion into ordered point sequences, and fusion with the real-time position coordinates of the robotic arm end effector to construct the current observation environment;
[0127] The current observation environment and target state node are input into the trained diffusion model, and the target action sequence of the end effector is generated through a multi-step inverse denoising process. This target action sequence contains the position and orientation information of the end effector in Cartesian space.
[0128] The robotic arm control system drives the joint motors to perform corresponding control actions based on the generated target motion sequence, thereby realizing the deformation and adjustment of the flexible linear object towards the target state.
[0129] This embodiment constructs an observation environment by effectively fusing the state information of a flexible linear object with the position of the robotic arm's end effector. Combined with the powerful noise prediction and denoising capabilities of the diffusion model, it achieves end-to-end modeling from complex environment perception to precise action sequence generation.
[0130] III. Reasoning Stage
[0131] During the inference phase, this invention utilizes a depth camera and a robotic system to achieve real-time acquisition and input construction of the state of a flexible linear object and the position of the robotic arm's end effector. Specifically, the system includes a depth camera mounted on a work platform and a robot end effector position output. This visual sensor can continuously acquire image information of the flexible linear object and perform state modeling of the deformable flexible linear object's current shape based on the previously mentioned visual algorithm, thereby extracting an ordered sequence of points representing the object's curved shape; simultaneously, it reads the position of the robotic arm's end effector in the robotic system and maps it onto the image using a hand-eye calibration matrix.
[0132] Subsequently, the system combines the sequence of flexible object state points acquired in the current frame with the position coordinates of the end effector to form the model input observation environment quantity, which is then input into the pre-trained diffusion model. This model employs an inverse diffusion process to gradually denoise random noise and generate a series of action sequences. This sequence describes the desired trajectory of the robotic arm's end effector in several future time steps in the form of ordered points.
[0133] The robotic arm control system executes corresponding actions point by point according to the sequence, driving the end effector to move along a specified trajectory, thereby achieving gradual manipulation of the flexible linear object. Because the generation process of this action sequence incorporates behavioral strategies learned from human hand teaching, the robot can reproduce shape-adjusting behavior consistent with human operation while maintaining the target shape constraints.
[0134] Through the above methods, the present invention achieves efficient, autonomous, and precise control of flexible linear objects, and can be widely applied to various robot operation scenarios such as cable management, medical catheter laying, and soft material assembly.
[0135] Compared with related technologies for manipulating flexible linear objects, the robot manipulation method based on a diffusion model proposed in this invention has the following significant advantages:
[0136] First, existing methods mostly rely on high-dimensional visual information (such as video frame sequences) for end-to-end modeling. The model training process has extremely high requirements for data volume and label accuracy, and often exhibits insufficient generalization ability when faced with complex deformations and diverse operational tasks. In contrast, this invention reduces the dimensionality of the state of a flexible linear object into a structured, ordered sequence of nodes through visual abstraction, and uses this sequence as the conditional input to the model. While maintaining the complete expression of the geometric features of the object's shape, it significantly reduces the input dimensionality, thereby improving the model training efficiency and generalization ability.
[0137] Secondly, traditional teaching methods generally require complex sensor deployments or robot playback mechanisms. This invention innovatively uses human hand teaching videos and lightweight vision modules such as Media Pipe to automatically extract fingertip trajectories, replacing the acquisition of real robot trajectories. This significantly reduces the physical and time costs of data acquisition, making the manipulation and learning of flexible objects more efficient, practical, and easily scalable. Combining the advantages of conditional diffusion models in modeling complex target distributions, this invention can achieve precise control over the shape of target objects, exhibiting good robustness and real-time performance.
[0138] See Figure 5 This invention provides a robot manipulation device based on a diffusion model, comprising:
[0139] The data acquisition module is used to collect information on the position of the index finger tip and the state information of the flexible linear object during the hand demonstration process to build training samples;
[0140] The model training module is used to train the training samples using the diffusion model. By training the noise prediction network in the diffusion model, the model learns the mapping from a high-noise state to a clear target state, thus obtaining a trained diffusion model. The training process of the diffusion model includes forward noise addition and inverse noise removal.
[0141] The motion control module is used to collect the state of the flexible linear object and the position of the robotic arm end effector, construct the current observation environment, input the current observation environment into the trained diffusion model, generate the target action sequence, and control the robotic arm to perform manipulation actions based on the target action sequence.
[0142] It is evident that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented in this device embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0143] See Figure 6 This invention provides a robot control system based on a diffusion model, comprising:
[0144] At least one processor;
[0145] At least one memory for storing at least one program;
[0146] When the at least one program is executed by the at least one processor, the at least one processor performs the method described above.
[0147] It is evident that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0148] Furthermore, this application also discloses a computer program product or computer program stored in a computer-readable storage medium. A processor of a computer device can read the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the computer device to perform the described method. Similarly, the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0149] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0150] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0151] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0152] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0153] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0154] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0155] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0156] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0157] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A robot manipulation method based on a diffusion model, characterized in that, The method includes the following steps: The method involves collecting information on the position of the index fingertip and the state of a flexible linear object during a human hand demonstration to construct training samples. Specifically, this includes: acquiring video images of the human hand demonstration; modeling the state of the flexible linear object in the video images using visual algorithms; obtaining the centerline of the flexible linear object using a skeleton extraction algorithm; uniformly selecting multiple sampling points on the centerline at preset intervals, and using these sampling points as candidate nodes for an ordered point sequence; using a Gaussian mixture model to treat each candidate node as a centroid of a Gaussian distribution; iteratively adjusting the position of the candidate nodes using an expectation-maximization algorithm to ensure that the candidate nodes conform to the actual shape of the flexible linear object; and smoothing each candidate node to obtain an ordered point sequence representing the curve shape of the flexible linear object. A diffusion model is used to train the training samples. The noise prediction network in the diffusion model learns the mapping from a high-noise state to a clear target state, thus obtaining a trained diffusion model. The training process of the diffusion model includes forward noise addition and inverse noise removal. The state of the flexible linear object and the position of the end effector of the robotic arm are collected to construct the current observation environment. The current observation environment is then input into the trained diffusion model to generate the target action sequence. Based on the target action sequence, the robotic arm is controlled to perform manipulation actions.
2. The method according to claim 1, characterized in that, The method of collecting the position information of the index finger tip and the state information of the flexible linear object during the teaching process to construct training samples includes: Obtain the position information of the index finger tip during the manual teaching process; The ordered point sequence of each flexible linear object, the position of the index finger tip, and the state nodes of the target are stitched together to form the observation environment; The observation environment of the current frame, the observation environment of the previous frame, and the state node of the target are used as inputs, and the action sequence formed by the trajectory of the index finger at the corresponding moment is used as training labels to form training samples.
3. The method according to claim 2, characterized in that, The acquisition of the index finger tip position information during the human hand teaching process includes: A hand tracking module was used to identify key points of the hand skeleton in the video image to obtain the position information of the index finger tip during the hand teaching process.
4. The method according to claim 1, characterized in that, The process involves training the training samples using a diffusion model, and learning the mapping from high-noise states to clear target states through the noise prediction network in the diffusion model, resulting in a trained diffusion model, including: In the forward noise addition stage, Gaussian noise is gradually added to the action sequence according to the preset noise scheduling strategy to generate noise action sequences at different noise levels. In the reverse denoising stage, the noise action sequence and the observation environment are used as inputs. The loss value is determined based on the predicted noise output by the noise prediction network. The parameters of the diffusion model are updated to reduce the loss value to the loss threshold, and the trained diffusion model is obtained.
5. The method according to claim 4, characterized in that, The process involves taking a noisy action sequence and the observation environment as input, determining the loss value based on the predicted noise output by the noise prediction network, and updating the parameters of the diffusion model to reduce the loss value to a loss threshold, thereby obtaining a trained diffusion model. This includes: Randomly sample the noisy action sequence, select one time step from all noisy time steps each time, and extract the noisy action sequence, the original action sequence, and the real noise corresponding to that time step. The noise action sequence, the embedding vector of the time step, and the observation environment are input into a noise prediction network based on a one-dimensional U-Net structure, and the predicted noise is output. Calculate the mean squared error loss between the predicted noise and the actual noise, and update the parameters of the noise prediction network through the backpropagation algorithm to reduce the mean squared error loss value to the loss threshold, thus obtaining the trained noise prediction network.
6. The method according to claim 1, characterized in that, The process involves acquiring the state of the flexible linear object and the position of the robotic arm's end effector to construct the current observation environment. This environment is then input into a trained diffusion model to generate a target action sequence. Based on this sequence, the robotic arm is controlled to perform manipulation actions, including: Real-time acquisition of current state images of flexible linear objects, conversion into ordered point sequences, and fusion with the real-time position coordinates of the robotic arm end effector to construct the current observation environment; The current observation environment and target state node are input into the trained diffusion model, and the target action sequence of the end effector is generated through a multi-step inverse denoising process. This target action sequence contains the position and orientation information of the end effector in Cartesian space. The robotic arm control system drives the joint motors to perform corresponding control actions based on the generated target motion sequence, thereby realizing the deformation and adjustment of the flexible linear object towards the target state.
7. A robot control device based on a diffusion model, characterized in that, include: The data acquisition module is used to collect information on the position of the index finger tip and the state information of the flexible linear object during the hand demonstration process to build training samples; Specifically, it is used to: acquire video images during the human demonstration process; model the state of a flexible linear object in the video image based on a visual algorithm; obtain the centerline of the flexible linear object through a skeleton extraction algorithm; uniformly select multiple sampling points on the centerline at preset intervals, and use these sampling points as candidate nodes for an ordered point sequence; use a Gaussian mixture model to treat each candidate node as a centroid of a Gaussian distribution; iteratively adjust the position of the candidate nodes using an expectation-maximization algorithm to make the candidate nodes fit the actual shape of the flexible linear object; and smooth each candidate node to obtain an ordered point sequence representing the curve shape of the flexible linear object. The model training module is used to train the training samples using the diffusion model. By training the noise prediction network in the diffusion model, the model learns the mapping from a high-noise state to a clear target state, thus obtaining a trained diffusion model. The training process of the diffusion model includes forward noise addition and inverse noise removal. The motion control module is used to collect the state of the flexible linear object and the position of the robotic arm end effector, construct the current observation environment, input the current observation environment into the trained diffusion model, generate the target action sequence, and control the robotic arm to perform manipulation actions based on the target action sequence.
8. A robot control system based on a diffusion model, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor performs the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a processor-executable program, characterized in that, The processor-executable program, when executed by the processor, is used to perform the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Object grabbing method based on de-noising diffusion model
CN118024253A
Robot motion trajectory planning method and device and robot
CN118404590A