Robot action generation method based on three-view strategy diffusion field and related device
Through the diffusion field method based on the three-view strategy, the robot's action distribution is generated using multi-view RGB-D images and multi-modal information, which solves the problems of perception limitations, insufficient flexibility in action generation and high computational complexity in the prior art, and achieves the effects of high precision, fast action generation and multi-modal information processing.
Patent Information
- Application Number
- CN202510454796.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has challenges in robot perception, action generation and computing efficiency, including limitations in environmental perception, insufficient flexibility in action generation, high computational complexity, and insufficient multimodal information fusion capabilities.
The diffusion field method based on the three-view strategy is adopted to generate the robot action distribution through the three-view projection module, the three-view converter module and the denoising network module. This method uses multi-view RGB-D images to perform three-view feature mapping, fuses visual, language and robot state information, and directly decodes the action distribution through a lightweight denoising network to reduce the iteration process.
It improves perception accuracy and action generation speed, enhances operation accuracy and multimodal information processing capabilities, reduces computing resource requirements and model training complexity, and is suitable for complex tasks and real-time operations.
Smart Images

Figure CN120056126A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of robot perception, decision-making and control, and particularly relates to a robot action generation method and related device based on a three-viewpoint policy diffusion field. Background Art
[0002] Multi-task robot operation is an important research direction in the current field of robotics, and is widely applied in fields such as industrial manufacturing, logistics sorting, and home services. These tasks require robots to possess comprehensive capabilities of environmental perception, decision-making planning, and precise execution.
[0003] Currently, with the improvement of task complexity and scenario diversity, existing technical solutions still face many challenges in aspects such as perception, action generation, and computational efficiency. Specifically, they are manifested as follows:
[0004] Environmental perception has limitations: Most existing robot operation technologies rely on global scene perception and lack the ability to focus on the geometric details of task-related areas. This global perception method is easily interfered with in complex or dynamic environments, resulting in the robot's difficulty in obtaining key task information, thereby affecting the operation accuracy.
[0005] Insufficient flexibility in action generation: Existing traditional action generation methods usually rely on discretized spatial models or implicit representations. Although they can adapt to certain single tasks, they show poor adaptability when dealing with multi-modal inputs or generating diverse actions.
[0006] High computational complexity: Although existing diffusion model-based methods have the ability to generate continuous action distributions, they need to generate target actions through multiple iterative denoising processes in high-dimensional spaces. Their high computational complexity severely restricts the real-time requirements in practical applications, especially in tasks that require quick responses.
[0007] Insufficient multi-modal information fusion ability: With the introduction of multi-modal inputs (such as vision, language, touch, etc.) in robot operation, how to effectively fuse this information has become a challenge. Existing technologies usually cannot balance information richness and computational efficiency, resulting in unsatisfactory task decision-making effects. Summary of the Invention
[0008] The purpose of the present invention is to provide a robot action generation method and related device based on a three-viewpoint policy diffusion field to solve one or more technical problems existing in existing technical solutions in terms of perception accuracy, computational efficiency, and action generation speed.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] In the first aspect of the present invention, a robot action generation method based on a three-view strategy diffusion field is provided, including the following steps:
[0011] Obtain a task instruction and multi-view RGB-D images observed by the robot;
[0012] According to the obtained task instruction and multi-view RGB-D images, use a pre-trained action generation model to generate actions and obtain a robot action distribution;
[0013] Among them, the action generation model includes:
[0014] A three-view projection module for inputting multi-view RGB-D images and performing three-view feature mapping processing to obtain three-view features;
[0015] A three-view transformer module for inputting three-view features and task instructions and performing feature fusion processing to obtain fused three-view features;
[0016] A denoising network module for inputting the fused three-view features and performing step-by-step denoising processing to obtain a robot action distribution.
[0017] In a further improvement of the present invention, in the three-view projection module, the steps of inputting multi-view RGB-D images and performing three-view feature mapping processing to obtain three-view features include:
[0018] Map the three-dimensional point cloud features to three orthogonal two-dimensional view feature planes, and use a convolutional neural network to encode each two-dimensional view feature plane to generate a high-dimensional representation to obtain three-view features;
[0019] Among them, the three orthogonal two-dimensional view feature planes are respectively a top view, a side view, and a front view;
[0020] The generation step of each two-dimensional view feature plane is to reconstruct a three-dimensional point cloud from the multi-view RGB-D images, and project the features of the three-dimensional point cloud by max-pooling to the two-dimensional plane. The specific formula is as follows:
[0021] V WD (x,z) = MaxPool(P(x,y,z));
[0022] V DH (y,z) = MaxPool(P(x,y,z));
[0023] V HW (x,y) = MaxaPool(P(x,y,z));
[0024] Wherein, V represents the planar feature obtained by projection; W, D, and H are the width, height, and depth of the two-dimensional view feature plane respectively; MaxPool represents max pooling; x, y, and z represent the normalized discrete three-dimensional coordinates;
[0025] The high-dimensional representation generated by encoding each two-dimensional view feature plane using a convolutional neural network is:
[0026] F WD ∈R W×D×C ;
[0027] F DH ∈R D×H×C ;
[0028] F HW ∈R H×W×C ;
[0029] Wherein, F represents the encoded feature map; C represents the number of feature channels.
[0030] A further improvement of the present invention lies in that in the three-view transformer module, the steps of executing the input three-view features and the task instruction and performing feature fusion processing to obtain the fused three-view features include:
[0031] Concatenate the three-view features with the multi-modal features of the multi-modal input containing the task instruction, process the concatenated features using the standard Transformer architecture, and perform layer normalization and adaptive adjustment to enhance the modeling ability of the robot state information to obtain the fused three-view features;
[0032] Among them, three-dimensional rotation position encoding is added to the three-view features before concatenation, and the formula is:
[0033] PE3D(i,j,k) = cat(PE1D(i),PE1D(j),PE1D(k));
[0034] Wherein, PE3D(·) represents the three-dimensional rotation position encoding function; i, j, and k are the spatial indices of the feature plane respectively; cat represents vector concatenation; PE1D(·) represents the one-dimensional rotation position encoding function.
[0035] A further improvement of the present invention lies in that in the denoising network module, the steps of executing the input fused three-view features and performing step-by-step denoising processing to obtain the robot action distribution include:
[0036] Input the fused three-view features and use the interpolation method to obtain the latent representation of the 6-degree-of-freedom action distribution at any point in space; among them, at any position (h, w, d) in the three-dimensional space, sample the corresponding features from the three-view feature plane, and the formula is:
[0037] c h,w,d = A(S(F HW , h, w), S(F DH , d, h), S(F WD , w, d));
[0038] In the formula, c is the obtained feature vector; A is summation; S is a bilinear interpolation sampling function, and h, w, and d represent continuous three-dimensional coordinates;
[0039] Conditioned on the latent representation of the 6-DOF action distribution at an arbitrary point, a lightweight denoising network is used to generate the target action distribution through an iterative process. The formula is:
[0040]
[0041] In the formula, a t-1 represents the noisy 6-DOF action at step t - 1; a t represents the noisy 6-DOF action at step t; ∈ θ represents the denoising network, α t and σ t are the noise intensity schedules, t represents the denoising time step, c is the latent diffusion field representation, and δ is the random noise.
[0042] A further improvement of the present invention is that the lightweight denoising network is specifically a multi-layer perceptron.
[0043] A further improvement of the present invention is that when generating the target action distribution through an iterative process, the initial action is sampled from a Gaussian noise distribution.
[0044] A further improvement of the present invention is that the training steps of the action generation model include:
[0045] Obtain a training sample set; wherein, each training sample in the training sample set includes: an RGB-D image sample, a task instruction sample, and a 6-DOF action demonstration label;
[0046] Using the obtained training sample set, update the parameters of the action generation model in a supervised training manner, and obtain a trained action generation model when reaching a preset convergence condition; wherein, for the selected training sample, use the RGB-D image sample and the task instruction sample as inputs and obtain a predicted action through the action generation model, and calculate the loss and update the model parameters based on the predicted action and the 6-DOF action demonstration label;
[0047] Among them, the expression of the overall loss function for calculating the loss is:
[0048] L = λ 1 L pose + λ 2 L State ;
[0049] In the formula, L is the overall loss function, and L pose is the action denoising loss for optimizing the action prediction accuracy of the denoising network, and L State is the state prediction loss for predicting the opening and closing state of the end effector; λ 1 , λ 2 is the hyperparameter for adjusting the weights of the two parts of losses L pose , L State .
[0050] In the second aspect of the present invention, a robot action generation system based on a three-view policy diffusion field is provided, including:
[0051] A data acquisition unit for acquiring task instructions and multi-view RGB-D images observed by the robot;
[0052] An action generation unit for generating an action according to the acquired task instructions and multi-view RGB-D images, and obtaining a robot action distribution by using a pre-trained action generation model;
[0053] Among them, the action generation model includes:
[0054] A three-view projection module for inputting multi-view RGB-D images and performing three-view feature mapping processing to obtain three-view features;
[0055] A three-view transformer module for inputting three-view features and task instructions and performing feature fusion processing to obtain fused three-view features;
[0056] A denoising network module for inputting the fused three-view features and performing step-by-step denoising processing to obtain a robot action distribution.
[0057] In the third aspect of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the robot action generation method based on a three-view policy diffusion field according to any one of the first aspects of the present invention.
[0058] In the fourth aspect of the present invention, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the robot action generation method based on a three-view policy diffusion field according to any one of the first aspects of the present invention.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] The present invention provides a robot motion generation method based on a three-view strategy diffusion field, which solves the problems existing in the prior art solutions in terms of perception accuracy, computational efficiency, motion generation speed, etc. Specifically and explanatorily, the present invention sets up a three-view projection module, enabling the robot to focus on the key regions related to the task, thereby capturing more refined geometric information and enhancing the operation accuracy; subsequent experiments show that the present invention significantly improves the success rate of high-precision tasks in multi-task operation scenarios and can effectively cope with environmental interference and detail requirements in complex tasks. The present invention sets up a small denoising network module, which can directly decode the target motion from the latent diffusion field without the multiple iteration process of the traditional diffusion model, optimizing the motion generation efficiency. This innovation greatly shortens the inference time and provides technical support for real-time operation. The present invention sets up a three-view transformer module, which effectively integrates visual, language, and robot state information, ensuring efficient and reliable motion generation ability even in complex multi-modal input scenarios and enhancing the multi-modal information processing ability.
[0061] The present invention adopts a design that combines explicit and implicit representations, which not only retains the flexible modeling ability for continuous motion distributions but also avoids computational bottlenecks in high-dimensional spaces, and has good scalability and adaptability. The present invention can dynamically adjust the perception and motion generation strategies according to task requirements and performs excellently in industrial robot assembly, domestic service robots, and real-time control tasks. In addition, through lightweight model design, the present invention reduces the computational resource requirements while significantly reducing the complexity of model training and enhancing the deployment efficiency in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art; obviously, the drawings in the following description are some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0063] Figure 1 It is a schematic flowchart of a robot motion generation method based on a three-view strategy diffusion field in an embodiment of the present invention;
[0064] Figure 2 It is a schematic framework structure diagram of a motion generation model in an embodiment of the present invention;
[0065] Figure 3 It is a schematic processing logic diagram of a motion generation model in an embodiment of the present invention;
[0066] Figure 4 It is a schematic diagram of the comparison result of the running time between the method of the present invention and different prior methods in an embodiment of the present invention;
[0067] Figure 5 It is a schematic diagram of a robot action generation system based on a three-view strategy diffusion field in an embodiment of the present invention. Detailed implementation manners
[0068] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention; obviously, the described embodiments of the technical solutions are part of the embodiments of the present invention, rather than all of the embodiments.
[0069] Based on the technical solutions disclosed in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0070] Please refer to Figure 1 , an embodiment of the present invention provides a robot action generation method based on a three-view strategy diffusion field, including the following steps:
[0071] Step 1, obtaining a task instruction and multi-view RGB-D images observed by the robot;
[0072] Step 2, according to the task instruction and multi-view RGB-D images obtained in Step 1, using a pre-trained action generation model to generate an action and obtain a robot action distribution;
[0073] Among them, the action generation model includes:
[0074] A three-view projection module, configured to input multi-view RGB-D images and perform three-view feature mapping processing to obtain three-view features;
[0075] A three-view transformer module, configured to input three-view features and a task instruction and perform feature fusion processing to obtain fused three-view features;
[0076] A denoising network module, configured to input the fused three-view features and perform step-by-step denoising processing to obtain a robot action distribution. In a further exemplary preferred technical solution, in the inference stage, an initial action can be sampled from a Gaussian noise distribution, and the final robot action distribution can be generated through the denoising network module; in addition, the generated action can be further executed by a motion planner (such as BiRRT) to ensure that the robot can complete high-precision tasks in multi-task operations.
[0077] In the technical solution disclosed in the embodiment of the present invention, the tri-view projection module maps the three-dimensional point cloud into three sets of orthogonal feature planes through the tri-view projection method, achieving efficient spatial perception ability and enhancing the detail capture of task-related regions; the tri-view transformer module combines vision, language instructions, and robot state perception to generate a potential diffusion field, completing the deep fusion of multi-modal information; the denoising network module can directly decode the 6-degree-of-freedom action distribution from the potential diffusion field through a lightweight denoising network, without the iterative denoising process of the traditional diffusion model, significantly improving the inference efficiency. Summarily, the technical solution of the embodiment of the present invention realizes the rapid generation of a 6-degree-of-freedom action distribution aligned with the observation space by decomposing the three-dimensional point cloud into three sets of orthogonal view feature planes and combining a lightweight denoising network, significantly improving the operation accuracy and inference efficiency. In summary, the embodiment of the present invention develops a robot operation framework that can accurately perceive task-related regions, support multi-modal information fusion, and have the ability to quickly generate actions, which is of great significance for improving the multi-task operation ability of robots and meeting the requirements of complex application scenarios.
[0078] Please refer to Figure 2 and Figure 3 In a specific embodiment of the present invention, in the tri-view projection module, the steps of executing the input multi-view RGB-D image and performing tri-view feature mapping processing to obtain tri-view features include:
[0079] (1) Map the three-dimensional point cloud features to three orthogonal two-dimensional view feature planes: top view, side view, and front view; wherein, each view feature plane is generated by the following steps: reconstruct the three-dimensional point cloud from the multi-view RGB-D image, and the point cloud contains the geometric coordinates and other features (such as color and depth, etc.) of each point; project the features of the three-dimensional point cloud by maximum pooling onto the two-dimensional plane, and the specific formula is as follows:
[0080] V wD (x,z) = MaxPool(P(x,y,z));
[0081] V DH (y,z) = MaxPool(P(x,y,z));
[0082] V HW (x,y) = MaxPool(P(x,y,z));
[0083] In the formula, V represents the three-plane features obtained by projection, E, D, and H are the width, height, and depth of the view feature plane respectively, MaxPool represents maximum pooling, and x, y, and z represent the normalized discrete three-dimensional coordinates;
[0084] Through this step, each plane retains the key geometric features and effectively reduces the computational complexity.
[0085] To further extract the deep information of the three-view features, the embodiment of the present invention uses a convolutional neural network (CNN) to encode each view feature plane to generate a high-dimensional representation:
[0086] F WD ∈R W×D×C ;
[0087] F DH ∈R D×H×C ;
[0088] F HW ∈R H×W×C ;
[0089] In the formula, F represents the encoded feature map; W, D, and H are the width, height, and depth of the view feature plane respectively; C represents the number of feature channels;
[0090] These high-dimensional representations provide sufficient geometric context information for subsequent action generation; in addition, through an efficient point cloud feature projection mechanism, the three-view features can effectively capture the geometric details of the local task area and adapt to diverse operation scenarios.
[0091] Please refer to Figure 2 and Figure 3 , in a specific embodiment of the present invention, the steps for the three-view transformer module to execute the input three-view features and task instructions and perform feature fusion processing to obtain the fused three-view features include:
[0092] Perform multi-modal feature fusion on the three-view features and multi-modal inputs (such as visual features, task instructions, and robot state information) to generate a unified multi-modal representation; specifically, the steps are as follows: 1) RGB-D visual feature extraction: Use a convolutional network to extract global visual features from RGB-D images; 2) Language instruction encoding: Use a pre-trained CLIP language encoder to encode language instructions into text features; 3) Encode robot state information (such as end effector position and attitude) into features through a multi-layer perceptron; 4) Concatenate the above multi-modal features with the three-view feature plane to form a unified input representation.
[0093] To enhance the model's understanding of the three-dimensional space, the embodiment of the present invention adds three-dimensional rotation position encoding (3D RoPE) to the three-view features before concatenation, and the specific implementation is:
[0094] PE3D(i,j,k) = cat(PE1D(i),PE1D(j),PE1D(k));
[0095] Wherein, i, j, and k are respectively the spatial indices of the feature plane, PE1D(·) represents a one-dimensional rotational position encoding function, PE3D(·) represents a three-dimensional rotational position encoding function, cat represents vector concatenation. Through three-dimensional position encoding, the model can retain the spatial context relationship between features, improving the accuracy and flexibility of action generation.
[0096] In the embodiment of the present invention, the fused features are processed by a three-view transformer. The transformer adopts a standard Transformer architecture and enhances the modeling ability of the robot state information through layer normalization and Adaptive LayerNorm. The finally output fused features provide a unified representation of multi-modal data, supporting subsequent action generation.
[0097] Please refer to Figure 2 and Figure 3 , in a specific embodiment of the present invention, in the denoising network module, the steps of performing input fused three-view features and gradually denoising to obtain the robot action distribution include:
[0098] Input the fused three-view features and use the interpolation method to obtain the latent representation of the 6-degree-of-freedom action distribution at any point in space; wherein, at any position (h, w, d) in the three-dimensional space, sample the corresponding features from the three-view feature plane, and the formula is:
[0099] c h,w,d = A(S(F HW , h, w), S(F DH , d, h), S(F WD , w, d));
[0100] Wherein, S is a bilinear interpolation sampling function, A is summation, c is the obtained feature vector, and h, w, and d represent continuous three-dimensional coordinates, used to integrate the three-view features into a unified latent representation; explanatorily, the latent diffusion field explicitly models the action distribution, providing accurate spatial features for subsequent action generation.
[0101] Taking the latent representation of the 6-degree-of-freedom action distribution at any point as a condition, a lightweight denoising network MLP (multi-layer perceptron) is used to generate the target action distribution through an iterative process. The specific formula is,
[0102]
[0103] Wherein, a t-1 represents the noisy 6DoF action at step t - 1; a t represents the noisy 6DoF action at step t; ∈ θ represents the denoising network, α t and σ tFor noise intensity scheduling, t represents the denoising time step, c represents the latent diffusion field, and δ represents random noise; this process can generate a high-precision 6-degree-of-freedom motion distribution within a small number of time steps.
[0104] In a specific embodiment of the present invention, the steps of network training and optimization include:
[0105] Obtain a training sample set, where each training sample includes: an RGB-D image sample, a task instruction sample, and a 6-degree-of-freedom motion demonstration label;
[0106] Using the obtained training sample set, in a supervised training manner, update the parameters of the motion generation model until a preset convergence condition is reached to obtain a trained motion generation model;
[0107] Explanatorily, during supervised training, for the selected training sample, the RGB-D image sample and the task instruction sample are used as inputs and the predicted motion is obtained through the model. Based on the predicted motion and the 6-degree-of-freedom motion demonstration label, the loss is calculated and the model parameters are updated. Specifically, exemplarily, the model uses the AdamW optimizer, the initial learning rate is 2.5×10 -4 , the total number of training steps is 30k, and by sampling the time step t multiple times, the stability and performance of model training are improved.
[0108] In the embodiment of the present invention, the training objectives include the following two loss functions:
[0109] The motion denoising loss is used to optimize the motion prediction accuracy of the denoising network,
[0110] L pose = E ∈,t [|∈ - ∈ θ (a t |t, c)| 2 ;
[0111] In the formula, E represents the mathematical expectation, ∈ represents the randomly sampled noise, ∈ θ represents the denoising network, a t represents the 6DoF motion with noise at the t-th step, t represents the denoising time step, and c represents the latent diffusion field representation;
[0112] The state prediction loss is used to predict the opening and closing state of the end effector,
[0113]
[0114] In the formula, BCE represents binary cross-entropy, a open and are the predicted end state and the true end state respectively.
[0115] The final loss function is:
[0116] L = λ 1 L pose + λ 2 L State ;
[0117] where λ 1 , λ 2 is a hyperparameter used to adjust the weights of the two parts of the loss.
[0118] In the specific embodiments of the present invention, in order to comprehensively evaluate the performance of the multi-task robot operation framework based on the three-viewpoint strategy diffusion field proposed in the embodiments of the present invention (hereinafter referred to as "the present invention"), extensive experiments were carried out in the RLBench simulation environment and real-world tasks, and the Average Success Rate (ASR) and inference efficiency were used as evaluation indicators. The following is a detailed analysis of the experimental results:
[0119] Analysis of real-world task performance: Table 1 shows the performance of the present invention in 6 real-world tasks and compares it with existing methods. These tasks include high-precision operation tasks such as "Put Fruit", "Push Buttons", and "Stack Blocks". The experimental results show that: in the "Put Fruit" task, the present invention achieved a 100% success rate; in the "Push Buttons" task, the success rate of the present invention reached 100%; in the "Stack Blocks" task, although the scenario is complex and requires high-precision operation, the present invention still achieved a 70% success rate; these results indicate that the present invention has strong generalization ability and stability when dealing with complex and diverse task scenarios in the real world.
[0120] Table 1. Results on hand-collected real-world data
[0121]
[0122] Multi-task Performance Analysis on the RLBench Dataset: Table 2 shows the comparison results of 18 multi-task experiments conducted by the present invention and various existing methods in the RLBench simulation environment. The experiments show that the present invention exhibits excellent performance in multi-task scenarios: the average success rate is 87.3%, which is 5.9 percentage points higher than the existing best method (RVT-2). In the "Push Buttons" task, the present invention achieved a success rate of 98.4%, which is 14.4% higher than that of 3D DiffuserActor. In the "Stack Blocks" task, the success rate of the present invention reached 75.2%, which is 19.3 percentage points higher than the existing methods. These results prove that the present invention can not only efficiently handle diverse task requirements but also maintain stable performance when the task complexity increases.
[0123] Table 2. Comparison results of different methods on 18 common tasks in the RLBench dataset
[0124]
[0125] Inference Efficiency Analysis: Figure 4 Shows the comparison of the inference efficiency between the present invention and other diffusion model methods. The results show that: the inference speed of the present invention is about 10 times faster than the existing diffusion model methods, while maintaining a high success rate. This result fully proves that the design of the hybrid explicit and implicit representation of the present invention greatly improves the inference efficiency and makes it more suitable for real-time operation scenarios.
[0126] Based on the above experimental results, the following conclusions can be drawn:
[0127] Excellent performance in real-world tasks: In multiple complex task scenarios, the success rate of the present invention is significantly better than the existing methods, especially reaching 100% in tasks such as "Place Fruit" and "Push Buttons";
[0128] Strong multi-task learning ability: In the RLBench dataset, the average success rate of the present invention is 87.3%, significantly exceeding the existing best method and performing outstandingly in high-complexity tasks;
[0129] High inference efficiency: Compared with the existing diffusion models, the present invention significantly reduces the inference time while maintaining the high-precision action generation ability, making it suitable for real-time robot control scenarios;
[0130] In summary, the embodiments of the present invention achieve a high-precision, multi-task, and high-efficiency robot operation framework through an innovative three-view strategy diffusion field design, which has broad application prospects and economic benefits. Through the three-view projection technology, this framework decomposes the three-dimensional point cloud into three sets of orthogonal feature planes and uses a three-view transformer to generate a potential diffusion field aligned with the observation space, thereby efficiently representing the probability distribution of 6-degree-of-freedom actions. The framework innovatively combines the advantages of explicit and implicit representations, realizes action sampling through a lightweight denoising network, simplifies the iterative denoising process of traditional diffusion models, and takes into account both inference efficiency and operation accuracy. The main technical features of the embodiments of the present invention include: a three-view projection module, which can significantly improve the spatial perception ability; a three-view transformer module, which fuses visual, language, and robot state information; and a denoising network module, which supports action decoding at any resolution. The framework of the present invention performs excellently in the RLBench benchmark test, with the average success rate increased by 5.9% and the inference speed increased by ten times compared with existing methods. Its design can be widely applied to multi-task learning of industrial robots, operation in complex scenarios, and real-time control, having important practical application value and economic benefits.
[0131] The following is the device embodiment of the present invention, which can be used to execute the method embodiment of the present invention. For details not disclosed in the device embodiment, please refer to the method embodiment of the present invention.
[0132] Please refer to Figure 5 , in the embodiments of the present invention, a robot action generation system based on a three-view strategy diffusion field is provided, including:
[0133] A data acquisition unit, configured to acquire a task instruction and multi-view RGB-D images observed by the robot;
[0134] An action generation unit, configured to generate an action according to the acquired task instruction and multi-view RGB-D images by using a pre-trained action generation model, and obtain a robot action distribution;
[0135] Wherein, the action generation model includes:
[0136] A three-view projection module, configured to input multi-view RGB-D images and perform three-view feature mapping processing to obtain three-view features;
[0137] A three-view transformer module, configured to input three-view features and a task instruction and perform feature fusion processing to obtain fused three-view features;
[0138] A denoising network module, configured to input the fused three-view features and perform step-by-step denoising processing to obtain a robot action distribution.
[0139] In an embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function. The processor described in the embodiment of the present invention can be used to execute the operations of the robot action generation method based on the three-viewpoint strategy diffusion field.
[0140] In an embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed Random Access Memory (RAM) or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the robot action generation method based on the three-viewpoint strategy diffusion field in the above embodiment.
[0141] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) that contain computer-usable program code.
[0142] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in Figure 1 one or more of the processes Figure 1 or a plurality of processes and / or blocks.
[0143] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means, and the instruction means implements the functions specified in Figure 1 one or more of the processes Figure 1 or a plurality of processes and / or blocks.
[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the processes Figure 1 or a plurality of processes and / or blocks.
[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of the present invention or make equivalent substitutions, and any modification or equivalent substitution that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A robot motion generation method based on a three-view strategy diffusion field, characterized in that: The following steps are involved: Obtain multi-view RGB-D images of task instructions and robot observations; According to the acquired task instructions and multi-view RGB-D images, the pre-trained action generation model is used to generate actions and obtain the robot action distribution; Wherein, the action generation model includes: A three-view projection module is used to input multi-view RGB-D images and perform three-view feature mapping processing to obtain three-view features; The three-view converter module is used to input three-view features and task instructions and perform feature fusion processing to obtain fused three-view features; The denoising network module is used to input the fused three-view features and perform step-by-step denoising to obtain the robot action distribution.
2. According to claim 1, a robot motion generation method based on a three-view strategy diffusion field is characterized in that: In the three-view projection module, the steps of inputting a multi-view RGB-D image and performing three-view feature mapping processing to obtain three-view features include: Map the 3D point cloud features to three orthogonal 2D view feature planes, use a convolutional neural network to encode each 2D view feature plane to generate a high-dimensional representation, and obtain three-view features; Among them, the three orthogonal 2D view feature planes are top view, side view and front view; The generation step of each 2D view feature plane is to reconstruct the 3D point cloud from the multi-view RGB-D image and project the feature of the 3D point cloud to the 2D plane using the maximum pooling method. The specific formula is as follows: V WD (x,z)=MaxPool(P(x,y,z)); V DH (y,z)=MaxPool(P(x,y,z)); V HW (x,y)=MaxPool(P(x,y,z)); Where V represents the plane feature obtained by projection; W, D, and H are the width, height, and depth of the two-dimensional view feature plane, respectively; MaxPool represents maximum pooling; x, y, and z represent normalized discrete three-dimensional coordinates; The high-dimensional representation generated by encoding each two-dimensional view feature plane using a convolutional neural network is: F WD ∈R W×D×C ; F DH ∈R D×H×C ; F HW ∈R H×W×C ; Where F represents the encoded feature map; C represents the number of feature channels.
3. The robot motion generation method based on the three-view strategy diffusion field according to claim 1 is characterized in that: In the three-view converter module, the steps of inputting three-view features and task instructions and performing feature fusion processing to obtain the fused three-view features include: The three-view features are concatenated with the multimodal features of the multimodal input containing the task instructions, and the concatenated features are processed using a standard Transformer architecture. The modeling capability of the robot state information is enhanced through layer normalization and adaptive adjustment to obtain the fused three-view features. Among them, the three-dimensional rotation position encoding is added to the three-view features before splicing, and the formula is: PE3D(i,j,k)=cat(PE1D(i),PE1D(j),PE1D(k)); Where PE3D(·) represents the three-dimensional rotation position encoding function; i, j, k are the spatial indices of the feature planes, respectively; cat represents vector concatenation; PE1D(·) represents the one-dimensional rotation position encoding function.
4. The robot motion generation method based on the three-view strategy diffusion field according to claim 1 is characterized in that: In the denoising network module, the steps of inputting the fused three-view features and performing step-by-step denoising to obtain the robot action distribution include: Input the fused three-view features and use the interpolation method to obtain the potential representation of the 6-DOF motion distribution at any point in the space; among them, at any position (h, w, d) in the three-dimensional space, the corresponding features are sampled from the three-view feature plane, and the formula is: c h,w,d =A(S(F HW ,h,w),S(F DH ,d,h),S(F WD ,w,d)); Where c is the obtained eigenvector; A is the sum; S is the bilinear interpolation sampling function, h, w and d represent continuous three-dimensional coordinates; Taking the potential representation of the 6-DOF action distribution of any point as a condition, a lightweight denoising network is used to generate the target action distribution through an iterative process, and the formula is: In the formula, a t-1 represents the noisy 6DoF action at step t-1; a t represents a 6DoF action with t steps of noise; ∈ θ represents the denoising network, α t and σ t is the noise intensity schedule, t represents the denoising time step, c is the potential diffusion field representation, and δ is the random noise.
5. The method for generating robot motion based on a three-view strategy diffusion field according to claim 4, characterized in that: The lightweight denoising network is specifically a multi-layer perceptron.
6. The method for generating robot motion based on a three-view strategy diffusion field according to claim 4, characterized in that: When generating the target action distribution through an iterative process, the initial actions are sampled from a Gaussian noise distribution.
7. The method for generating robot motion based on a three-view strategy diffusion field according to claim 1, characterized in that: The training steps of the action generation model include: Acquire a training sample set; wherein each training sample in the training sample set includes: an RGB-D image sample, a task instruction sample, and a 6-DOF action demonstration label; Using the acquired training sample set, the parameters of the action generation model are updated in a supervised training manner, and a trained action generation model is obtained when the preset convergence condition is reached; wherein, for the selected training samples, RGB-D image samples and task instruction samples are used as input and the predicted action is obtained through the action generation model, and the loss is calculated based on the predicted action and the 6-DOF action demonstration label and the model parameters are updated; Among them, the expression of the overall loss function for calculating the loss is: L=λ1L pose +λ2L State ; Where L is the overall loss function, L pose is the motion denoising loss used to optimize the motion prediction accuracy of the denoising network, L State is the state prediction loss used to predict the opening and closing state of the end effector; λ1 and λ2 are used to adjust L pose , L State Hyperparameters for the weights of the two-part loss.
8. A robot action generation system based on a three-view strategy diffusion field, characterized in that: include: A data acquisition unit, used to acquire task instructions and multi-view RGB-D images observed by the robot; An action generation unit is used to generate actions based on the acquired task instructions and multi-view RGB-D images using a pre-trained action generation model to obtain the robot action distribution; Wherein, the action generation model includes: A three-view projection module is used to input multi-view RGB-D images and perform three-view feature mapping processing to obtain three-view features; The three-view converter module is used to input three-view features and task instructions and perform feature fusion processing to obtain fused three-view features; The denoising network module is used to input the fused three-view features and perform step-by-step denoising to obtain the robot action distribution.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the robot action generation method based on the three-view strategy diffusion field according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the robot action generation method based on the three-view strategy diffusion field as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Water pouring service robot control method based on dynamic model reinforcement learning
CN113031437A
Image generation method and device based on pose guidance, medium and equipment
CN117745956A
Intelligent three-dimensional structure reconstruction method based on three-plane feature representation and visual angle condition diffusion model
CN117893691A
Robot closed-loop joint optimization method and system based on data driving
CN119328777A
Image synthesis using diffusion models created from single or multiple view images
US20240135630A1
Cited By
Robot action generation method and related device
CN120422249A