Robot action generation method based on continuous state decomposition and related device

By introducing a diffusion embedding-based action generation method in the robot system, using the large language model and the SD2Actor model to process continuous states, the problem of lack of fineness and flexibility in robot action generation in the prior art is solved, and higher generalization ability and operation accuracy are achieved.

CN120146198AActive Publication Date: 2025-06-13XI AN JIAOTONG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510326975.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-13
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle continuous states, resulting in a lack of fineness and flexibility in performing targeted actions, especially in complex and dynamic environments.

Method used

The robot action generation method based on diffusion embedding is adopted, language instructions are analyzed through a large language model, language features, scene features and key target state values ​​are embedded and embedded with the SD2Actor model, and refined robot actions are generated through the diffusion model.

Benefits of technology

It improves the generalization ability and operation accuracy of the robot in diversified tasks, can effectively handle continuous states and generate accurate robot operation actions in new states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146198A_ABST
    Figure CN120146198A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent equipment, and discloses a robot action generation method based on continuous state decomposition and a related device. The robot action generation method based on continuous state decomposition comprises the following steps: acquiring a language instruction of a user and scene observation of a robot; analyzing the language instruction through a large language model to obtain a key target state value; and based on the language instruction, the scene observation and the key target state value, performing action generation by utilizing a trained SD2Actor model to obtain a robot action. According to the method, the problem of difficulty in mapping between language instructions and target actions in current robot operation can be solved, continuous states can be effectively processed, accurate robot operation actions can be generated in a new state, and the generalization ability and operation precision of a robot in diversified tasks are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of embodied intelligence, and particularly relates to a robot action generation method and related device based on continuous state decomposition. Background Art

[0002] Manipulation tasks under language conditions, especially manipulation tasks based on continuous target states, mainly involve two aspects: the robot's understanding of the object state and the execution of accurate target actions. This requires the robot system to maintain a mapping from language instructions to precise target states in a continuous world. However, this process often faces the following two challenges, including:

[0003] 1) The robot needs to comprehensively understand the scene, which includes various information such as geometric details, scene layout, and visual appearance. In order to accurately execute tasks, the robot needs to be able to identify and locate various objects in the environment, which is crucial for its subsequent decision-making and action execution.

[0004] 2) The robot needs to execute increasingly refined actions to conform to the desired target state, which means that the robot not only needs to understand language instructions but also needs to have a high degree of flexibility and precision to be able to respond quickly in complex and dynamic environments.

[0005] Currently, with the increase in task complexity, the robot must consider environmental changes and uncertainties when executing target actions to achieve higher operation accuracy. However, existing technical solutions mainly emphasize abstract language instructions, which assume discrete object states and relatively large-range movements, ignoring the continuity and generalization of target states. Summary of the Invention

[0006] The purpose of the present invention is to provide a robot action generation method and related device based on continuous state decomposition to solve one or more of the above-mentioned technical problems. The technical solution disclosed by the present invention is specifically a robot operation solution that realizes continuous state decomposition through diffusion embedding, solves the problem of difficult mapping between language instructions and target actions in current robot operations, can effectively process continuous states, and generate accurate robot operation actions in new states, improving the generalization ability and operation accuracy of the robot in diverse tasks.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] In the first aspect of the present invention, a robot action generation method based on continuous state decomposition is provided, including the following steps:

[0009] Obtain the user's language instructions and the robot's scene observations;

[0010] Parse the language instruction through a large language model to obtain the key target state value;

[0011] Based on the language instruction, the scene observation, and the key target state value, use the trained SD 2 Actor model to generate actions and obtain the robot actions;

[0012] Among them, the SD 2 Actor model includes:

[0013] A feature extraction module for extracting the language features of the language instruction and the scene features of the scene observation;

[0014] A feature fusion module for fusing the language features, the scene features, and the corresponding embeddings of the key target state values by corresponding embedding to obtain the fused features;

[0015] A diffusion model for using the fused features as conditional guidance for the noise reduction process and generating refined robot actions by gradually reducing the noise;

[0016] The SD 2 During the training process of the Actor model, initialize some learnable embeddings, retrieve the corresponding embeddings through the language instruction, fuse them with the language features and interact with the scene features, and then use them as the conditional input of the diffusion model. The corresponding robot actions are predicted through the diffusion model; among them, an iterative update method of supervised training is adopted, and the overall loss function includes the orthogonality loss and the noise reduction loss of the diffusion model;

[0017] The steps for obtaining the corresponding embedding of the key target state value are as follows: if the key target state value is included in the training sample set of the SD 2 Actor model, the corresponding embedding of the key target state value is directly obtained through retrieval; if the key target state value is not included in the training sample set of the SD 2 Actor model, then linearly decompose the key target state value through a large language model, and then obtain the linear combination of the known state embeddings through a linear combination method and use it as the corresponding embedding of the key target state value.

[0018] A further improvement of the present invention lies in

[0019] In the feature fusion module, the steps of fusing the language features, the scene features, and the corresponding embeddings of the key target state values to obtain the fused features include:

[0020] First, perform self-attention calculation on the language features and the embeddings corresponding to the key target state values to obtain the preliminary fused features; then perform cross-attention calculation on the preliminary fused features and the scene features to obtain the fused features corresponding to the states in the scene.

[0021] A further improvement of the present invention lies in that

[0022] The overall loss function is expressed as:

[0023]

[0024] In the formula, represents the overall loss function; represents the denoising loss; represents the orthogonality loss; γ 0 represents the weight coefficient of the orthogonality loss;

[0025]

[0026] In the formula, and respectively represent the losses of the position, rotation, and switch state of the target action;

[0027]

[0028] In the formula, represents the orthogonal loss between embeddings of different tasks; represents the orthogonal loss between embeddings within the same task.

[0029] A further improvement of the present invention lies in that

[0030] During the calculation of the denoising loss,

[0031]

[0032] In the formula, f, represent the prediction network and the diffusion model network. f is used to predict the fixture state, is used to predict translation and rotation; o represents the scene observation; e′ represents the state embedding; represents the translation and rotation amounts at the t - th time step; BCE represents the categorical cross - entropy loss function; a open represents the fixture state, and the fixture state is open or closed.

[0033] A further improvement of the present invention lies in that

[0034] During the calculation of the orthogonality loss,

[0035]

[0036] In the formula, γ 1 、γ 2 represent the weight factors; E i represents the i - th embedding within the task; E meanRepresents the average of all embeddings for each task; I is the identity matrix; σ(·) represents calculating the spectral norm.

[0037] A further improvement of the present invention lies in

[0038] In the step of linearly decomposing the key target state value through the large language model and then obtaining the linear combination of known state embeddings as the embedding corresponding to the key target state value by means of linear combination,

[0039] s novel = λ 1 s 1 + λ 2 s 2 +…+ λ i s i +…+ λ m s m ;

[0040] λ 1 + λ 2 +…+ λ i +…+ λ m = 1;

[0041] In the formula, s novel represents the new target state value, and the embedding corresponding to s novel is not in the training sample set; λ i represents the i-th coefficient obtained by decomposing with the large language model; s i represents the known i-th state value, and the embedding corresponding to s i is in the training sample set; m is the total number of state values;

[0042] e novel = λ 1 e 1 + λ 2 e 2 +…+ λ i e i +…+ λ m e m ;

[0043] In the formula, e novel represents the embedding corresponding to the new target state value obtained by linearly combining embeddings; e i represents the known i-th state embedding, retrieved from the model according to s i .

[0044] A further improvement of the present invention lies in

[0045] The robot action a generated by the diffusion model novel is expressed as:

[0046]

[0047] In the formula, a t is Gaussian noise, and α t represents the denoising coefficient, ∈ θ represents the diffusion network, and σ t represents the Gaussian distribution, and δ represents the noise intensity.

[0048] In the second aspect of the present invention, a robot motion generation system based on continuous state decomposition is provided, including:

[0049] A data acquisition module for acquiring the user's language instructions and the robot's scene observations;

[0050] A state value parsing module for parsing the language instructions through a large language model to obtain key target state values;

[0051] An action generation module for generating robot actions based on the language instructions, the scene observations, and the key target state values by using a trained SD 2 Actor model to obtain robot actions;

[0052] Among them, the SD 2 Actor model includes:

[0053] A feature extraction module for extracting the language features of the language instructions and the scene features of the scene observations;

[0054] A feature fusion module for fusing the language features, the scene features, and the key target state values by corresponding embedding to obtain the fused features;

[0055] A diffusion model for using the fused features as conditional guidance for the denoising process to generate refined robot actions through step-by-step denoising;

[0056] During the training process of the SD 2 Actor model, some learnable embeddings are initialized, the corresponding embeddings are retrieved through the language instructions, and after being fused with the language features and interacting with the scene features, they are used as the conditional input of the diffusion model, and the corresponding robot actions are predicted through the diffusion model; among them, a supervised training iterative update method is adopted, and the overall loss function includes the orthogonality loss and the denoising loss of the diffusion model;

[0057] The obtaining step of the corresponding embedding of the key target state value is as follows: if the key target state value is included in the training sample set of the SD 2 Actor model, the corresponding embedding of the key target state value is directly obtained through retrieval; if the key target state value is not included in the SD 2In the training sample set of the Actor model, the key target state value is linearly decomposed by the large language model, and then through the way of linear combination, a linear combination of known state embeddings is obtained and used as the embedding corresponding to the key target state value.

[0058] In the third aspect of the present invention, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the method for generating robot actions based on continuous state decomposition according to any one of the first aspects of the present invention.

[0059] In the fourth aspect of the present invention, there is provided a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method for generating robot actions based on continuous state decomposition according to any one of the first aspects of the present invention.

[0060] Compared with the prior art, the present invention has the following beneficial effects:

[0061] In the prior art, rough spatial actions are generated without paying attention to the target state. Aiming at the problem of how to determine consistent spatial actions according to specific state information, the present invention specifically discloses a method for robot operation based on continuous state decomposition through diffusion embedding. By using learnable embeddings to represent continuous target state features, the attention to state features is enhanced, and precise action generation in the continuous space is promoted. In addition, aiming at how to solve the generalization problem of new states, the present invention uses the large language model to analyze and extract the state information in the language instruction, and performs linear decomposition and linear combination (affine combination) of the embedding to obtain new state features. This method has the characteristics of high efficiency and convenience and can be generalized to the entire continuous space. The above two improved parts can improve the generalization ability and operation accuracy of the robot in diverse tasks. Summarily, the present invention utilizes the relationship between object states and the relationship between states and language instructions to enhance state representation and promote refined action generation; uses the large language model to parse the state information in the language instruction and generates corresponding actions through the diffusion model; the present invention can effectively process continuous states and generate accurate robot operation actions in new states, improving the generalization ability and operation accuracy of the robot in diverse tasks. Description of the Drawings

[0062] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art; obviously, the drawings in the following description are some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0063] Figure 1 It is a schematic flowchart of a robot motion generation method based on continuous state decomposition in an embodiment of the present invention;

[0064] Figure 2 In an embodiment of the present invention, SD 2 It is a schematic diagram of the processing logic of the Actor model;

[0065] Figure 3 It is a schematic diagram of the general architecture for generalizing to a new state through state decomposition in an embodiment of the present invention;

[0066] Figure 4 It is a schematic diagram of the training and inference architecture of a robot operation method for generalizing a new state through state decomposition in an embodiment of the present invention;

[0067] Figure 5 It is a schematic diagram of the detailed structure of the state decomposition method in an embodiment of the present invention;

[0068] Figure 6 It is a schematic diagram of the constraint of the orthogonal loss between embeddings in an embodiment of the present invention;

[0069] Figure 7 It is a schematic diagram of the result visualization in the ARNOLDBenchmark in an embodiment of the present invention;

[0070] Figure 8 It is a schematic diagram of a robot motion generation system based on continuous state decomposition in an embodiment of the present invention. Detailed implementation manners

[0071] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention; obviously, the described embodiments of the technical solutions are part of the embodiments of the present invention, rather than all of the embodiments.

[0072] Based on the technical solutions disclosed in the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0073] Please refer to Figure 1 , a robot motion generation method based on continuous state decomposition provided by an embodiment of the present invention includes the following steps:

[0074] Step 1, obtain the user's language instruction and the robot's scene observation;

[0075] Step 2, parse the language instruction through a large language model to obtain the key target state value;

[0076] Step 3, based on the language instruction, the scene observation, and the key target state value, use the trained SD 2 Actor model to generate actions, and obtain the robot actions corresponding to the language instruction and the scene observation;

[0077] Among them, the SD 2 Actor model includes:

[0078] A feature extraction module for extracting the language features of the language instruction and the scene features of the scene observation;

[0079] A feature fusion module for fusing the language features, scene features, and state values by corresponding embedding to obtain the fused features; in an exemplary preferred technical solution, when performing feature fusion, first perform self-attention calculation on the embeddings corresponding to the language features and the key target state values to obtain the preliminary fused features; then perform cross-attention calculation on the preliminary fused features and the scene features to obtain the fused features corresponding to the state in the scene; in a further explanatory technical solution, if the key target state value is included in the training sample set of the SD 2 Actor model, the embedding corresponding to the key target state value is directly obtained by retrieval; if the key target state value is not included in the training sample set of the SD 2 Actor model, then linearly decompose the key target state value through a large language model to obtain relevant known state values, and then retrieve the corresponding embeddings according to the state values, and use the linear combination of these known state embeddings as the embedding corresponding to the key target state value;

[0080] A diffusion model for using the fused features as external conditions to guide the noise reduction process and generating refined robot actions by gradually reducing noise.

[0081] In the technical solution disclosed in the embodiments of the present invention, the state decomposition and state combination are uniformly applied to the solution for improving the operation performance of the robot, providing a brand-new idea for the robot to process continuous states. Through the organic combination of state decomposition and state combination, the robot can more flexibly cope with complex and changeable environments and tasks, improving the overall performance and adaptability of the robot system. Specifically, when decomposing the state, the large language model (LLMs) is used to parse the state information in the language instruction. This decomposition method not only improves the accuracy of state recognition but also enhances the robot's ability to understand complex language instructions, enabling the robot to execute user instructions more accurately. When combining the states, the state features obtained by linear combination are used to generate actions adapted to the new state. This combination method can make full use of the information of the known states, quickly generate actions adapted to the new state, and improve the response speed and operation accuracy of the robot; through the unified framework of state decomposition and combination, the robot can better adapt to unseen or changing states, significantly improving its generalization ability in continuous states, which means that the robot can perform excellently in a wider range of scenarios and tasks, reducing its dependence on specific environments and tasks.

[0082] In one embodiment of the present invention, a robot action generation method for continuous state decomposition through diffusion embedding is provided to achieve fine-grained operation of the robot, which specifically includes the following steps:

[0083] Step 1, given a set of demonstration data of robot operations, which includes scene observations, language instructions, and the robot's proprioception. To facilitate training, the key frames are extracted as training data; the main framework of the present invention adopts the structure of a diffusion model to facilitate capturing the distribution of continuous actions. In the training stage, first, CLIP is used to extract scene features and language features; to generate fine-grained actions, the present invention uses learnable embeddings to model the target state of the object and performs self-attention with the language features to obtain a preliminary fusion feature. Subsequently, the preliminary fusion feature and the scene feature are cross-attentioned to obtain a secondary fusion feature as an external condition to guide the diffusion model to generate more fine-grained actions.

[0084] Step 2, to force the model to better distinguish different target states, the present invention introduces an orthogonality loss, adds a soft margin between the embeddings, and forces them to maintain an appropriate distance from each other, so that each embedding has a unique corresponding state value, aligning the embedding space and the continuous action space; finally, using the denoising loss and the orthogonality loss function of the diffusion model, iterative optimization of the network is achieved in the training stage.

[0085] Step 3, under the supervision of Step 1 and Step 2, a trained embedding and diffusion model are obtained. Then, when encountering an instruction containing a new state during the inference phase, the present invention uses a large language model to analyze and extract the corresponding target state, and further linearly decomposes the instruction state. Subsequently, a new feature is obtained by fusing with a linear combination, which is used as a new state condition to interact with the language feature and the scene feature and input into the diffusion model, enabling the diffusion model to generalize to new actions. Finally, by reasonably using the decomposition and combination of states, the correct understanding of each new state is achieved during the inference phase, and the corresponding actions are generated.

[0086] In a specific exemplary technical solution, in Step 1, the specific steps of using the diffusion model and the embedding include:

[0087] Step 1.1, a set of operation task demonstrations are given, including language instructions l, scene observations o, and the agent's proprioceptive perception c. These demonstration data usually come from the actual operations of multiple tasks. Through these demonstrations, the robot can learn the mapping relationship from language instructions to action execution. To convert these data into effective training data that the model can learn, key frames are extracted from each operation task. These key frames are time points in the scene that contain the main information, which can effectively represent the changes in object states and the key steps of the task, reducing redundant data and improving training efficiency.

[0088] Step 1.2, in order to better handle the continuous action state distribution, the framework of the present invention adopts a diffusion model structure to iteratively denoise the noisy action a, which can be approximated as generating complex target data from simple noise by gradually removing noise. Moreover, the diffusion model can model an action space with continuity and high complexity. Especially when facing diverse target states in a task, it can effectively infer the action transitions between different states.

[0089] Step 1.3, in order to improve the accuracy and refinement of action generation, the present invention introduces learnable embeddings to model the target states of objects. Different from traditional methods, the states of objects are not just simple descriptive features, but rich state information is represented by learnable embeddings. These embeddings can be fused with various features of the object, including the scene information where the object is located, language instructions, and proprioceptive perception information, etc., and these information are used as conditions and input into the diffusion model.

[0090] Explanatorily, the state information and the language feature are subjected to self-attention, and after cross-attention with the scene feature, the final state fusion feature is obtained. The specific process is as follows:

[0091]

[0092] In the formula, ⊕ represents a concatenation operation, WK and W V represent the linear transformation matrices of Key and Value; K ⊕ and V ⊕ represent the linear transformation matrices of Key and Value; represents projecting the state representation into the feature space of the language, which is composed of multiple layers of MLP; finally, each group of K and V is fused with the scenario, and noise is predicted through the multi-layer transformers of the diffusion model.

[0093] In a specific embodiment of the present invention, in step 2, the main framework for action generation has been obtained. In order to enhance the model's modeling of different target states and distinguish them, the present invention uses an orthogonality constraint to limit the distance between embeddings, which specifically includes the following steps:

[0094] Step 2.1, the main goal of the present invention is to distinguish the states of tasks, but the distances between the state embeddings within a task and between tasks should be slightly different. To ensure this, the present invention imposes different constraints on the states within a task and between tasks. Specifically, all the states belonging to the same task are taken out respectively, the inner product of the transpose of the task embedding and itself is calculated, and the difference between the result and the identity matrix represents the correlation between the embeddings. Through the activation function, the non-linear relationship is strengthened to obtain the loss within the task.

[0095] Step 2.2, for the calculation of the states of different tasks, since the number of states is large, in order to simplify the network calculation, the mean value of all the states within each task is calculated, and their losses of the state values are calculated using them;

[0096] The specific formula is as follows:

[0097]

[0098] In the formula, γ 1 and γ 2 are weight factors (hyperparameters) that control the influence of the loss on the training process; E i represents the i-th embedding within the task; E mean represents the average value of all the embeddings of each task; I is the identity matrix; σ(E) represents calculating the spectral norm of E; represents the orthogonality loss; represents the orthogonality loss between tasks and within tasks.

[0099] In a specific embodiment of the present invention, in step 3, the state decomposition acts on each new state, and the purpose is to construct and fuse into a new state representation by inferring the combined features between the states, which specifically includes the following steps:

[0100] Step 3.1, in the inference stage, use a large language model (exemplarily, optionally ChatGPT) to analyze and extract the target state information in the new instruction. The large language model automatically identifies the key target state value s in the task by parsing the input natural language instruction. novel After obtaining the embedding of the target state, the linear decomposition of the state is carried out next. Specifically, the target state in the instruction is usually not a simple and fully visible state, but needs to be approximated by the known object state and environmental information. Therefore, by novel decomposing s 1 into a linear combination of multiple known state embeddings {s 2 …s K}, this problem is handled.

[0101] The specific formula is as follows:

[0102] s novel = λ 1 s 1 + λ 2 s 2 + … + λ i s i + … + λ m s m ;

[0103] λ 1 + λ 2 + … + λ i + … + λ m = 1;

[0104] In the formula, s novel represents the new target state value, and the embedding corresponding to s novel is not in the training sample set; λ i represents the i-th coefficient obtained by decomposing with the large language model; s i represents the known i-th state value, and the embedding corresponding to s i is in the training sample set; m is the total number of state values.

[0105] Step 3.2, use a set of known state embeddings {e 1 , e 2 … e K}, and the corresponding actions of these states have been learned through the diffusion model in the training stage. By weighted combination of these known state embeddings, we obtain an approximate representation e novel of the target state. This weighted combination is achieved through affine combination, where the weight coefficient λ i of each known state embedding is provided by the large model. The specific formula is as follows:

[0106] e novel = λ1 e 1 + λ 2 e 2 + … + λ i e i + … + λ m e m ;

[0107] In the formula, e novel represents the embedding corresponding to the new target state value obtained by the embedded linear combination; e i represents the known i-th state embedding, retrieved from the model according to s i

[0108] In this way, not only can the information of the existing states be utilized, but also the generation of the new target state can be ensured to conform to the pattern learned during training.

[0109] Step 3.3, through the above affine combination, the embedding e novel of the new target state is obtained, and the weighted sum of its various state features further enhances the representation ability of the target state, enabling it to cover a wider state space. When the diffusion model receives this new state condition, it can predict the action a novel corresponding to the target state according to the learned action generation rule.

[0110]

[0111] In the formula, a t is Gaussian noise, α t represents the denoising coefficient, ∈ θ represents the diffusion network, σ t represents the Gaussian distribution, and δ represents the noise intensity.

[0112] The above method of the embodiment of the present invention realizes effective generalization in new environments and tasks by combining the process of state decomposition and combination with the action generation model.

[0113] Please refer to Figures 2 to 4 , a robot operation method based on state decomposition provided by the embodiment of the present invention includes the following steps:

[0114] Step 1, given the language instruction l, the scene observation o, and the robot's proprioception c. We extract the actions and observations of adjacent key frames as a set of training data, that is, according to the observation at the current moment and the action to be predicted at the next moment.

[0115] ​To fully integrate scene and instruction information, CLIP is used as the visual and language encoder to obtain their respective feature encodings, and cross-attention is performed among them for feature fusion. To effectively model the action distribution, we adopt the method of diffusion model. Specifically, using the above features as conditions, the real noisy actions are gradually denoised. During this process, two-way interactions are carried out among the instructions, sampled scene features, and actions to obtain a more comprehensive feature, and this action feature passes through multiple layers of MLP to obtain the corresponding predicted noise.

[0116] Step 2, in the training stage, to enhance the modeling of state information and generate refined actions, we initialize some learnable embeddings, retrieve the corresponding embeddings through the state values, and fuse them with the language features and interact with the scene features. The specific approach is as follows:

[0117]

[0118] Among them, ⊕ represents the concatenation operation, W K and W V represent the linear transformation matrices of Key and Value, K ⊕ and V ⊕ represent the linear transformation matrices of Key and Value, represents projecting the state representation into the language feature space, which consists of multiple layers of MLP. The concatenated state features are fed into the diffusion model to further refine action generation.

[0119] Step 3, during the inference process, referring to Figure 5 , a large language model is used to understand and extract the target state information in the new instruction. By parsing the input natural language instruction, the language model can automatically identify and extract the key target state value s novel . After obtaining the embedding representation of the target state, next we perform a linear decomposition on this state. Specifically, we represent the target state as a weighted linear combination of multiple known states {s 1 , s 2 … s m} to better model the complex target state. Specifically, this process can be represented by the following formula:

[0120] s novel = λ 1 s 1 + λ 2 s 2 +…+ λ m s m ;

[0121] λ 1 + λ 2 +…+ λm = 1;

[0122] Among them, λ i represents the decomposition coefficient of the state. We make the sum of these coefficients equal to 1, so that our state combinations are also in the same embedding space.

[0123] Step 4, according to {s 1 , s 2 … s m}, retrieve a corresponding set of known state embeddings {e 1 , e 2 … e m}. These states have learned the corresponding state features through the diffusion model during the training phase. By linearly combining these known state embeddings according to the decomposition coefficients, we obtain an approximate representation e novel of the target state. The specific mathematical expression of this process is as follows:

[0124]

[0125] Among them, e i is the embedding of the i-th known state, and λ i is the decomposition coefficient corresponding to each state embedding, representing the contribution degree of different known states to the target state. Through this affine combination, we can effectively approximate the new state e novel , providing conditional input for subsequent action generation.

[0126] Step 5, during the network optimization training process, perform end-to-end joint training on the above state features and the diffusion model. At the same time, in order to increase the distinguishability of state information and its correspondence with actions, for all states within the same task, refer to Figure 6 , we first calculate the inner product of the transpose of the task embedding with itself, and use the difference between this result and the identity matrix as the correlation metric between the embeddings. Then, we strengthen this non-linear relationship through an activation function to obtain the loss within the task.

[0127]

[0128] Through the above steps, we ensure the effective distinction of states within and between tasks, enable the model to better learn the features of different task states, and at the same time enhance the generalization ability of the model in multiple task environments.

[0129] Combined with the loss of modeling actions by the diffusion model It can be expressed as:

[0130]

[0131] Among them, and respectively represent the losses of the position, rotation, and switch state of the target action; BCE represents the categorical cross-entropy loss function; f and represent the prediction network and the diffusion model network.

[0132] The present invention evaluates the proposed method using the average success rate and evaluates the state decomposition model of the embodiments of the present invention on ARNOLD. This dataset provides camera inputs from 5 perspectives, and the resolution of each input is defaulted to 128×128. The present invention uses a 7-degree-of-freedom Franka Emika Panda manipulator, equipped with a parallel gripper to perform tasks. The dataset contains 40 different objects and 20 diverse scenarios, covering 8 tasks, and the task target states have different variabilities. Each task is divided into a training set, a test set, a novel and an any state set in the original way. The test set only contains the target states observed during the training process, the novel set contains a new, unobserved target state, and the any set contains multiple target states (including observed and unobserved target states), covering the entire operation state space. In this process, the present invention mainly uses the novel and any sets to verify the generalization ability of the model. When a task instance meets the success condition for 2 seconds, the task is considered successful. This requires that the deviation between the current state and the target state is within an acceptable threshold.

[0133] Table 1. Results of the basic performance, new state performance, and continuous state performance on the ARNOLD Benchmark

[0134]

[0135] Table 1 shows the experimental results of different methods used in the ARNOLD simulator. It can be seen from Table 1 that the robot operation method based on state decomposition proposed by the present invention has obvious advantages in terms of basic performance, generalization performance of new states, and generalization performance of continuous states. Figure 7 is the visualization result of the method of the present invention in the ARNOLD experiment. The experimental results show the effectiveness of the relationship state and action alignment ability of the present invention in continuous scenarios. In summary, compared with other methods, the method of the present invention has achieved a great improvement in the robot operation task of state generalization.

[0136] The embodiments of the present invention disclose a robot operation method for realizing continuous state decomposition through diffusion embedding. This technology utilizes the combination of a diffusion model and large language models (LLMs) to achieve accurate mapping from language instructions to robot actions. Its technical solutions include state modeling, state decomposition and combination, action generation, and orthogonality constraints, and can effectively process and generate transitional actions for continuous states. SD 2The Actor supports zero-shot learning and does not require task-specific fine-tuning. It can directly generate robot operation actions suitable for new states, featuring multi-modal fusion and efficient state decomposition and recombination. The technology of this invention is mainly applied to robot operation tasks, human-computer interaction, and education and research fields, and is particularly suitable for handling diverse objects in dynamic environments. SD 2 The Actor significantly improves the generalization ability and operation accuracy of robots in unseen states, reduces the dependence on large-scale training data, improves task execution efficiency, and enhances the practical applicability of the model. The continuous state decomposition method based on diffusion embedding disclosed in the embodiments of this invention is the first work to unify state decomposition and state combination for improving robot operation performance. It analyzes the state information in language instructions through a large language model, decomposes the new state into multiple known states, and generates actions adapted to the new state by affine-combining the decomposed state features. The combination of the two can significantly improve the generalization ability and operation accuracy of robots in continuous states, can be directly integrated into existing robot operating systems, and is trained end-to-end through standard action generation and state prediction losses.

[0137] The following is the device embodiment of this invention, which can be used to execute the method embodiment of this invention. For details not disclosed in the device embodiment, please refer to the method embodiment of this invention.

[0138] Please refer to Figure 8 , in the embodiments of this invention, a robot action generation system based on continuous state decomposition is provided, including:

[0139] A data acquisition module for acquiring the user's language instructions and the robot's scene observations;

[0140] A state value parsing module for parsing the language instructions through a large language model to obtain key target state values;

[0141] An action generation module for generating robot actions based on the language instructions, the scene observations, and the key target state values by using the trained SD 2 Actor model;

[0142] wherein, the SD 2 Actor model includes:

[0143] A feature extraction module for extracting the language features of the language instructions and the scene features of the scene observations;

[0144] A feature fusion module for fusing the language features, the scene features, and the key target state values by corresponding embedding to obtain fused features;

[0145] A diffusion model is used to take the fused features as conditions to guide the noise reduction process, and generate refined robot actions by gradually reducing noise.

[0146] The SD 2 During the training process of the Actor model, some learnable embeddings are initialized. The corresponding embeddings are retrieved through language instructions, fused with language features and interacted with scene features, and then used as the conditional input of the diffusion model. The corresponding robot actions are predicted through the diffusion model. Among them, a supervised training iterative update method is adopted, and the overall loss function includes an orthogonality loss and a noise reduction loss of the diffusion model.

[0147] The step of obtaining the embedding corresponding to the key target state value is as follows: if the key target state value is included in the SD 2 in the training sample set of the Actor model, the embedding corresponding to the key target state value is directly obtained through retrieval; if the key target state value is not included in the SD 2 in the training sample set of the Actor model, the key target state value is linearly decomposed by a large language model, and then through a linear combination method, a linear combination of known state embeddings is obtained and used as the embedding corresponding to the key target state value.

[0148] In an embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program. The computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used to execute the operations of the robot action generation method based on continuous state decomposition.

[0149] In an embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in a computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the operating system of the terminal. Moreover, one or more instructions suitable for being loaded and executed by the processor are stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM (Random Access Memory) or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the robot action generation method based on continuous state decomposition in the above embodiment.

[0150] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, optical memories, etc.) containing computer-usable program codes.

[0151] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0152] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and this instruction device implements the functions in the process Figure 1One or more processes and / or boxes Figure 1 The functions specified in one or more boxes.

[0153] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 One or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes.

[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A robot motion generation method based on continuous state decomposition, characterized in that: The following steps are involved: Obtain the user's language instructions and the robot's scene observations; Parsing the language instruction through a large language model to obtain a key target state value; Based on the language instructions, the scene observations and the key target state values, the trained SD 2 The Actor model generates actions to obtain robot actions; Among them, the SD 2 The Actor model includes: A feature extraction module, used to extract language features of language instructions and scene features of scene observations; The feature fusion module is used to fuse the language features, scene features and key target state values ​​into corresponding embeddings to obtain fused features; A diffusion model is used to guide the noise reduction process by taking the fused features as conditions, and to generate refined robot actions by gradually reducing noise; The SD 2 During the training of the Actor model, some learnable embeddings are initialized, and the corresponding embeddings are retrieved through language instructions. After being fused with language features and interacting with scene features, they are used as conditional inputs of the diffusion model, and the corresponding robot actions are predicted through the diffusion model. In this process, an iterative update method of supervised training is adopted, and the overall loss function includes orthogonalization loss and denoising loss of the diffusion model. The key target state value corresponding to the embedded acquisition step is: if the key target state value is included in SD 2 In the training sample set of the Actor model, the corresponding embedding of the key target state value is directly obtained by retrieval; if the key target state value is not included in the SD 2 In the training sample set of the Actor model, the key target state value is linearly decomposed through the large language model, and then the linear combination of the known state embeddings is obtained through linear combination and used as the corresponding embedding of the key target state value.

2. A robot motion generation method based on continuous state decomposition according to claim 1, characterized in that: In the feature fusion module, the steps of embedding and fusing language features, scene features, and key target state values ​​to obtain fused features include: First, the embeddings corresponding to the language features and key target state values ​​are self-attentioned to obtain preliminary fusion features. Then, the preliminary fusion features are cross-attentioned with the scene features to obtain the fused features of the corresponding states in the scene.

3. The method for generating robot motion based on continuous state decomposition according to claim 1, characterized in that: The overall loss function is expressed as: In the formula, represents the overall loss function; represents the noise reduction loss; represents the orthogonalization loss; γ0 represents the weight coefficient of the orthogonalization loss; In the formula, and Represent the losses of position, rotation and switch state of the target action respectively; In the formula, represents the orthogonal loss between embeddings of different tasks; represents the orthogonal loss between embeddings within the same task.

4. A robot motion generation method based on continuous state decomposition according to claim 3, characterized in that: In the calculation process of noise reduction loss, In the formula, f, represents the prediction network and diffusion model network, f is used to predict the fixture state, Used to predict translation and rotation; o represents scene observation; e ′ Representation state embedding; represents the translation and rotation at time step t; BCE represents the classification cross entropy loss function; a open Indicates the fixture status, which is open or closed.

5. The method for generating robot motion based on continuous state decomposition according to claim 3, characterized in that: In the calculation process of orthogonalization loss, In the formula, γ1 and γ2 represent weight factors; E i represents the i embeddings within the task; E mean represents the average of all embeddings for each task; I is the identity matrix; σ(·) represents the calculation of the spectral norm.

6. The method for generating robot motion based on continuous state decomposition according to claim 1, characterized in that: In the step of linearly decomposing the key target state value through the large language model, and then obtaining the linear combination of the known state embeddings through linear combination and using them as the corresponding embedding of the key target state value, s novel =λ1s1+λ2s2+…+λ i s i +…+l m s m ; λ1+λ2+…+λ i +…+l m =1; In the formula, s novel represents the new target state value, s novel The corresponding embedding is not in the training sample set; λ i represents the i-th coefficient decomposed by the large language model; s i represents the known i-th state value, s i The corresponding embedding is in the training sample set; m is the total number of state values; e novel =λ1e1+λ2e2+…+λ i e i +…+l m e m ; In the formula, e novel represents the embedding corresponding to the new target state value obtained by embedding linear combination; e i represents the known i-th state embedding, according to s i Retrieved from the model.

7. A robot motion generation method based on continuous state decomposition according to claim 6, characterized in that: Robot actions generated by diffusion model a novel It is expressed as: In the formula, a t is Gaussian noise, α t represents the denoising coefficient, ∈ θ represents the diffusion network, σ t represents Gaussian distribution, and δ represents noise intensity.

8. A robot motion generation system based on continuous state decomposition, characterized in that: include: Data acquisition module, used to obtain the user's language instructions and the robot's scene observation; A state value parsing module, used to parse the language instruction through a large language model to obtain a key target state value; The action generation module is used to generate an action based on the language instruction, the scene observation and the key target state value using the trained SD 2 The Actor model generates actions to obtain robot actions; Among them, the SD 2 The Actor model includes: A feature extraction module, used to extract language features of language instructions and scene features of scene observations; The feature fusion module is used to fuse the language features, scene features and key target state values ​​into corresponding embeddings to obtain fused features; A diffusion model is used to guide the noise reduction process by taking the fused features as conditions, and to generate refined robot actions by gradually reducing noise; The SD 2 During the training of the Actor model, some learnable embeddings are initialized, and the corresponding embeddings are retrieved through language instructions. After being fused with language features and interacting with scene features, they are used as conditional inputs of the diffusion model, and the corresponding robot actions are predicted through the diffusion model. In this process, an iterative update method of supervised training is adopted, and the overall loss function includes orthogonalization loss and denoising loss of the diffusion model. The key target state value corresponding to the embedded acquisition step is: if the key target state value is included in SD 2 In the training sample set of the Actor model, the corresponding embedding of the key target state value is directly obtained by retrieval; if the key target state value is not included in the SD 2 In the training sample set of the Actor model, the key target state value is linearly decomposed through the large language model, and then the linear combination of the known state embeddings is obtained through linear combination and used as the corresponding embedding of the key target state value.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the robot motion generation method based on continuous state decomposition according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the robot motion generation method based on continuous state decomposition according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Robot task generation method and device based on pre-training language model and medium

    CN116402164A

  • Robot task learning and evaluation method and device under variable language instruction

    CN117972430A

  • Exhibition hall robot visual language navigation method based on large model

    CN119309580A

  • Robot device, operation preparation device and method, control program and recording medium

    JP2002331481A