Body execution model action generation method and system based on large-scale diffusion coding and decoding

Through hierarchical training of large-scale diffusion codec models, combined with Stable Diffusion and Diffusion Policy, the problem of high cost of data set acquisition of embodied execution models and insufficient generalization capabilities is solved, and the smoothness and generalization capabilities of robot action generation are improved.

CN120372183APending Publication Date: 2025-07-25HARBIN INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510522271.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing embodied execution models are expensive to collect data sets and lack generalization capabilities, resulting in poor smoothness and generalization performance of robot action generation, especially in new scenarios and new tasks.

Method used

The embodied execution model of large-scale diffusion codec is adopted, and the combination of large-scale diffusion encoder and small-scale diffusion decoder is used to perform hierarchical training, generating high-dimensional action coded vectors and decoding into continuous and smooth action trajectories.

Benefits of technology

It improves the smoothness and generalization ability of robot action generation, can train efficiently in the absence of data sets, adapt to diverse scenarios, and improves the success rate of robots in different tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372183A_ABST
    Figure CN120372183A_ABST
Patent Text Reader

Abstract

The invention provides a body execution model action generation method and system based on large-scale diffusion coding and decoding. The method comprises the following steps: step 1, preprocessing an original data set, including cleaning and denoising of data; 2, constructing a body execution model based on large-scale diffusion coding and decoding, and training the model by using the preprocessed data set; and step 3, generating an action track by using the trained model. The large-scale diffusion model used in the method can map the multi-modal environment information into the high-dimensional semantic vector, so that the small-scale diffusion model can decode and generate the action sequence according to the semantic vector information. According to the method, a small-scale diffusion decoder is used, so that the generated action track has better effects in the aspects of smoothness and continuity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of embodied intelligence technology, and particularly to an action generation method and system for an embodied execution model based on large-scale diffusion encoding and decoding. It is applied to improve the ability of the current embodied intelligence large model to generate intelligent robot execution actions, enabling the robot to perform tasks more generally and commonly. Background Art

[0002] Embodied execution is a key technology to make robots truly "move", which affects the safety, aesthetics, and rationality of robot operation. Through the embodied execution technology, a robot can generate precise rotation angles for controlling each joint servo, enabling the robot to achieve different postures in space, and thus complete autonomous actions such as walking, standing, grasping, and carrying. The core of this technology lies in the precise understanding and control of the robot's multi-joint system. However, due to the extremely complex structure of the robot body, it often contains more than a dozen or even dozens of servos, which are responsible for the rotation or movement of each joint or component respectively. It is extremely difficult for an execution model to simultaneously control these servos and generate a smooth and physically constrained action sequence.

[0003] The current embodied execution methods are mainly divided into two categories. One category is to use the autoregressive large model technology with strong current general capabilities and reasoning capabilities to generate robot execution actions end-to-end. By means of imitation learning, the large model fits various discretized action trajectories, endowing the embodied intelligence large model with the ability to control the robot to complete various tasks. The other category is to use a diffusion model that is good at generating continuous space information, and through reinforcement learning or imitation learning, the model continuously learns and fits the actions of performing a specific task, enabling the model to master specific skills at a relatively low cost.

[0004] The former uses an end-to-end large model and is trained with a large number of action datasets in real scenarios. Therefore, when facing different tasks, the large model can generate corresponding actions and has a certain generalization ability. However, the cost of collecting datasets in the field of embodied intelligence is extremely high. The movement of real-scenario robots requires real-time human operation to complete, which consumes a large amount of manpower. During the data collection process of real robots, failures often occur, resulting in damage to the robots. Repairing the robots and continuously re-collecting data lead to huge consumption of funds and time in the entire collection process. In addition, considering the characteristics that the dataset itself should have multiple scenarios and multiple tasks, various scenarios also need to be built for the robots in reality, which also consumes a large amount of time and manpower. The lack of embodied intelligence-related datasets caused by the high cost makes the embodied intelligence large model far from being able to be trained with a sufficient amount of datasets like the large language models in traditional natural language processing, computer vision fields, and image-text multimodal large models. There are significant problems in aspects such as the generalization performance of the embodied intelligence large model in generating robot actions, the smoothness, and the success rate of task execution. Moreover, the fundamental reason for the lack of datasets is difficult to effectively solve in a short time.

[0005] To address the above problems of the first type of method, the second type of method, that is, the idea of using a diffusion model to generate robot execution actions, realizes that only a small amount of datasets can be used to generate smoother and more successful robot actions on specific tasks. The reason is that the process of the diffusion model itself in restoring Gaussian noise can better fit the information in the continuous space. While the autoregressive large model discretizes the robot actions, it will cause loss of accuracy and affect the action effect. In addition, compared with the autoregressive model, the diffusion model is also easier to train, enabling the diffusion model to master a certain skill with only a small amount of human demonstration data. However, there are still many deficiencies in this method nowadays. One is that the model scale of the existing methods based on the diffusion model is small; the other is that a method similar to overfitting is adopted in training. Although the model can well reproduce the demonstration actions, it hardly has generalization ability and the model cannot be used in new scenarios and new tasks. Summary of the Invention

[0006] The object of the present invention is to solve the problems in the prior art, and a method and system for generating actions of an embodied execution model based on large-scale diffusion encoding and decoding are proposed. The method of the present invention will fully combine the technical advantages of the two types of methods in the background technology, that is, the large model scale makes the generalization ability strong, and the diffusion decoding makes the actions smoother, and a new idea of a hierarchical diffusion model is proposed.

[0007] The present invention is realized through the following technical solutions. The present invention proposes a method for generating actions of an embodied execution model based on large-scale diffusion encoding and decoding, and the method includes:

[0008] Step 1: Preprocess the original dataset, including data cleaning and denoising;

[0009] Step 2: Construct an embodied execution model based on large-scale diffusion encoding and decoding, and use the preprocessed dataset to train the model;

[0010] Step 3: Use the trained model to generate action trajectories.

[0011] Furthermore, the embodied execution model based on large-scale diffusion encoding and decoding consists of a large-scale diffusion encoder and a small-scale diffusion decoder. The large-scale diffusion encoder is responsible for encoding environmental information and mapping it to a high-dimensional vector space containing action semantics. The small-scale diffusion decoder is responsible for decoding this high-dimensional vector space into action trajectories.

[0012] Furthermore, the large-scale diffusion encoder is a large-scale action encoder based on Stable Diffusion; during the encoding process, the original action vector first undergoes a step-by-step noise addition operation and becomes a noise vector that does not contain any information at all. Then, through the denoising process of Stable Diffusion guided by images and text instructions, it is restored to a new high-dimensional vector.

[0013] Furthermore, in the inference stage of the encoding process, Stable Diffusion independently restores the noise vector to the action encoding vector in the current scene, thereby achieving the purpose of using Stable Diffusion as an action encoder to extract action semantic features and generate action encodings.

[0014] Furthermore, the small-scale diffusion decoder is a small-scale action decoder based on Diffusion Policy; this action decoder adopts the diffusion model in the open-source work Diffusion Policy. By injecting the action encoding vector with rich action meanings generated by StableDiffusion as guiding information into the diffusion model of DiffusionPolicy, it assists in generating real action trajectory parameters.

[0015] Furthermore, in the decoding process, the training and inference processes are specifically as follows: during training, input the real action trajectory and the high-dimensional action encoding generated by the encoder into the embodied execution model, add noise to the action trajectory, and let the embodied execution model refer to the high-dimensional action encoding to denoise and restore the action trajectory; during inference, refer to the high-dimensional action encoding and restore the noise information to the action trajectory.

[0016] Furthermore, during the model training process, a step-by-step training strategy is adopted to train the embodied execution model. First, the decoder is frozen, and the encoder is independently trained to optimize the parameters of the encoder, enabling it to fully learn the feature representation of the input information, encoding the multi-modal information into high-dimensional encoded vectors to ensure accurate capture of task-related information. After that, the focus is on training the decoder to learn to gradually denoise from the high-dimensional semantic embeddings of the encoder and generate continuous, smooth, and accurate motion trajectories that meet the task requirements. Finally, the two are jointly trained as a whole to further optimize the collaborative work of the encoder and decoder.

[0017] The present invention also proposes an action generation system for an embodied execution model based on large-scale diffusion encoding and decoding. The system includes:

[0018] A preprocessing module: preprocesses the original dataset, including data cleaning and denoising;

[0019] A construction and training module: constructs an embodied execution model based on large-scale diffusion encoding and decoding and trains the model using the preprocessed dataset;

[0020] A generation module: generates action trajectories using the trained model.

[0021] The present invention also proposes an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the action generation method for the embodied execution model based on large-scale diffusion encoding and decoding are implemented.

[0022] The present invention also proposes a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the action generation method for the embodied execution model based on large-scale diffusion encoding and decoding are implemented.

[0023] Compared with the prior art, the beneficial effects of the present invention are:

[0024] (1) Stronger action representation and generation capabilities. Compared with large models with an autoregressive architecture, the principle of Gaussian denoising makes the diffusion model itself more suitable for understanding and generating data such as robot execution action trajectories that have continuity (temporal continuity, spatial continuity) and non-discrete information representation. Therefore, the small-scale diffusion decoder used in the method of the present invention can make the generated action trajectories have better effects in terms of smoothness and continuity.

[0025] (2) Stronger generalization ability for action generation. It is one of the important reasons why current autoregressive large model technologies have strong generalization ability that they can map the semantics of various modal information into a high-dimensional vector space, and then achieve the understanding and reasoning of high-level semantics by fusing and processing these high-dimensional vectors. Similar to autoregressive large models, the large-scale diffusion model used in the method of the present invention can map multi-modal environmental information into high-dimensional semantic vectors, enabling the small-scale diffusion model to decode and generate action sequences based on this semantic vector information. By learning the semantic representation of such high-dimensional vectors, the model can capture the connections between action semantics in various environments and tasks. Thus, when facing new environments and new tasks, the model can also generate new corresponding high-dimensional action semantic vectors, making the generalization ability of the model of the present invention stronger. In addition, since diffusion models are easier to train than autoregressive models, in the current situation where relevant datasets for embodied intelligence are lacking, training diffusion models of the same scale is more efficient than autoregressive models, which is beneficial to subsequent technical implementation and industrialization. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0027] Figure 1 It is an architecture diagram of an embodied execution model based on large-scale diffusion encoding and decoding according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0029] The main task of embodied execution is to convert the results of perception and planning into specific action behaviors, enabling the robot to flexibly and autonomously complete multi-step tasks in the real environment, which directly affects the adaptability and flexibility of the robot's action behaviors in the real environment and is the core for the embodied intelligence to realize its application value.

[0030] Combined with Figure 1 , the present invention proposes an action generation method for an embodied execution model based on large-scale diffusion encoding and decoding, and the method includes:

[0031] Step 1: Preprocess the original dataset, including data cleaning and denoising;

[0032] Step 2: Construct an embodied execution model based on large-scale diffusion encoding and decoding, and train the model using the preprocessed dataset;

[0033] Step 3: Generate action trajectories using the trained model.

[0034] The embodied execution model based on large-scale diffusion encoding and decoding consists of a large-scale diffusion encoder and a small-scale diffusion decoder. The large-scale diffusion encoder is responsible for encoding environmental information and mapping it to a high-dimensional vector space containing action semantics. The small-scale diffusion decoder is responsible for decoding the high-dimensional vector space into action trajectories.

[0035] The large-scale diffusion encoder is a large-scale action encoder based on Stable Diffusion; during the encoding process, the original action vector first undergoes a step-by-step noise addition operation and becomes a noise vector that does not contain any information. Then, through the denoising process of Stable Diffusion with image and text instruction guidance, it is restored to a new high-dimensional vector. During the inference stage of the encoding process, Stable Diffusion independently restores the noise vector to the action encoding vector in the current scene, thereby achieving the purpose of using Stable Diffusion as an action encoder to extract action semantic features and generate action encodings.

[0036] Specifically, the method of the present invention uses the currently open-source Stable Diffusion model, which, after fine-tuning, serves as the encoding model for action trajectories and is responsible for encoding and compressing action trajectory data into high-dimensional encoding vectors containing action semantic information.

[0037] Furthermore, in the training phase, the method of the present invention uses the open-source model Octo based on the Transformer architecture, taking the pictures seen by the current robot and the received text instructions as inputs to generate high-dimensional encoded vectors containing action semantics. To enhance the feature extraction ability of the Octo model when encoding actions, the method of the present invention expands the parameter scale of the Octo model and conducts more sufficient training on it. The action encoded vector generated by the expanded model is the vector that Stable Diffusion in this method needs to learn to generate. Consistent with the noise addition and denoising process of the classical diffusion model, this vector first undergoes a step-by-step noise addition operation and becomes a noise vector that does not contain any information at all, and then undergoes a denoising process guided by images, text instructions, etc. by StableDiffusion to be restored to a new high-dimensional vector. At the same time, due to the L2 regularization loss function set during the training process to evaluate the distribution difference between the new vector and the original action vector, StableDiffusion can continuously learn the distribution of the original action vector, and the newly generated high-dimensional vector is almost the same as the action vector generated by the Octo model in terms of the action semantics it represents, thereby enabling Stable Diffusion to obtain the ability to generate high-dimensional action encoded vectors.

[0038] In the inference phase, the Octo model will no longer participate in this process. Since information such as images and text instructions are injected during the denoising process of Stable Diffusion, it can refer to this information and independently restore the noise vector to the action encoded vector in the current scenario, thereby achieving the purpose of using Stable Diffusion as an action encoder to extract action semantic features and generate action encodings.

[0039] Since the powerful ability of Stable Diffusion itself in understanding image information and generating high-dimensional semantic vectors can be transferred to the VLA (image-text-action) task, and the diffusion model is easier and more efficient to train, the encoder designed by this method has more excellent effects when encoding action semantic information and can enhance the overall robot action generation effect.

[0040] The small-scale diffusion decoder is a small-scale action decoder based on Diffusion Policy; this action decoder adopts the diffusion model in the open-source work Diffusion Policy. By using the action encoding vectors with rich action meanings generated by Stable Diffusion as guiding information, it is injected into the diffusion model of Diffusion Policy to assist in generating real action trajectory parameters. During the decoding process, the training and inference processes are as follows: During training, the real action trajectory and the high-dimensional action encoding generated by the encoder are input into the embodied execution model. The action trajectory is added with noise, and the embodied execution model refers to the high-dimensional action encoding to denoise and restore the action trajectory; during inference, referring to the high-dimensional action encoding, the noise information is restored to the action trajectory.

[0041] During the model training process, the method described in the present invention uses the largest open-source robot operation dataset Open X-Embodiment so far for model training. At the same time, in order to improve the data quality, the original dataset is preprocessed, including data cleaning and denoising, removing datasets in the original Open X-Embodiment that do not contain image streams, have too high a degree of duplication, too low image resolution, or involve overly specialized tasks; and data reallocation. According to the diversity of tasks and environments, different weights are assigned to each dataset. Then, zero-padding is used for the missing camera channels to ensure the consistency of data dimensions, and the gripper action spaces in different datasets are aligned.

[0042] After that, using the entire preprocessed dataset, the embodied execution model is trained using a step-by-step training strategy; gradually improving the performance of the encoder and decoder. First, freeze the decoder and independently train the encoder to optimize the parameters of the encoder so that it can fully learn the feature expressions of the input information, encoding multi-modal information (such as images, actions, etc.) into high-dimensional encoding vectors to ensure that task-related information can be accurately captured; then, focus on the training of the decoder so that it can learn to denoise step by step from the high-dimensional semantic embedding of the encoder to generate continuous, smooth, and accurate motion trajectories that meet the task requirements; finally, jointly train the two as a whole to further optimize the collaborative work of the encoder and decoder, making the entire model more coordinated, thereby improving the accuracy and success rate of task execution, enhancing its generalization ability in diverse scenarios, and ultimately creating an efficient and highly adaptable embodied execution model.

[0043] The present invention also proposes an embodied execution model action generation system based on large-scale diffusion encoding and decoding, and the system includes:

[0044] A preprocessing module: preprocess the original dataset, including data cleaning and denoising;

[0045] Construction and training module: Construct an embodied execution model based on large-scale diffusion encoding and decoding, and train the model using the preprocessed dataset;

[0046] Generation module: Use the trained model to generate action trajectories.

[0047] Through the use of an action encoder based on the Stable Diffusion model, the present invention can effectively improve the model's performance in representing action semantics and quickly and accurately obtain action encoding vectors; through the use of the high-quality Open X-Embodiment dataset for training, the embodied execution model can be fully trained, effectively improving the model's ability to generate action trajectories; through the use of a distributed training strategy, both the encoder and decoder of the embodied execution model can be independently and fully trained, enabling them to be combined with high adaptability without destroying the original image understanding ability of the Stable Diffusion model.

[0048] The method of the present invention can be applied to any scenario where a robot needs to perform behavioral actions, such as a robot walking, running, jumping, grasping an object, operating an object, navigating, patrolling, etc. It widely affects the effects of various types of robots performing various tasks and is of great significance for the research and development and application of current intelligent robots, enabling robots to replace humans to complete various tasks such as home services, industrial manufacturing, medical care, and special production, promoting the low-cost, intelligent, and general-purpose development of robots.

[0049] The present invention also proposes an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method for generating actions of the embodied execution model based on large-scale diffusion encoding and decoding.

[0050] The present invention also proposes a computer-readable storage medium for storing computer instructions, and when the computer instructions are executed by a processor, they implement the steps of the method for generating actions of the embodied execution model based on large-scale diffusion encoding and decoding.

[0051] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DRRAM). It should be noted that the memory of the method described in the present invention is intended to include but not limited to these and any other suitable types of memory.

[0052] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a high-definition digital video disc (DVD)), or a semiconductor medium (such as a solid state disc (SSD)), etc.

[0053] In the implementation process, the steps of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by the hardware processor or executed by the combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0054] It should be noted that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in software form. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0055] The above has introduced in detail the method and system for generating actions of the embodied execution model based on large-scale diffusion encoding and decoding proposed by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. An action generation method for an embodied execution model based on large-scale diffusion encoding and decoding, characterized in that The method includes the following steps: Step 1: Preprocess the original dataset, including data cleaning and denoising; Step 2: Construct an embodied execution model based on large-scale diffusion encoding and decoding, and use the preprocessed dataset to train the model; Step 3: Use the trained model to generate action trajectories.

2. The method according to claim 1, wherein The embodied execution model based on large-scale diffusion encoding and decoding consists of a large-scale diffusion encoder and a small-scale diffusion decoder. The large-scale diffusion encoder is responsible for encoding environmental information and mapping it to a high-dimensional vector space containing action semantics. The small-scale diffusion decoder is responsible for decoding the high-dimensional vector space into action trajectories.

3. The method according to claim 2, wherein The large-scale diffusion encoder is a large-scale action encoder based on StableDiffusion; during the encoding process, the original action vector first undergoes a step-by-step noise addition operation and becomes a noise vector that does not contain any information at all. Then, through the denoising process of Stable Diffusion guided by images and text instructions, it is restored to a new high-dimensional vector.

4. The method according to claim 3, characterized in that In the inference stage of the encoding process, StableDiffusion independently restores the noise vector to the action encoding vector in the current scene, so as to achieve the purpose of using StableDiffusion as an action encoder to extract action semantic features and generate action encodings.

5. The method according to claim 4, wherein The small-scale diffusion decoder is a small-scale action decoder based on DiffusionPolicy; this action decoder adopts the diffusion model in the open-source work Diffusion Policy. By injecting the action encoding vector with rich action meanings generated by Stable Diffusion as guiding information into the diffusion model of Diffusion Policy, it assists in generating real action trajectory parameters.

6. The method according to claim 5, wherein In the decoding process, the training and inference processes are specifically as follows: during training, input the real action trajectory and the high-dimensional action encoding generated by the encoder into the embodied execution model, add noise to the action trajectory, and let the embodied execution model refer to the high-dimensional action encoding to denoise and restore the action trajectory; during inference, refer to the high-dimensional action encoding and restore the noise information to the action trajectory.

7. The method according to claim 6, characterized in that, During the model training process, a step-by-step training strategy is adopted to train the embodied execution model; first, freeze the decoder and independently train the encoder to optimize the parameters of the encoder, enabling it to fully learn the feature expressions of the input information, encode multi-modal information into high-dimensional encoding vectors, and ensure that it can accurately capture task-related information; After that, focus on the training of the decoder, enabling it to learn to gradually denoise from the high-dimensional semantic embedding of the encoder and generate continuous, smooth, and accurate motion trajectories that meet the task requirements; finally, conduct overall joint training on the two to further optimize the collaborative work of the encoder and the decoder.

8. An embodied execution model action generation system based on large-scale diffusion encoding and decoding, characterized in that The system includes: Preprocessing module: Preprocess the original dataset, including data cleaning and denoising; Construction and training module: Construct an embodied execution model based on large-scale diffusion encoding and decoding, and use the preprocessed dataset to train the model; Generation module: Use the trained model to generate action trajectories.

9. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, the steps of the method according to any one of claims 1-7 are implemented.