Model training method and device, training data enhancement method and device and electronic equipment

By combining real and simulated data to train the embodied world model, training data containing language commands, videos, and motion trajectories is generated, solving the problem of insufficient training data for embodied intelligent agents and achieving efficient generation of high-quality training data.

CN121963301APending Publication Date: 2026-05-01PAXINI TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PAXINI TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2025-12-22
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies struggle to generate large amounts of high-quality training data for embodied intelligent agents. Real data collection is costly and has low scalability, simulation data performs poorly in the real world, and internet video data lacks motion trajectories and has numerous licensing issues.

Method used

By combining real and simulated data, hypothetical instructions are generated from the initial frames of real videos to train an embodied world model. This generates training data containing language instructions, videos, and motion trajectories, reducing the gap between simulation and reality and improving data quality and scalability.

Benefits of technology

It provides a large amount of high-quality training data for intelligent agents, improves the efficiency and quality of training data generation, and enhances the training effect and robustness of embodied world models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963301A_ABST
    Figure CN121963301A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method, a training data enhancement method and device and electronic equipment, and the method comprises the steps: obtaining real data which comprises a real instruction, a real video and a real motion track; simulation data are obtained, the simulation data comprise a hypothetical instruction, a simulation video and a simulation action track, and the hypothetical instruction is generated according to an initial frame of a real video; according to the real data and the simulation data, the initial body-equipped world model is trained, a trained body-equipped world model is obtained, and the trained body-equipped world model is used for generating training data containing language instructions, videos and action tracks for model training of the intelligent agent. Therefore, by improving the training effect of the body world model and applying the body world model to the training data generation of the intelligent agent, the generation efficiency and data quality of the training data of the intelligent agent are improved, and the problem of providing a large amount of high-quality training data for the intelligent agent is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Model training methods, training data augmentation methods, devices and electronic equipment Technical Field

[0001] This application relates to the field of embodied intelligence technology, and in particular to a model training method, a training data augmentation method, an apparatus, and an electronic device. Background Technology

[0002] Embodied intelligence refers to the ability of an intelligent agent to learn and perform tasks through interaction with the physical environment. It emphasizes that intelligent agents not only need to rely on abstract computational models, but also on multimodal perception, kinematic principles, and environmental interaction to achieve intelligent behavior in recognizing the surrounding environment, learning task-related knowledge, and performing tasks.

[0003] In embodied intelligence, the training data for agents comes in three forms: real data acquired through teleoperation of the agent, simulation data generated in a simulation environment, and massive amounts of internet video data. Real data has high quality but high acquisition costs and low scalability; simulation data has high scalability but is affected by the differences between the simulation world and the real world, resulting in agents trained on simulation data performing poorly in the real world; internet video data is the most abundant but lacks motion trajectories, making it difficult to use directly for agent training.

[0004] Therefore, how to generate a large amount of high-quality training data for intelligent agents is an urgent problem to be solved. Summary of the Invention

[0005] The main objective of this application is to propose a model training method, a training data augmentation method, a device, and an electronic device that at least solves the problem of how to generate a large amount of high-quality training data for embodied intelligent agents.

[0006] In a first aspect, this application provides a model training method, comprising: acquiring real data, the real data including real instructions, real videos, and real motion trajectories, wherein the real videos and real motion trajectories are obtained by data collection during the process of an agent executing the real instructions in a real environment; acquiring simulation data, the simulation data including hypothetical instructions, simulation videos, and simulation motion trajectories, wherein the hypothetical instructions are generated based on the initial frames of the real videos, and the simulation videos and simulation motion trajectories are obtained by simulating the agent executing the hypothetical instructions; training an initial embodied world model based on the real data and the simulation data to obtain a trained embodied world model, wherein the initial embodied world model includes an initial video generation model and an initial motion generation model, and the trained embodied world model includes a trained video generation model and a trained motion generation model; wherein the trained embodied world model is used to generate training data containing language instructions, videos, and motion trajectories for model training of the agent.

[0007] Secondly, this application provides a training data augmentation method, comprising: acquiring real video frames from the original training data of an agent's model, wherein the real video frames are obtained by video capture of the agent executing real instructions in a real environment; generating hypothetical instructions based on the real video frames, wherein the hypothetical instructions belong to task instructions that the agent supports executing in the task scenario described in the real video frames; inputting the real video frames and the hypothetical instructions into a trained embodied world model, generating hypothetical videos and hypothetical action trajectories through the trained embodied world model, wherein the trained embodied world model is trained according to the model training method described in the first aspect; and determining the hypothetical instructions, the hypothetical videos, and the hypothetical action trajectories as new training data for the agent's model.

[0008] Thirdly, this application provides a model training apparatus, comprising: a first acquisition unit for acquiring real data, the real data including real instructions, real videos, and real motion trajectories, wherein the real videos and real motion trajectories are obtained by data collection during the process of an agent executing the real instructions in a real environment; a second acquisition unit for acquiring simulation data, the simulation data including hypothetical instructions, simulation videos, and simulation motion trajectories, wherein the hypothetical instructions are generated based on the initial frames of the real videos, and the simulation videos and simulation motion trajectories are obtained by simulating the agent executing the hypothetical instructions; and a training unit for training an initial embodied world model based on the real data and the simulation data to obtain a trained embodied world model, wherein the initial embodied world model includes an initial video generation model and an initial motion generation model, and the trained embodied world model includes a trained video generation model and a trained motion generation model; wherein the trained embodied world model is used to generate training data containing language instructions, videos, and motion trajectories for the model training of the agent.

[0009] Fourthly, this application provides a training data augmentation apparatus, comprising: an acquisition unit, configured to acquire real video frames from the original training data of an agent's model, wherein the real video frames are obtained by video capture during the execution of real instructions by the agent in a real environment; an instruction generation unit, configured to generate hypothetical instructions based on the real video frames, wherein the hypothetical instructions belong to task instructions that the agent supports executing in the task scenario described in the real video frames; a data generation unit, configured to input the real video frames and the hypothetical instructions into a trained embodied world model, and generate hypothetical videos and hypothetical action trajectories through the trained embodied world model, wherein the trained embodied world model is trained according to the model training method described in the first aspect; and a training data determination unit, configured to determine the hypothetical instructions, the hypothetical videos, and the hypothetical action trajectories as new training data for the agent's model.

[0010] Fifthly, this application provides an electronic device, including: one or more processors; and a memory storing one or more programs thereon, which, when executed by the one or more processors, cause the one or more processors to implement, as described in the first aspect, the model training method, or the training data augmentation method described in the second aspect.

[0011] Sixthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements, for example, the model training method described in the first aspect, or the training data augmentation method described in the second aspect.

[0012] In a seventh aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements, for example, the model training method described in the first aspect, or the training data augmentation method described in the second aspect.

[0013] The model training method, training data augmentation method, apparatus, and electronic device proposed in this application acquire real data, including real instructions, real videos, and real motion trajectories. The real videos and motion trajectories are obtained by data collection during the execution of real instructions by an agent in a real environment. Simulation data is also acquired, including hypothetical instructions, simulation videos, and simulation motion trajectories. The hypothetical instructions are generated based on the initial frames of the real videos, and the simulation videos and motion trajectories are obtained by simulating the agent's execution of the hypothetical instructions, reducing the gap between simulation and reality and improving the data quality of the simulation data. Based on the aforementioned real and simulation data, an initial embodied world model is trained to obtain a trained embodied world model, improving the training effect of the embodied world model. The initial embodied world model includes an initial video generation model and an initial motion generation model, while the trained embodied world model includes a trained video generation model and a trained motion generation model. The trained embodied world model is used to generate training data containing language instructions, videos, and motion trajectories for the agent's model training. Thus, by utilizing an embodied world model with high-quality training data generation capabilities, the problem of providing a large amount of high-quality training data for the agent is solved. Attached Figure Description

[0014] Figure 1 is a schematic flowchart of the model training method provided in the embodiments of this application.

[0015] Figure 2 is a schematic diagram of the process of training an initial embodied world model in the model training method provided in the embodiments of this application.

[0016] Figure 3 is a flowchart illustrating the first stage of the training process in the model training method provided in the embodiments of this application.

[0017] Figure 4 is an example of the framework for model training in the first stage.

[0018] Figure 5 is a flowchart illustrating the second stage of the training process in the model training method provided in the embodiments of this application.

[0019] Figure 6 is an example diagram of the framework for model training in the second stage.

[0020] Figure 7 is a flowchart illustrating the training process in the third stage of the model training method provided in the embodiments of this application.

[0021] Figure 8 shows an example of the framework for model training in the third stage.

[0022] Figure 9 is a flowchart illustrating the training data augmentation method provided in an embodiment of this application.

[0023] Figure 10 is a schematic diagram of the structure of the embodied world model provided in the embodiment of this application.

[0024] Figure 11 is a schematic diagram of the structure of the model training device provided in the embodiment of this application.

[0025] Figure 12 is a schematic diagram of the training data augmentation device provided in an embodiment of this application.

[0026] Figure 13 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0027] To enable those skilled in the art to better understand the technical solutions of this application, the technical solutions provided in this application will be described in detail below with reference to the accompanying drawings.

[0028] Exemplary embodiments will be described more fully below with reference to the accompanying drawings; however, the described exemplary embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will enable those skilled in the art to fully understand the scope of this application.

[0029] As used herein, the term "and / or" includes any and all combinations of one or more related enumerated purposes.

[0030] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of a feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded.

[0031] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0032] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this application, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined in the embodiments of this application.

[0033] The technological implementation and development of embodied intelligence has always faced many challenges, one of which is how to provide a large amount of high-quality data for the training of intelligent agents.

[0034] In some approaches, using real-world data as training data for intelligent agents suffers from high data acquisition costs and low scalability. Using simulation data from simulated environments as training data for intelligent agents results in a simulation-to-reality gap (sim2real gap), leading to poor performance of the trained agents in the real world. While there is a wealth of internet video data related to embodied intelligence, internet video data faces licensing issues and lacks motion trajectory data of intelligent agents, making it unsuitable for direct use in training.

[0035] In other approaches, training data for the agent is generated using vision-language-action models (VLA models), such as visual-language models (VLMs) and motion expert models. However, these models have limited generalization capabilities and cannot accurately understand and respond to natural language instructions of different types, styles, and complexities, resulting in poor quality of the generated training data. Alternatively, training data can be generated for the agent using video-generated world models. These world models are trained using massive amounts of internet video data and have good instruction following and generalization capabilities, but they cannot generate high-quality motion trajectories.

[0036] It is clear that none of the above methods can provide the agent with a large amount of high-quality training data.

[0037] This application provides a model training method, training data augmentation method, apparatus, and electronic device. It acquires real data, including real commands, real videos, and real motion trajectories; acquires simulation data, including hypothetical commands, simulation videos, and simulation motion trajectories, where the hypothetical commands are generated from the initial frames of the real videos; and combines the real and simulation data to train an embodied world model that includes a video generation model and a motion generation model. On one hand, simulation data compensates for the shortcomings of insufficient real data and limited scalability; on the other hand, it utilizes the initial frames of real videos to generate hypothetical commands in the simulation data, reducing the gap between simulation and reality. Thus, it provides a large amount of rich, high-quality training data for training the embodied world model, improving the training effect and model capability. The trained embodied world model is used to generate training data for an agent, including language commands, videos, and motion trajectories, improving the generation efficiency and quality of training data for the agent, and achieving the generation of a large amount of high-quality training data for the agent's model training.

[0038] The embodiments of this application can be executed by an electronic device, which may include a terminal and / or a server.

[0039] Please refer to Figure 1, which is a flowchart illustrating the model training method provided in an embodiment of this application. This model training method is applied to a terminal, and as shown in Figure 1, it includes at least the following steps S101 to S103.

[0040] S101, Acquire real data. Real data includes real instructions, real videos, and real action trajectories. Real videos and real action trajectories are obtained by collecting data during the process of the agent executing real instructions in a real environment.

[0041] Among them, real commands are control commands issued to intelligent agents in real environments (i.e., real physical environments, real physical worlds). Real commands are language commands, that is, they are carried out in natural language. They can be the result of voice command recognition and conversion, directly input text commands, or language commands converted from other interaction methods such as brain-computer interfaces and gesture recognition.

[0042] Among them, real video is a video that records the dynamic process of an intelligent agent responding to and executing real instructions in a real environment; real motion trajectory is the motion trajectory of an intelligent agent executing real instructions in a real environment, such as the motion path of the agent's body and / or end effector in the real environment.

[0043] In a real-world environment, the agent can be remotely operated. During this operation, real commands are issued and data is collected, resulting in real video and motion trajectories. This yields real data containing real commands, video, and motion trajectories. Multiple sets of real data can be generated. Within a single set, the commands, video, and motion trajectories are interconnected, indicating that the video and motion trajectories within that set were collected during the agent's execution of the commands specified in that set.

[0044] In this step, real data can be obtained from a data acquisition device in a real environment, such as a functional module with data acquisition capabilities deployed by the agent itself, or other devices deployed in the same real environment as the agent; or, real data can be obtained from a local or remote storage database; or, real data can be received from user input.

[0045] Step S102: Obtain simulation data. The simulation data includes hypothetical instructions, simulation videos, and simulation action trajectories. The hypothetical instructions are generated based on the initial frames of the real video. The simulation videos and simulation action trajectories are obtained by simulating the execution of hypothetical instructions by the agent.

[0046] The initial frame of the real video can be either the first N frames of the real video, or the first N frames of the real video showing the environment, where N is greater than or equal to 1. The initial frame of the real video describes the initial environmental state of the real environment, in which there are objects that support agent operations. Hypothetical instructions are the language instructions that the agent can execute in this initial environmental state; they can be obtained by analyzing the language instructions that the agent can execute in this initial environmental state.

[0047] An initial frame can be analyzed to obtain one or more hypothetical instructions. One hypothetical instruction can correspond to one or more simulation data. Therefore, one piece of real data can be used to generate one or more simulation data.

[0048] In this way, by establishing the relationship between hypothetical instructions and the initial frames of real videos, a scene association between real data and simulation data is established, enabling the simulation of more language instructions executed by the agent in the initial environmental state. This reduces data acquisition costs, effectively increases the amount of training data used for training the embodied world model, reduces the difference between simulation and reality in the simulation data, improves the data quality of the simulation data, and provides a large amount of high-quality training data for the embodied world model.

[0049] In some embodiments, hypothetical instructions may not include real instructions. Based on the initial frames of real video data, multiple language instructions that the agent can execute in the initial state described by the initial frames can be analyzed. The remaining language instructions, excluding the real instructions in the real data, are identified as hypothetical instructions. In this way, hypothetical instructions compensate for language instructions that cannot be executed in the real environment, providing more training data for the training of the embodied world model.

[0050] For example, in real data, the initial frame of a real video describes the initial environmental state of "three items A, B, and C are placed on the table". The real instruction is to pick up A, so the hypothetical instructions could be to pick up B and pick up C.

[0051] Among them, simulation video is a video recording the dynamic process of an agent responding to and executing hypothetical instructions in a simulation environment; simulation action trajectory is the action trajectory of an agent executing hypothetical instructions in a simulation environment.

[0052] After receiving hypothetical instructions, the process of executing these instructions by the agent is simulated and data is collected in a simulation environment, resulting in simulation videos and motion trajectories. This yields real data containing hypothetical instructions, simulation videos, and motion trajectories. Multiple simulation data sets can be generated. Within the same set, the hypothetical instructions, simulation videos, and motion trajectories are interconnected, indicating that the simulation videos and motion trajectories within the same set were collected during the simulation of the agent executing the hypothetical instructions within that set.

[0053] In this step, simulation data can be obtained from a simulation device, which can be the same device as the device currently used for model training or a different device; alternatively, simulation data can be obtained from a local or remote storage database; or, simulation data input by the user can be received.

[0054] S103, Based on real data and simulation data, train the initial embodied world model to obtain the trained embodied world model. The initial embodied world model includes the initial video generation model and the initial action generation model. The trained embodied world model includes the trained video generation model and the trained action generation model. The trained embodied world model is used to generate training data containing language instructions, videos and action trajectories for the model training of the intelligent agent.

[0055] The video generation model generates videos recording the dynamic process of the agent executing language commands, while the action generation model generates the action trajectories of the agent executing the language commands. After inputting language commands into the embodied world model, the video generation model generates videos of the agent executing the voice commands, and the action generation model generates the action trajectories of the agent executing the language commands. Thus, the embodied world model can provide training data for the agent model, including not only language commands and videos but also action estimations, improving the completeness of the training data. Moreover, by training with both real and simulated data, the performance of the embodied world model is effectively improved. Using the embodied world model to generate training data for the agent achieves a balance between data quality and generation efficiency.

[0056] In this step, real instructions from real data are used as input data during the training process, and real videos and real motion trajectories from real data are used as the corresponding label data. Hypothetical instructions from simulation data are used as input data, and simulated videos and simulated motion trajectories from simulation data are used as the corresponding label data. The initial embodied world model is trained to obtain the trained embodied world model. That is, the initial video generation model and the initial motion generation model are trained to obtain the trained video generation model and the trained motion generation model.

[0057] In this embodiment, real data and simulated data are used as training data for the embodied world model. On the one hand, real data provides high-quality training data for the embodied world model; on the other hand, hypothetical instructions are generated by referring to the initial frames of real videos in the real data, and then further simulated data is generated. This improves the data quality of the simulated data and compensates for the high cost and insufficient scalability of real data collection. This provides a large amount of rich, high-quality training data for the embodied world model, improving the model training effect. Through multiple iterative training of the embodied world model, its robustness and generalization are improved. The embodied world model includes a video generation model and an action generation model. It can not only generate videos of the agent executing language instructions, but also generate the action trajectories of the agent executing language instructions. Thus, by applying the trained embodied world model to the generation of training data for the agent, the generation efficiency and quality of training data are improved, ensuring the integrity of the training data and providing a large amount of high-quality training data for the model training of the agent.

[0058] In some embodiments, multiple initial environmental states are prepared in advance. The process of generating real data may include: for each initial environmental state, when the real environment is in that initial environmental state, controlling the agent to execute real instructions and collecting data on the dynamic process of the agent executing real instructions to obtain the real data corresponding to that initial environmental state; wherein, the multiple initial environmental states are generated in a simulation environment and replicated in a real environment. For example, multiple initial environmental states are randomly generated using a simulation platform, and then replicated in a real environment. For each initial environmental state, real data corresponding to the initial environmental state is collected through teleoperation of the agent. Thus, on the one hand, a rich variety of initial environmental states are generated through the simulation environment, improving the diversity of real data; on the other hand, initial environmental states that can be realized in both simulation and real environments are provided, so that hypothetical instructions that the agent can support execution can be determined by referring to the initial frames of real video.

[0059] In some embodiments, the process of generating hypothetical instructions may include: extracting and analyzing image features from the initial frame of a real video using an instruction generation model (e.g., a Virtual Model) to generate hypothetical instructions. The instruction generation model possesses the capability to generate corresponding language instructions from scene images, and can be trained in advance to ensure it possesses this capability. Therefore, utilizing the instruction generation model improves the efficiency and accuracy of hypothetical instruction generation, ensuring that language instructions executable by the agent are obtained within the initial environmental state described by the initial frame.

[0060] Please refer to Figure 2, which is a flowchart illustrating the training process of the initial embodied world model in the model training method provided in this application embodiment. As shown in Figure 2, the training of the initial embodied world model is a phased training process, which includes at least the following steps S201 to S203: S201, in the first phase, the initial video generation model is trained multiple times based on real data and simulation data to obtain the video generation model after the first phase training.

[0061] In the initial embodied world model, the initial action generation model follows the initial video generation model. The ability of the video generation model to generate high-quality video has a decisive impact on the ability of the action generation model to generate high-quality motion trajectories. Therefore, the initial video generation model is trained first in the first stage.

[0062] In this step, both real and simulated data are used to train the initial video generation model. The real videos from the real data set and the simulated videos from the simulated data set serve as sample labels during the training process. These labels are compared with the videos generated by the model, allowing for supervised multi-round training of the initial video generation model to obtain the first-stage trained model. Thus, the real and simulated data provide a wealth of high-quality training data for the initial video generation model in the first stage. Utilizing this high-quality training data for supervised multi-round training effectively improves the training performance of the initial video generation model in the first stage.

[0063] In some embodiments, in the first stage, i is greater than or equal to 1, as shown in Figure 3. The i-th round of training of the initial video generation model includes steps S2011 to S2024: S2011, in real data and simulation data, determine the first input sample and the first sample label corresponding to the first input sample. The first input sample includes the initial frame of the real command, the real motion trajectory and the real video. The first sample label includes the real video. Alternatively, the first input sample includes the initial frame of the hypothetical command, the simulated motion trajectory and the real video. The first sample label includes the initial frame of the simulated video and the real video. S2012, input the first input sample into the video generation model for the i-th round of training, and perform video generation in the video generation model for the i-th round of training to obtain the first output video. S2013, determine the first video generation loss value according to the first output video and the first sample label. S2014, adjust the parameters of the video generation model for the i-th round of training according to the first video generation loss value to obtain the video generation model after the i-th round of training.

[0064] In this model, one real data point or one simulated data point can correspond to a pair of sample data in the first stage. A pair of sample data includes a first input sample and a first sample label corresponding to that first input sample. In this way, based on multiple real data points and multiple simulated data points, multiple pairs of sample data can be obtained, providing rich input samples and sample labels for multiple rounds of training of the initial video generation model in the first stage, thereby improving the training effect of the video generation model.

[0065] In the i-th round of training, the first input sample and the first sample label come from real data or simulated data.

[0066] In S2011, when the first input sample and the first sample label come from real data: the first input sample includes real instructions, real action trajectories, and the initial frame of a real video; wherein, the real instructions and the initial frame of the real video provide environmental information for the agent to execute language instructions, so that the video generation model trained in the i-th round can accurately generate video based on the process of the agent executing real instructions in the initial environmental state described by the initial frame of the real video; wherein, the real action trajectory provides action information for the agent to execute real instructions, which is beneficial to improving the video quality generated by the video generation model; the first sample label includes real video, so that the video generation model trained in the i-th round can generate videos that are closer to real videos.

[0067] In S2011, when the first input sample and the first sample label come from simulation data: the first input sample corresponds to the hypothetical instruction, the simulated action trajectory, and the initial frame of the real video; wherein, the hypothetical instruction and the initial frame of the real video provide environmental information for the agent to execute the language instruction, so that the video generation model trained in the i-th round can accurately generate video based on the process of the agent executing the hypothetical instruction in the initial environmental state described by the initial frame of the real video; wherein, the simulated action trajectory is the video generation process of the video generation model, providing action information of the agent executing the hypothetical instruction, which is beneficial to improving the quality of the video generated by the video generation model; the first sample label includes the initial frame of the simulated video and the real video, so that the video generated by the video generation model after the i-th round of training is close to the simulated video in terms of video content and close to the initial frame of the real video in terms of video style, and will not lack realism due to the use of simulated video for training the video generation model.

[0068] In S2012, the first input sample is input into the video generation model trained in the i-th round, and video generation is performed in the video generation model trained in the i-th round to obtain the first output video. That is, when the first input sample comes from real data, real instructions, real motion trajectories, and the initial frame of real video can be input into the video generation model trained in the i-th round; when the first input sample comes from simulation data, hypothetical instructions, simulated motion trajectories, and the initial frame of real video can be input into the video generation model trained in the i-th round.

[0069] In one possible implementation, before the video generation model undergoing the i-th round of training is generated by inputting the first input sample into the video generation model and before the first output video is obtained, i.e. before S2012, the i-th round of training further includes: generating a first random number; if the first random number is less than a first threshold, removing the real motion trajectory from the first input sample, or removing the simulated motion trajectory from the first input sample.

[0070] In this implementation, if the first random number is less than a first threshold, and the first input sample contains real motion trajectories (i.e., the first input sample comes from real data), then the real motion trajectories are removed from the first input sample; if the first input sample contains simulated motion trajectories (i.e., the first input sample comes from simulated data), then the simulated motion trajectories are removed from the first input sample. The first threshold is a set value, and the first random number is a random value within a set range. For example, the first threshold is 0.3, and the first random number is a random value between 0 and 1. Thus, motion trajectories are removed from the input samples of the video generation model with a certain probability, enabling the video generation model to initially possess the ability to generate videos without relying on motion trajectories, preparing it for generating high-quality videos completely independent of motion trajectories.

[0071] In one possible implementation, the embodied world model further includes an encoder. Before the first input sample is input into the video generation model undergoing the i-th round of training, and before video generation is performed in the i-th round of training to obtain the first output video (i.e., before S2012), the i-th round of training further includes: inputting the first input sample into the encoder to obtain a first encoded vector. Then, the first encoded vector is input into the video generation model undergoing the i-th round of training, and video generation is performed in the i-th round of training. Thus, by encoding the input sample through the encoder, an encoded vector that can be directly used for feature learning is provided to the video generation model.

[0072] The encoder may include a video encoder and a text encoder. The video encoder can be used to encode video frames (such as the initial frame of a real video) in the first input sample. The text encoder can be used to encode language instructions (such as real instructions or hypothetical instructions) and motion trajectories (motion trajectories can be presented in text form, such as real motion trajectories or simulated motion trajectories) in the first input sample.

[0073] In S2013, a first video generation loss value is determined based on the first output video and the first sample label. When the first input sample comes from real data, the first sample label includes the real video, and the first video generation loss value is determined based on the difference between the first output video and the real video. When the first input sample comes from simulated data, the first sample label includes the initial frame of both the simulated video and the real video, and the first video generation loss value is determined based on the difference between the first output video and the simulated video, as well as the difference between the first output video and the initial frame. This ensures that the video generated by the video generation model after the i-th training round is close to the simulated video in content and close to the real video in style at the initial frame, preventing the video generated by the video generation model from lacking realism due to training with simulated videos.

[0074] In one possible implementation, S2013 includes: when the first input sample comes from real data, obtaining a first video generation loss value by calculating the reconstruction loss between the first output video and the real video; when the first input sample comes from simulation data, obtaining a first video generation loss value by calculating the reconstruction loss between the first output video and the simulation video, the semantic consistency loss between the first output video and the simulation video, and the perceptual loss between the video frames of the first output video and the initial frames of the real video.

[0075] Among them, the reconstruction loss compares the differences between the two (such as the first output video and the real video, or the first output video and the simulated video in this implementation scheme) at the pixel level, the semantic consistency loss compares the differences between the two (such as the first output video and the real video, or the first output video and the simulated video in this implementation scheme) in image semantics and text semantics, and the perceptual loss is a loss designed for "visual perception", which compares the differences between the two in visual perception, making up for the shortcomings of the reconstruction loss and the semantic consistency loss, which do not pay attention to visual experience.

[0076] The video frames of the first output video are time-series images in the first output video. On the timeline of the first output video, each timestamp can correspond to a video frame.

[0077] In this implementation, when the first input sample and the first sample label come from real data, the first sample label includes real video. By comparing the differences between the first output video and the real video at the pixel level, the reconstruction loss between the first output video and the real video is calculated to obtain the first video generation loss value. The first video generation loss value is the reconstruction loss value between the first output video and the real video. Using the reconstruction loss between the first output video and the real video as the model training loss can help the video generated after the video generation model is trained to be closer to the real video.

[0078] In this implementation, when the first input sample and the first sample label come from simulation data, the first sample label includes the initial frames of the simulation video and the real video. The reconstruction loss between the first output video and the simulation video is calculated by comparing the pixel-level differences. The semantic consistency loss between the first output video and the simulation video is obtained by comparing their semantic differences. The visual perception loss between the first output video and the initial frames of the real video is obtained by comparing the visual perception differences between each video frame in the first output video and the initial frames of the real video. The first video generation loss is obtained by weighting the reconstruction loss, semantic consistency loss, and visual perception loss between the first output video and the initial frames of the real video. Therefore, by combining the reconstruction loss, semantic consistency loss, and visual perception loss between the first output video and the initial frames of the real video, the video generation model, after training, generates videos that are close to the simulation video in terms of image pixels and semantics, and close to the real video in terms of visual perception, ensuring the realism of the videos generated by the video generation model.

[0079] For example, a semantic-level comparison can be performed on each video frame in the first output video and each video frame in the simulation video using a pre-trained cross-modal model to obtain the semantic consistency loss between the first output video and the simulation video. The cross-modal model can be, for example, a contrastive language-image pre-training (CLIP) model.

[0080] For example, a pre-trained visual feature extraction model can be used to perform high-level visual feature extraction on each video frame in the first output video to obtain the visual features of each video frame in the first output video. High-level visual feature extraction can also be performed on the initial frame of the real video to obtain the visual features of the initial frame. The visual features of each video frame in the first output video can be compared with the visual features of the initial frame to obtain the visual perception loss between the first output video and the initial frame of the real video.

[0081] In S2014, the first video generation loss value is used as the model training loss. An optimization algorithm is employed to adjust the parameters of the video generation model trained in the i-th round, resulting in the video generation model after the i-th round of training. The optimization algorithm is not restricted here. After S2014, it can be determined whether the i-th round of training is the last round of training in the first stage. For example, it can be determined whether i equals the training count threshold corresponding to the first stage. If yes, then the i-th round of training is determined to be the last round of training in the first stage; otherwise, it is determined not to be the last round of training in the first stage. Alternatively, it can be determined whether the first video generation loss value is less than the loss threshold corresponding to the first stage. If yes, then the i-th round of training is determined to be the last round of training in the first stage; otherwise, it is determined not to be the last round of training in the first stage. If the i-th round of training is the last round of training in the first stage, the video generation model after the first stage of training can be obtained. If the i-th round of training is not the last round of training in the first stage, i can be incremented by one, and the process jumps to S2011 to execute the next round of training in the first stage. Thus, through multiple rounds of iterative training, the video generation quality of the video generation model can be improved.

[0082] As an example, Figure 4 shows a framework example for model training in the first stage. As shown in Figure 4, firstly, the initial frame of the real video can be input into the instruction generation model (Figure 4 uses VLM as an example) to generate hypothetical instructions; in the simulation environment, simulation video and simulation motion trajectory can be generated based on the hypothetical instructions, wherein the hypothetical instructions, simulation video, and simulation motion trajectory form simulation data; in the first stage, the initial frame containing the hypothetical instructions, simulation motion trajectory (which can be removed according to the aforementioned embodiment) and real video can be used as the first input sample and input into the video generation model to obtain the first output video of the video generation model. Based on the first output video, simulation video, and the initial frame of the real video, the first video generation loss value is calculated, and the parameters of the video generation model are adjusted according to the first video generation loss value.

[0083] S202, In the second stage, the initial motion generation model is trained multiple times based on real data, simulation data, and the video generation model trained in the first stage to obtain the motion generation model trained in the second stage. The model parameters of the video generation model trained in the first stage remain unchanged.

[0084] After completing the first stage of training, the video generation model trained in the first stage is used in the second stage to assist in the multi-round training of the initial action generation model.

[0085] In this process, after the video generation model is trained in the first stage, an initial action generation model is connected to form the embodied world model trained in the first stage. Connecting the initial action generation model after the video generation model trained in the first stage means that the output data of at least one network layer (which may include hidden layers and / or output layers) in the video generation model is used as the input data of at least one network layer (which may include input layers and / or hidden layers) in the action generation model.

[0086] In the second stage, using both real and simulated data, the embodied world model trained in the first stage is trained in multiple rounds. In each round of training, the model parameters of the video generation model trained in the first stage remain unchanged, while the parameters of the initial motion generation model are adjusted. This process of training the initial motion generation model in multiple rounds yields the motion generation model trained in the second stage, which is the embodied world model trained in the second stage. The embodied world model trained in the second stage includes the video generation model trained in the first stage and the motion generation model trained in the second stage.

[0087] Thus, through the first and second stages, the initial video generation model and the initial action generation model are trained separately. The first stage focuses on improving the video generation capability of the video generation model, while the second stage focuses on improving the action trajectory generation capability of the action generation model. Moreover, the video generation model trained in the first stage can assist the training of the action generation model in the second stage, effectively improving the training effect of the video generation model and the action generation model, as well as the consistency between the generated video and the action.

[0088] In this step, multiple input samples and corresponding sample labels for the second stage are obtained based on real and simulated data. For ease of distinction, these input samples are referred to as second input samples, and their labels as second sample labels. The second input samples may include the input data required by the video generation model, and the second sample labels may include real or simulated motion trajectories. By inputting the second input samples into the video generation model trained in the first stage, and through the layer-by-layer processing of the video generation model trained in the first stage and the layer-by-layer processing of the motion generation model to be trained in the second stage, the motion trajectory generated by the motion generation model is obtained. The training loss value of the motion generation model is determined by comparing the real or simulated motion trajectory in the second sample labels with the motion trajectory generated by the motion generation model. Based on the training loss value, the parameters of the motion generation model are adjusted. In this way, the motion generation model is trained in supervised multiple rounds to obtain the video generation model trained in the second stage. Thus, by using real and simulated data, a large amount of high-quality training data was provided for the training of the initial motion generation model in the second stage. Based on this high-quality training data, with the assistance of the video generation model trained in the first stage, the initial motion generation model was trained in supervised multiple rounds, which effectively improved the training effect of the initial motion generation model in the second stage.

[0089] In some embodiments, in the second stage, j is greater than or equal to 1, as shown in Figure 5. The j-th round of training of the initial action generation model includes steps S2021~S2024: S2021, in real data and simulation data, determine the second input sample and the second sample label corresponding to the second input sample. The second input sample includes real instructions, real action trajectories and initial frames of real videos, and the second sample label includes real action trajectories; or, the second input sample includes hypothetical instructions, simulated action trajectories and initial frames of real videos, and the second sample label includes simulated action trajectories; S2022, input the second input sample into the video generation model trained in the first stage, after the first stage of training... In the video generation model, video generation is performed to obtain the hidden layer output features of the video generation model; in S2023, a randomly generated noise vector is input into the input layer of the action generation model trained in the j-th round, and the hidden layer output features are used as the generation condition input to the hidden layer of the action generation model trained in the j-th round. Action trajectory is generated in the action generation model trained in the j-th round to obtain the first output action trajectory; in S2024, the first trajectory generation loss value is determined based on the first output action trajectory and the second sample label; in S2025, the parameters of the action generation model trained in the j-th round are adjusted based on the first trajectory generation loss value to obtain the action generation model after the j-th round of training.

[0090] In this model, one real data point or one simulated data point can correspond to a pair of sample data in the second stage. A pair of sample data includes a second input sample and a second sample label corresponding to that second input sample. In this way, based on multiple real data points and multiple simulated data points, multiple pairs of sample data can be obtained, providing rich input samples and sample labels for multiple rounds of training of the initial action generation model in the second stage, thereby improving the training effect of the video generation model.

[0091] In the j-th round of training, the second input sample and the second sample label come from real data or simulation data.

[0092] In S2021, when the second input sample and the second sample label come from real data: the second input sample includes real instructions, real action trajectories, and the initial frame of a real video; wherein, the real instructions and the initial frame of the real video provide environmental information for the agent to execute language instructions, so that the video generation model trained in the first stage can accurately generate video for the process of the agent executing real instructions in the initial environmental state described by the initial frame of the real video, so that the action generation model trained in the j-th round can accurately generate action trajectories for this process; wherein, the real action trajectory is the video generation process of the video generation model and the action trajectory generation process of the action generation model, providing action information for the agent to execute real instructions, so as to improve the video quality generated by the video generation model and the action trajectory quality generated by the action generation model; the second sample label includes real action trajectories, so that the action generation model trained in the j-th round generates action trajectories that are closer to real action trajectories.

[0093] In S2021, when the second input sample comes from simulation data: the second input sample corresponds to the hypothetical instruction, simulated action trajectory, and the initial frame of the real video; wherein, the hypothetical instruction and the initial frame of the real video provide environmental information for the agent to execute language instructions, so that the video generation model trained in the first stage can accurately generate video based on the process of the agent executing the hypothetical instruction in the initial environmental state described by the initial frame of the real video, thereby enabling the action generation model trained in the j-th round to accurately generate action trajectories based on the process of the agent executing real instructions in the initial environmental state described by the initial frame of the real video; wherein, the simulated action trajectory is the video generation process of the video generation model and the action trajectory generation process of the action generation model, providing action information of the agent executing the hypothetical instruction, which is beneficial to improving the video quality generated by the video generation model; the second sample label includes the simulated action trajectory, so that the action generation model trained in the j-th round generates action trajectories that are closer to the simulated action trajectories.

[0094] In S2022, the second input sample is fed into the video generation model trained in the first stage, and video generation is performed in the video generation model to obtain the hidden layer output features of the video generation model. That is, when the second input sample comes from real data, real instructions, real motion trajectories, and the initial frame of real video can be input into the video generation model trained in the first stage; when the second input sample comes from simulation data, hypothetical instructions, simulated motion trajectories, and the initial frame of real video can be input into the video generation model trained in the first stage.

[0095] In one possible implementation, before inputting the second input sample into the video generation model trained in the first stage, generating video in the video generation model trained in the first stage, and obtaining the hidden layer output features of the video generation model, i.e. before S2022, the j-th round of training also includes: generating a second random number; if the second random number is less than a second threshold, removing the real motion trajectory from the second input sample, or removing the simulated motion trajectory from the second input sample.

[0096] In this implementation, if the second random number is less than the second threshold, and the second input sample contains real motion trajectories (i.e., the second input sample comes from real data), then the real motion trajectories are removed from the second input sample; if the second input sample contains simulated motion trajectories (i.e., the second input sample comes from simulated data), then the simulated motion trajectories are removed from the second input sample. The second threshold is a set value, and the second random number is a random value within a set range. For example, the second threshold is 0.3, and the second random number is a random value between 0 and 1. Thus, motion trajectories are removed from the input samples of the video generation model with a certain probability, enabling the motion generation model to initially possess the ability to generate motion trajectories based on the hidden layer output features of the video generation model without relying on motion trajectories. This prepares the embodied world model for generating videos and motion trajectories of intelligent agents based on language commands and initial frames in the absence of motion trajectories.

[0097] In one possible implementation, the embodied world model also includes an encoder. The encoder may include a video encoder and a text encoder. Specific details can be found in the description of the foregoing embodiments, and will not be repeated here.

[0098] In S2023, a randomly generated noise vector is input into the input layer of the action generation model trained in the j-th round. The hidden layer output features of the video generation model trained in the first stage are used as generation conditions and input into the hidden layer of the action generation model trained in the j-th round. In the action generation model trained in the j-th round, the action trajectory is generated based on the randomly generated noise vector and the hidden layer output features to obtain the first output action trajectory.

[0099] In one possible implementation, the video generation model trained in the first stage can include multiple hidden layers, and the action generation model trained in the j-th round can also include multiple hidden layers. A correspondence between the hidden layers in the video generation model trained in the first stage and the hidden layers in the action generation model trained in the j-th round can be established beforehand. Based on this correspondence, the hidden layer output features of the hidden layers in the video generation model trained in the first stage can be input into the corresponding hidden layers in the action generation model trained in the j-th round. In the action generation model trained in the j-th round, the hidden layer output features of the hidden layers in the video generation model trained in the first stage can be fused sequentially with the noise vector in the input layer of the action generation model trained in the j-th round. Based on the feature vector obtained through feature fusion of multiple hidden layers, a trajectory is generated to obtain the first output action trajectory. Therefore, based on the hidden layer correspondence between the video generation model and the action generation model, and based on the feature extraction capability of the video generation model, the video generation model provides hidden layer features of different granularities to the action generation model, providing rich feature information for the action trajectory generation and improving the quality of the generated action trajectory.

[0100] In S2024, a first trajectory generation loss value is determined based on the first output motion trajectory and the second sample label. When the second input sample and the second sample label come from real data, the second sample label includes the real motion trajectory, and the first trajectory generation loss value is determined based on the difference between the first output motion trajectory and the real motion trajectory. When the second input sample and the second sample label come from simulated data, the second sample label includes the simulated motion trajectory, and the first trajectory generation loss value is determined based on the difference between the first output motion trajectory and the simulated motion trajectory. Thus, in the second stage, with multiple rounds of training on the initial motion generation model, the output motion trajectory of the motion generation model gradually approaches the real motion trajectory or the simulated motion trajectory, effectively improving the trajectory generation accuracy of the motion generation model.

[0101] In one possible implementation, S2024 includes: when the second input sample comes from real data, obtaining a first trajectory generation loss value by calculating the reconstruction loss between the first output motion trajectory and the real motion trajectory; when the second input sample comes from simulated data, obtaining a first trajectory generation loss value by calculating the reconstruction loss between the first output motion trajectory and the simulated motion trajectory. Thus, by using the reconstruction loss between the motion trajectory in the second sample label and the first output motion trajectory as the model training loss, the motion generation model acquires the ability to reconstruct real or simulated motion trajectories as the number of training iterations increases, effectively improving the trajectory generation accuracy of the motion generation model.

[0102] The reconstruction loss between the first output motion trajectory and the real motion trajectory can be: a reconstruction loss based on point-by-point error, such as the mean squared error between the spatial coordinates of the first output motion trajectory at multiple time points and the spatial coordinates of the real motion trajectory at those multiple time points; or a loss based on the sequence as a whole, such as extracting the shape features of the first input motion trajectory, extracting the shape features of the real motion trajectory, comparing the shape features of the first input motion trajectory with the shape features of the real motion trajectory, and obtaining the reconstruction loss between the first output motion trajectory and the real motion trajectory.

[0103] The reconstruction loss between the first output motion trajectory and the simulated motion trajectory can be referred to in the above examples of the reconstruction loss between the first output motion trajectory and the real motion trajectory, and will not be repeated here.

[0104] In S2025, based on the loss value generated from the first trajectory, an optimization algorithm is used to adjust the parameters of the action generation model trained in the j-th round, resulting in the action generation model after the j-th round of training. The optimization algorithm is not restricted here. After S2025, it can be determined whether the j-th round of training is the last round of training in the first stage. For example, it can be determined whether j equals the training count threshold corresponding to the second stage. If yes, then the j-th round of training is determined to be the last round of training in the second stage; otherwise, it is determined not to be the last round of training in the second stage. Alternatively, it can be determined whether the loss value generated from the first trajectory is less than the loss threshold corresponding to the second stage. If yes, then the j-th round of training is determined to be the last round of training in the second stage; otherwise, it is determined not to be the last round of training in the second stage. If the j-th round of training is the last round of training in the second stage, the action generation model after the second stage of training can be obtained. If the j-th round of training is not the last round of training in the second stage, j can be incremented by one, and the process jumps to S2021 to execute the next round of training in the second stage. Thus, through multiple rounds of iterative training, the accuracy of motion trajectories generated by the motion generation model can be improved.

[0105] As an example, Figure 6 shows a framework example for model training in the second stage. As shown in Figure 6, firstly, the initial frame of the real video can be input into the instruction generation model (Figure 6 uses VLM as an example) to generate hypothetical instructions; then, in the simulation environment, a simulation video and a simulation motion trajectory can be generated based on the hypothetical instructions, wherein the hypothetical instructions, the simulation video, and the simulation motion trajectory form simulation data; next, taking the second input sample and the second sample label as coming from the simulation data as an example, in the second stage, the initial frame containing the hypothetical instructions, the simulation motion trajectory (which can be removed according to the aforementioned embodiment), and the real video is used as the second input sample and input into the video generation model to obtain the hidden layer output features of the video generation model. The hidden layer output features are then input into the motion generation model to obtain the first output motion trajectory of the motion generation model; by comparing the first output motion trajectory and the simulation motion trajectory, the first trajectory generation loss value is obtained, and the parameters of the motion generation model are adjusted according to the first trajectory generation loss value.

[0106] S203, In the third stage, based on real data, simulation data, and the action generation model trained in the second stage, the video generation model trained in the first stage is trained multiple times to obtain the video generation model trained in the third stage. The model parameters of the action generation model trained in the second stage remain unchanged. The trained embodied world model includes the video generation model trained in the third stage and the action generation model trained in the second stage.

[0107] In this training model, after completing the second phase, the video generation model trained in the first phase and the action generation model trained in the second phase are used in the third phase. In the third phase, the action generation model trained in the second phase assists in the multi-round training of the video generation model trained in the first phase. In the first phase, the video generation model training lacks feedback from the action generation model's generated motion trajectories. Considering that the video generation model influences the action trajectory generation of the action generation model, the action generation model trained in the second phase is introduced in the third phase to assist in the training of the video generation model. The motion trajectories generated by the action generation model trained in the second phase serve as some feedback data, further enhancing the video generation capabilities of the video generation model. This allows the video generation model and the action generation model to be combined, enabling not only the generation of high-quality videos of the agent executing language commands but also the accurate generation of the agent's motion trajectories during the execution of language commands.

[0108] In this step, based on real and simulated data, multiple input samples and corresponding sample labels for the third stage are obtained. For ease of distinction, these input samples are referred to as third input samples, and their labels as third sample labels. The third input samples may include the input data required by the video generation model, and the third sample labels may include real video and real motion trajectories, or simulated video and simulated motion trajectories. By inputting the third input samples into the video generation model to be trained in the third stage, and through the layer-by-layer processing of the video generation model and the layer-by-layer processing of the motion generation model trained in the second stage, the video generated by the video generation model and the motion trajectories generated by the motion generation model are obtained. Based on the third sample labels, the video generated by the video generation model, and the motion trajectories generated by the motion generation model, the training loss value of the video generation model can be determined, and the parameters of the video generation model can be adjusted based on this training loss value. By executing the above process multiple times, supervised multi-round training of the video generation model is performed, resulting in the video generation model trained in the third stage, and thus the embodied world model is completed. Thus, by using real and simulated data, a large amount of high-quality training data is provided for the training of the video generation model in the third stage. Based on this high-quality training data, with the assistance of the action generation model trained in the second stage, supervised multi-round training is carried out on the video generation model trained in the first stage, which effectively improves the training effect of the video generation model in the third stage.

[0109] In some embodiments, in the third stage, k is greater than or equal to 1, as shown in Figure 7. The k-th round of training of the video generation model after the first stage training includes steps S2031~S2035: S2031, in real data and simulation data, determine the third input sample and the third sample label corresponding to the third input sample. The third input sample includes the initial frame of the real instruction, the real motion trajectory, and the real video. The third sample label includes the real motion trajectory and the real video. Alternatively, the third input sample includes the hypothetical instruction, the simulated motion trajectory, and the initial frame of the real video. The third sample label includes the simulated motion trajectory, the simulated video, and the initial frame of the real video. S2032, input the third input sample into the video generation model for the k-th round of training. In step S2033, video generation is performed in the model to obtain the hidden layer output features and the second output video of the video generation model; in step S2034, a randomly generated noise vector is input into the input layer of the action generation model after the second stage of training, and the hidden layer output features are used as generation conditions to be input into the hidden layer of the action generation model after the second stage of training. Action trajectory is generated in the action generation model after the second stage of training to obtain the second output action trajectory; in step S2035, the second video generation loss value and the second trajectory generation loss value are determined based on the second output video, the second output action trajectory and the third sample label; in step S2036, the parameters of the video generation model trained in the kth round are adjusted based on the second video generation loss value and the second trajectory generation loss value to obtain the video generation model after the kth round of training.

[0110] In this model, one real data point or one simulated data point can correspond to a pair of sample data in the third stage. A pair of sample data includes a third input sample and its corresponding third sample label. In this way, based on multiple real data points and multiple simulated data points, multiple pairs of sample data can be obtained, providing rich input samples and sample labels for multiple rounds of training of the video generation model in the third stage, thereby improving the training effect of the video generation model.

[0111] In the k-th round of training, the third input sample and the third sample label are derived from real data or simulation data.

[0112] In S2031, when the third input sample and the third sample label come from real data: the third input sample includes real instructions, real action trajectories, and the initial frame of a real video; wherein, the real instructions and the initial frame of the real video provide environmental information for the agent to execute language instructions, so that the video generation model trained in the kth round can accurately generate video for the process of the agent executing real instructions in the initial environmental state described by the initial frame of the real video, and enable the action generation model trained in the second stage to accurately generate action trajectories for this process; wherein, the real action trajectory provides action information for the agent to execute real instructions, which is beneficial to improving the video quality generated by the video generation model; the third sample label includes real video and real action trajectory, which can be used to determine the model loss of the video generation model trained in the kth round by referring to the real video, and the model loss of the action generation model trained in the second stage by referring to the real action trajectory, so as to combine the model loss of the video generation model trained in the kth round and the model loss of the action generation model trained in the second stage to adjust the parameters of the video generation model trained in the kth round.

[0113] In S2031, when the third input sample and the third sample label come from simulation data: the third input sample corresponds to the hypothetical instruction, the simulated action trajectory, and the initial frame of the real video; wherein, the hypothetical instruction and the initial frame of the real video provide environmental information for the agent to execute the language instruction, so that the video generation model trained in the kth round can accurately generate video for the process of the agent executing the real instruction in the initial environmental state described by the initial frame of the real video, so that the action generation model trained in the second stage can accurately generate the action trajectory for this process; wherein, the simulated action trajectory provides action information of the agent executing the hypothetical instruction, which is beneficial to improving the video quality generated by the video generation model trained in the kth round; the third sample label includes the simulation video, the initial frame of the real video, and the simulated action trajectory. On the one hand, the model loss of the video generation model trained in the kth round can be determined by referring to the initial frame of the simulation video and the real video, and on the other hand, the model loss of the action generation model trained in the second stage can be determined by referring to the simulated action trajectory. In order to combine the model loss of the video generation model trained in the kth round and the model loss of the action generation model trained in the second stage, the parameters of the video generation model trained in the kth round can be adjusted to improve the model training effect.

[0114] In one possible implementation, before inputting the third input sample into the video generation model for the kth round of training, and before generating the hidden layer output features and the second output video of the video generation model for the kth round of training (i.e., before S2032), the kth round of training further includes: generating a third random number; if the third random number is less than a third threshold, removing the real motion trajectory from the third input sample, or removing the simulated motion trajectory from the third input sample; wherein the third threshold increases with the increase of the number of training rounds in the third stage.

[0115] In this implementation, if the third random number is less than the third threshold, and the third input sample contains real motion trajectories (i.e., the third input sample comes from real data), then the real motion trajectories are removed from the third input sample. If the third input sample contains simulated motion trajectories (i.e., the third input sample comes from simulated data), then the simulated motion trajectories are removed from the third input sample. The third threshold is a set value, and the third random number is a random value within a set range. As the number of training rounds increases, the third random number gradually increases until it reaches the third threshold and then remains constant. Ultimately, this allows the video generation model to generate high-quality videos without relying on motion trajectories at all.

[0116] S2032 and S2033 can be referred to as S2022 and S2023 in the aforementioned embodiments, and will not be described again.

[0117] In S2034, the second video generation loss value and the second trajectory generation loss value are determined based on the second output video, the second output motion trajectory, and the third sample label. When the third input sample and the third sample label are from real data, the third sample label includes the real video and the real motion trajectory. The second video generation loss value can be determined based on the difference between the second output video and the real video, and the second trajectory generation loss value can be determined based on the difference between the second output motion trajectory and the real motion trajectory. When the third input sample and the third sample label are from simulated data, the third sample label includes the simulated video, the initial frame of the real video, and the simulated motion trajectory. The second video generation loss value can be determined based on the difference between the second output video and the simulated video, and the difference between the second output video and the initial frame. The second trajectory generation loss value can be determined based on the difference between the second output motion trajectory and the simulated motion trajectory.

[0118] The determination of the second video generation loss value based on the difference between the second output video and the real video, the determination of the second trajectory generation loss value based on the difference between the second output motion trajectory and the real motion trajectory, the determination of the second video generation loss value based on the difference between the second output video and the simulated video and the difference between the second output video and the initial frame, and the determination of the second trajectory generation loss value based on the difference between the second output motion trajectory and the simulated motion trajectory can be referred to the following processes in the foregoing embodiments: determining the first video generation loss value based on the difference between the first output video and the real video, determining the first trajectory generation loss value based on the difference between the first output motion trajectory and the real motion trajectory, determining the first video generation loss value based on the difference between the first output video and the simulated video and the difference between the first output video and the initial frame, and determining the first trajectory generation loss value based on the difference between the first output motion trajectory and the simulated motion trajectory, will not be elaborated further.

[0119] In one possible implementation, S2034 includes: when the third input sample comes from real data, obtaining a second video generation loss value by calculating the reconstruction loss between the second output video and the real video, and obtaining a second trajectory generation loss value by calculating the reconstruction loss between the second output motion trajectory and the real motion trajectory; when the third input sample comes from simulated data, obtaining a second video generation loss value by calculating the reconstruction loss between the second output video and the simulated video, the semantic consistency loss between the second output video and the simulated video, and the perceptual loss between the video frames of the second output video and the initial frame of the real video, and obtaining a second trajectory generation loss value by calculating the reconstruction loss between the second output motion trajectory and the simulated motion trajectory. Thus, on the one hand, by applying the motion trajectory reconstruction loss to the training of the video generation model, after the third stage of training, the video generation model can not only ensure that it generates high-quality videos of the agent executing language instructions, but also output accurate feature data to the action generation model so that the action generation model can generate accurate motion trajectories of the agent executing language instructions. On the other hand, when the third input data comes from simulation data, by combining the reconstruction loss between the second output video and the simulation video, the semantic consistency loss between the second output video and the simulation video, and the perceptual loss between the video frames of the second output video and the initial frames of the real video, the second video generation loss value of the video generation model is determined. This ensures that after training, the video generation model generates videos that are close to the simulation video in terms of image pixels and image semantics, and close to the real video in terms of visual perception, thus guaranteeing the realism of the videos generated by the video generation model.

[0120] The calculation process of the second video generation loss value and the second trajectory generation loss value can refer to the calculation process of the first video generation loss value and the first trajectory generation loss value in the foregoing embodiments, and will not be repeated here.

[0121] As an example, Figure 8 shows a framework example for model training in the third stage. As shown in Figure 8, firstly, the initial frame of the real video is input into the instruction generation model (Figure 8 uses VLM as an example) to generate hypothetical instructions. In the simulation environment, simulation videos and simulation motion trajectories can be generated based on the hypothetical instructions, where the hypothetical instructions, simulation videos, and simulation motion trajectories form simulation data. Next, taking the third input sample and the third sample label as coming from the simulation data as an example, in the third stage, the initial frame containing the hypothetical instructions, simulation motion trajectories (which can be removed according to the aforementioned embodiment, gradually stopping the input of simulation motion trajectories), and the real video is used as the third input sample and input into the video generation model to obtain the hidden layer output features and the second output video of the video generation model. The hidden layer output features are input into the motion generation model to obtain the second output motion trajectory of the motion generation model. Based on the second output video, the simulation video, and the initial frame of the real video, the second video generation loss value is calculated. By comparing the second output motion trajectory and the simulation motion trajectory, the second trajectory generation loss value is obtained. The parameters of the motion generation model are adjusted according to the second video generation loss value and the second trajectory generation loss value.

[0122] Please refer to Figure 9, which is a flowchart illustrating the training data augmentation method provided in this application embodiment. As shown in Figure 6, the training data augmentation method includes at least the following steps S901 to S904: S901, obtaining real video frames from the original training data of the agent's model. The real video frames are obtained by video capture during the process of the agent executing real instructions in a real environment.

[0123] The model of the intelligent agent is the model that controls the intelligent agent to execute language instructions.

[0124] The original training data can be obtained by teleoperating the agent in a real environment. The original training tool includes real video recording the dynamic process of the agent executing real commands; the real video frames can be the initial frames of the real video. The data acquisition in the real environment and the initial frames of the real video are described in the foregoing embodiments and will not be repeated here.

[0125] In this step, multiple real video frames can be obtained from the original training data to provide the simulation environment with a variety of real video frames, thus providing a rich real-world reference for the simulation environment.

[0126] S902 generates hypothetical instructions based on real video frames. These hypothetical instructions are task instructions that the agent can execute in the task scenario described by the real video frames.

[0127] In this step, one or more hypothetical instructions can be generated based on a single real video frame. If multiple real video frames are available, multiple hypothetical instructions can be generated. The process of generating hypothetical instructions based on real video frames can be referred to the relevant descriptions in the preceding embodiments, and will not be repeated here.

[0128] S903 inputs real video frames and hypothetical instructions into the trained embodied world model, and generates hypothetical videos and hypothetical action trajectories through the trained embodied world model.

[0129] The trained embodied world model is obtained by training according to the model training method provided in any of the foregoing embodiments.

[0130] In this step, real video frames and hypothetical instructions are input into the trained embodied world model. The real video frames provide the embodied world model with initial environmental information for the agent to execute the hypothetical instructions. In the embodied world model, based on the real video frames and hypothetical instructions, a hypothetical video can be generated by a video generation model, and a hypothetical action trajectory can be generated by an action generation model.

[0131] The video generation model, the motion generation model, and the data input-output relationship between the video generation model and the motion generation model can be referred to the description in the foregoing embodiments, and will not be repeated here.

[0132] S904 identifies hypothetical instructions, hypothetical videos, and hypothetical action trajectories as new training data for the agent's model.

[0133] In this step, a hypothetical instruction generates one new training data point. Through the above steps, a large number of hypothetical instructions can be generated. Based on this large number of hypothetical instructions, a large number of new training data points can be generated, providing rich training data for the agent's model.

[0134] In this embodiment, the embodied world model trained using the model training method provided in any of the foregoing embodiments generates new training data for the agent model. On the one hand, as can be seen from the foregoing embodiments, the embodied world model can generate training data containing language instructions, videos, and motion trajectories, and can also ensure video quality, video naturalness, and motion trajectory accuracy, thereby improving the data quality of the training data. On the other hand, in the process of generating new training data, the hypothetical instructions are generated based on the real environment indicated by the real video frames, and the instructions are not generated outside the real environment. The hypothetical instructions are executable and conform to the real situation, so that the generated training data is also usable data that conforms to the real situation.

[0135] As an example, Figure 10 is a schematic diagram of the structure of the embodied world model provided in an embodiment of this application. As shown in Figure 10, the embodied world model includes a video generation model, an action generation model, a video encoder, and a text encoder. The video encoder encodes real video frames to obtain the encoded vector of the real video frames, and the text encoder encodes hypothetical instructions to obtain the encoded vector of the hypothetical instructions. The encoded vectors of the real video frames and the encoded vectors of the hypothetical instructions are input into the video generation model. In the video generation model, features are extracted through multiple hidden layers to obtain the output features of the hidden layers. The output features of the last hidden layer are input into the output layer to obtain the hypothetical video. In the action generation model, a randomly generated noise vector and the output features of the hidden layers from the video generation model can be input. Features are extracted through multiple hidden layers in the action generation model, and finally, after passing through the output layer, the hypothetical action trajectory is obtained.

[0136] Please refer to Figure 11, which is a schematic diagram of the model training device provided in this application embodiment. As shown in Figure 11, the model training device 1100 includes: a first acquisition unit 1101, used to acquire real data, including real instructions, real videos, and real motion trajectories. The real videos and real motion trajectories are obtained by data collection during the process of an agent executing real instructions in a real environment; a second acquisition unit 1102, used to acquire simulation data, including hypothetical instructions, simulation videos, and simulation motion trajectories. The hypothetical instructions are generated based on the initial frame of the real video, and the simulation videos and simulation motion trajectories are obtained by simulating the agent executing hypothetical instructions; and a training unit 1103, used to train an initial embodied world model based on the real data and simulation data to obtain a trained embodied world model. The initial embodied world model includes an initial video generation model and an initial motion generation model, and the trained embodied world model includes a trained video generation model and a trained motion generation model. The trained embodied world model is used to generate training data containing language instructions, videos, and motion trajectories for the model training of the agent.

[0137] In some embodiments, the initial embodied world model is trained in stages. The training unit is specifically used for: in the first stage, training the initial video generation model multiple times based on real data and simulation data to obtain the video generation model trained in the first stage; in the second stage, training the initial motion generation model multiple times based on real data, simulation data, and the video generation model trained in the first stage to obtain the motion generation model trained in the second stage, while the model parameters of the video generation model trained in the first stage remain unchanged; in the third stage, training the video generation model trained in the first stage multiple times based on real data, simulation data, and the motion generation model trained in the second stage to obtain the video generation model trained in the third stage, while the model parameters of the motion generation model trained in the second stage remain unchanged; wherein, the trained embodied world model includes the video generation model trained in the third stage and the motion generation model trained in the second stage.

[0138] In some embodiments, in the first stage, i is greater than or equal to 1, and the i-th round of training of the initial video generation model includes: determining a first input sample and a first sample label corresponding to the first input sample in real data and simulation data, wherein the first input sample includes real instructions, real motion trajectories and initial frames, and the first sample label includes real video; or, the first input sample includes hypothetical instructions, simulated motion trajectories and initial frames, and the first sample label includes simulated video and initial frames; inputting the first input sample into the video generation model undergoing the i-th round of training, performing video generation in the video generation model undergoing the i-th round of training, and obtaining a first output video; determining a first video generation loss value based on the first output video and the first sample label; and adjusting the parameters of the video generation model undergoing the i-th round of training based on the first video generation loss value, to obtain the video generation model after the i-th round of training.

[0139] In some embodiments, the first input sample is input into the video generation model for the i-th round of training. Before the video generation model for the i-th round of training generates the first output video, the i-th round of training further includes: generating a first random number; if the first random number is less than a first threshold, removing the real motion trajectory from the first input sample, or removing the simulated motion trajectory from the first input sample.

[0140] In some embodiments, determining a first video generation loss value based on a first output video and a first sample label includes: when the first input sample comes from real data, obtaining a first video generation loss value by calculating the reconstruction loss between the first output video and the real video; when the first input sample comes from simulated data, obtaining a first video generation loss value by calculating the reconstruction loss between the first output video and the simulated video, the semantic consistency loss between the first output video and the simulated video, and the perceptual loss between the video frames of the first output video and the initial frame.

[0141] In some embodiments, in the second stage, j is greater than or equal to 1, and the j-th round of training of the initial action generation model includes: determining a second input sample and a second sample label corresponding to the second input sample in real data and simulation data, wherein the second input sample includes a real instruction, a real action trajectory, and an initial frame, and the second sample label includes the real action trajectory; or, the second input sample includes a hypothetical instruction, a simulated action trajectory, and an initial frame, and the second sample label includes the simulated action trajectory; inputting the second input sample into the video generation model trained in the first stage, and performing video generation in the video generation model trained in the first stage to obtain the hidden layer output features of the video generation model; inputting a randomly generated noise vector into the input layer of the action generation model trained in the j-th round, and using the hidden layer output features as generation condition inputs into the hidden layer of the action generation model trained in the j-th round, and performing action trajectory generation in the action generation model trained in the j-th round to obtain a first output action trajectory; determining a first trajectory generation loss value based on the first output action trajectory and the second sample label; and adjusting the parameters of the action generation model trained in the j-th round based on the first trajectory generation loss value to obtain the action generation model trained in the j-th round.

[0142] In some embodiments, before inputting the second input sample into the video generation model trained in the first stage and generating video in the video generation model trained in the first stage to obtain the hidden layer output features of the video generation model, the j-th round of training further includes: generating a second random number; if the second random number is less than a second threshold, removing the real motion trajectory from the second input sample, or removing the simulated motion trajectory from the second input sample.

[0143] In some embodiments, determining the trajectory generation loss value based on the first output motion trajectory and the second sample label includes: when the second input sample comes from real data, obtaining the first trajectory generation loss value by calculating the reconstruction loss between the first output motion trajectory and the real motion trajectory; and when the second input sample comes from simulation data, obtaining the first trajectory generation loss value by calculating the reconstruction loss between the first output motion trajectory and the simulated motion trajectory.

[0144] In some embodiments, in the third stage, k is greater than or equal to 1, and the k-th round of training of the video generation model after the first stage training includes: determining a third input sample and a corresponding third sample label from real data and simulated data; the third input sample includes a real instruction, a real motion trajectory, and an initial frame; the third sample label includes a real motion trajectory and a real video; or, the third input sample includes a hypothetical instruction, a simulated motion trajectory, and an initial frame; the third sample label includes a simulated motion trajectory, a simulated video, and an initial frame; inputting the third input sample into the video generation model undergoing the k-th round of training, and generating a video in the video generation model undergoing the k-th round of training to obtain the video. The hidden layer output features and the second output video of the generation model are generated. A randomly generated noise vector is input into the input layer of the action generation model after the second stage of training. The hidden layer output features are used as generation conditions and input into the hidden layer of the action generation model after the second stage of training. Action trajectories are generated in the action generation model after the second stage of training to obtain the second output action trajectory. Based on the second output video, the second output action trajectory, and the third sample label, the second video generation loss value and the second trajectory generation loss value are determined. Based on the second video generation loss value and the second trajectory generation loss value, the parameters of the video generation model trained in the k-th round are adjusted to obtain the video generation model after the k-th round of training.

[0145] In some embodiments, after determining the third input sample and the third sample label corresponding to the third input sample in real data and simulation data, the k-th round of training further includes: generating a third random number; if the third random number is less than a third threshold, removing the real motion trajectory from the third input sample, or removing the simulated motion trajectory from the third input sample; wherein the third threshold increases with the increase of the number of training rounds in the third stage.

[0146] In some embodiments, determining the second video generation loss value and the second trajectory generation loss value based on the second output video, the second output motion trajectory, and the third sample label includes: when the third input sample comes from real data, obtaining the second video generation loss value by calculating the reconstruction loss between the second output video and the real video, and obtaining the second trajectory generation loss value by calculating the reconstruction loss between the second output motion trajectory and the real motion trajectory; when the third input sample comes from simulated data, obtaining the second video generation loss value by calculating the reconstruction loss between the second output video and the simulated video, the semantic consistency loss between the second output video and the simulated video, and the perceptual loss between the video frames of the second output video and the initial frame, and obtaining the second trajectory generation loss value by calculating the reconstruction loss between the second output motion trajectory and the simulated motion trajectory.

[0147] The model training device 1100 provided in this application embodiment can execute the technical solution shown in the embodiment of the above model training method. Its implementation principle and beneficial effects are similar, and will not be described again here.

[0148] Please refer to Figure 12, which is a schematic diagram of the training data augmentation device provided in an embodiment of this application. As shown in Figure 12, the training data augmentation device 1200 includes: an acquisition unit 1201, used to acquire real video frames from the original training data of the agent's model, wherein the real video frames are obtained by video capture during the process of the agent executing real instructions in a real environment; an instruction generation unit 1202, used to generate hypothetical instructions based on the real video frames, wherein the hypothetical instructions belong to the task instructions that the agent supports executing in the task scenario described by the real video frames; a data generation unit 1203, used to input the real video frames and hypothetical instructions into the trained embodied world model, and generate hypothetical videos and hypothetical action trajectories through the trained embodied world model, wherein the trained embodied world model is trained according to the model training method provided in any of the above embodiments; and a training data determination unit 1204, used to determine the hypothetical instructions, hypothetical videos, and hypothetical action trajectories as new training data for the agent's model.

[0149] The training data augmentation device 1200 provided in this application embodiment can execute the technical solution shown in the above-described training data augmentation method embodiment. Its implementation principle and beneficial effects are similar, and will not be repeated here.

[0150] Please refer to Figure 13, which is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. As shown in Figure 13, the electronic device 1300 includes: one or more processors 1310; and a memory 1320 storing one or more programs. When the one or more programs are executed by the one or more processors 1310, the one or more processors 1310 implement the model training method described in any of the above embodiments.

[0151] Memory 1320, as a non-transitory network system, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory 1320 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory 1320 may optionally include remotely located memories 1320 relative to processor 1310, which can be connected to processor 1310 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0152] The memory 1320 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1320 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1320 and is called and executed by the processor 1310.

[0153] The processor 1310 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0154] In some embodiments, the electronic device further includes: an input / output interface for inputting and outputting information; a communication interface for communication and interaction between the device and other devices, which can be implemented via wired means (e.g., USB, Ethernet cable, etc.) or wireless means (e.g., mobile network, WIFI, Bluetooth, etc.); and a bus for transmitting information between various components of the device (e.g., processor 1310, memory 1320, input / output interface, and communication interface); wherein the processor 1310, memory 1320, input / output interface, and communication interface can be interconnected within the device via the bus.

[0155] One embodiment of this application also provides a computer-readable storage medium storing computer-executable instructions for performing the model training method or training data augmentation method described in any of the embodiments above.

[0156] An embodiment of this application also provides a computer program product, including a computer program or computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer program or computer instructions from the computer-readable storage medium and executes the computer program or computer instructions, causing the computer device to perform a model training method or training data augmentation method as described in any of the above embodiments.

[0157] The system architecture and application scenarios described in this application are intended to more clearly illustrate the technical solutions of this application and do not constitute a limitation on the technical solutions provided in this application. Those skilled in the art will understand that as system architectures evolve and new application scenarios emerge, the technical solutions provided in this application are also applicable to similar technical problems.

[0158] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0159] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0160] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0161] The above description, with reference to the accompanying drawings, illustrates some embodiments of this application, but does not limit the scope of the invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and spirit of this invention should be considered within the scope of this application.

Claims

1. A model training method, characterized in that, include: Acquire real data, which includes real commands, real videos, and real action trajectories. The real videos and real action trajectories are obtained by collecting data during the process of the intelligent agent executing the real commands in a real environment. Simulation data is acquired, including hypothetical instructions, simulation videos, and simulation motion trajectories. The hypothetical instructions are generated based on the initial frames of the real video, and the simulation videos and motion trajectories are obtained by simulating the agent executing the hypothetical instructions. Based on the real data and the simulation data, an initial embodied world model is trained to obtain a trained embodied world model. The initial embodied world model includes an initial video generation model and an initial motion generation model, and the trained embodied world model includes a trained video generation model and a trained motion generation model. The trained embodied world model is used to generate training data containing language instructions, videos, and motion trajectories for the agent's model training.

2. The model training method according to claim 1, characterized in that, The initial embodied world model is trained in stages. The training of the initial embodied world model based on the real data and the simulation data to obtain the trained embodied world model includes: in a first stage, training the initial video generation model multiple times based on the real data and the simulation data to obtain the video generation model trained in the first stage; in a second stage, training the initial motion generation model multiple times based on the real data, the simulation data, and the video generation model trained in the first stage to obtain the motion generation model trained in the second stage, where the model parameters of the video generation model trained in the first stage remain unchanged; in a third stage, training the video generation model trained in the first stage multiple times based on the real data, the simulation data, and the motion generation model trained in the second stage to obtain the video generation model trained in the third stage, where the model parameters of the motion generation model trained in the second stage remain unchanged; wherein, the trained embodied world model includes the video generation model trained in the third stage and the motion generation model trained in the second stage.

3. The model training method according to claim 2, characterized in that, In the first stage, i is greater than or equal to 1. The i-th round of training of the initial video generation model includes: determining a first input sample and a first sample label corresponding to the first input sample from the real data and the simulation data. The first input sample includes the real instruction, the real motion trajectory, and the initial frame. The first sample label includes the real video. Alternatively, the first input sample includes the hypothetical instruction, the simulated motion trajectory, and the initial frame. The first sample label includes the simulated video and the initial frame. The first input sample is input into the video generation model undergoing the i-th round of training. Video generation is performed in the video generation model undergoing the i-th round of training to obtain a first output video. A first video generation loss value is determined based on the first output video and the first sample label. The parameters of the video generation model undergoing the i-th round of training are adjusted based on the first video generation loss value to obtain the video generation model after the i-th round of training.

4. The model training method according to claim 3, characterized in that, In the step of inputting the first input sample into the video generation model for the i-th round of training, before generating the video and obtaining the first output video in the video generation model for the i-th round of training, the i-th round of training further includes: generating a first random number; if the first random number is less than a first threshold, removing the real motion trajectory from the first input sample, or removing the simulated motion trajectory from the first input sample.

5. The model training method according to claim 3, characterized in that, The step of determining the first video generation loss value based on the first output video and the first sample label includes: when the first input sample comes from the real data, obtaining the first video generation loss value by calculating the reconstruction loss between the first output video and the real video; when the first input sample comes from the simulation data, obtaining the first video generation loss value by calculating the reconstruction loss between the first output video and the simulation video, the semantic consistency loss between the first output video and the simulation video, and the perceptual loss between the video frame of the first output video and the initial frame.

6. The model training method according to any one of claims 2 to 5, characterized in that, In the second stage, j is greater than or equal to 1. The j-th round of training of the initial action generation model includes: determining a second input sample and a second sample label corresponding to the second input sample from the real data and the simulation data. The second input sample includes the real instruction, the real action trajectory, and the initial frame. The second sample label includes the real action trajectory. Alternatively, the second input sample includes the hypothetical instruction, the simulated action trajectory, and the initial frame. The second sample label includes the simulated action trajectory. The second input sample is input into the video generation model trained in the first stage. In the model, video generation is performed to obtain the hidden layer output features of the video generation model; a randomly generated noise vector is input into the input layer of the action generation model trained in the j-th round, and the hidden layer output features are used as generation conditions input into the hidden layer of the action generation model trained in the j-th round. Action trajectory generation is performed in the action generation model trained in the j-th round to obtain a first output action trajectory; a first trajectory generation loss value is determined based on the first output action trajectory and the second sample label; based on the first trajectory generation loss value, the parameters of the action generation model trained in the j-th round are adjusted to obtain the action generation model after the j-th round of training.

7. The model training method according to claim 6, characterized in that, Before inputting the second input sample into the video generation model trained in the first stage, and generating video in the video generation model trained in the first stage to obtain the hidden layer output features of the video generation model, the j-th round of training further includes: generating a second random number; if the second random number is less than a second threshold, removing the real motion trajectory from the second input sample, or removing the simulated motion trajectory from the second input sample.

8. The model training method according to claim 6, characterized in that, The step of determining the trajectory generation loss value based on the first output motion trajectory and the second sample label includes: when the second input sample comes from the real data, obtaining the first trajectory generation loss value by calculating the reconstruction loss between the first output motion trajectory and the real motion trajectory; and when the second input sample comes from the simulation data, obtaining the first trajectory generation loss value by calculating the reconstruction loss between the first output motion trajectory and the simulated motion trajectory.

9. The model training method according to any one of claims 2 to 5, characterized in that, In the third stage, k is greater than or equal to 1. The k-th round of training of the video generation model after the first stage includes: determining a third input sample and a corresponding third sample label from the real data and the simulation data; the third input sample includes the real instruction, the real motion trajectory, and the initial frame; the third sample label includes the real motion trajectory and the real video; or, the third input sample includes the hypothetical instruction, the simulated motion trajectory, and the initial frame; the third sample label includes the simulated motion trajectory, the simulated video, and the initial frame; inputting the third input sample into the video generation model undergoing the k-th round of training; and performing video generation within the video generation model undergoing the k-th round of training. The process involves generating the hidden layer output features and the second output video of the video generation model. A randomly generated noise vector is input into the input layer of the action generation model after the second stage of training. The hidden layer output features are used as generation conditions and input into the hidden layer of the action generation model after the second stage of training. Action trajectories are generated in the action generation model after the second stage of training to obtain the second output action trajectory. Based on the second output video, the second output action trajectory, and the third sample label, a second video generation loss value and a second trajectory generation loss value are determined. Based on the second video generation loss value and the second trajectory generation loss value, the parameters of the video generation model trained in the k-th round are adjusted to obtain the video generation model after the k-th round of training.

10. The model training method according to claim 9, characterized in that, After determining the third input sample and the third sample label corresponding to the third input sample in the real data and the simulation data, the k-th round of training further includes: generating a third random number; if the third random number is less than a third threshold, removing the real motion trajectory from the third input sample, or removing the simulated motion trajectory from the third input sample; wherein the third threshold increases with the increase of the number of training rounds in the third stage.

11. The model training method according to claim 9, characterized in that, The step of determining the second video generation loss value and the second trajectory generation loss value based on the second output video, the second output motion trajectory, and the third sample label includes: when the third input sample comes from the real data, obtaining the second video generation loss value by calculating the reconstruction loss between the second output video and the real video, and obtaining the second trajectory generation loss value by calculating the reconstruction loss between the second output motion trajectory and the real motion trajectory; when the third input sample comes from the simulation data, obtaining the second video generation loss value by calculating the reconstruction loss between the second output video and the simulation video, the semantic consistency loss between the second output video and the simulation video, and the perceptual loss between the video frames of the second output video and the initial frame, and obtaining the second trajectory generation loss value by calculating the reconstruction loss between the second output motion trajectory and the simulation motion trajectory.

12. A training data augmentation method, characterized in that, include: Real video frames are obtained from the original training data of the agent model. These real video frames are obtained by video capture during the process of the agent executing real instructions in a real environment. Based on the real video frames, hypothetical instructions are generated, which belong to the task instructions that the agent supports executing in the task scenario described in the real video frames; the real video frames and the hypothetical instructions are input into the trained embodied world model, and hypothetical videos and hypothetical action trajectories are generated through the trained embodied world model, which is trained by the model training method according to any one of claims 1 to 11; the hypothetical instructions, the hypothetical videos, and the hypothetical action trajectories are determined as new training data for the agent's model.

13. A model training device, characterized in that, include: The first acquisition unit is used to acquire real data, which includes real instructions, real videos, and real action trajectories. The real videos and real action trajectories are obtained by data collection during the process of the intelligent agent executing the real instructions in a real environment. The second acquisition unit is used to acquire simulation data, which includes hypothetical instructions, simulation videos, and simulation action trajectories. The hypothetical instructions are generated based on the initial frames of the real video, and the simulation videos and simulation action trajectories are obtained by simulating the execution of the hypothetical instructions by the intelligent agent. The training unit is used to train an initial embodied world model based on the real data and the simulation data to obtain a trained embodied world model. The initial embodied world model includes an initial video generation model and an initial motion generation model. The trained embodied world model includes a trained video generation model and a trained motion generation model. The trained embodied world model is used to generate training data containing language instructions, videos, and motion trajectories for the model training of the intelligent agent.

14. An electronic device comprising: One or more processors; A memory having stored one or more programs that, when executed by one or more processors, cause the one or more processors to implement, as in any one of claims 1 to 11, the model training method, or the training data augmentation method of claim 12.