Task processing method based on multi-modal world model, storage medium and equipment

Through the multimodal world model processing multimodal data, the problem of a single application scenario of the existing world model is solved, the universality and diversity processing capabilities of multimodal data are realized, and the prediction and decoding of multiple tasks can be carried out.

CN120337112APending Publication Date: 2025-07-18北京极佳视界科技有限公司

Patent Information

Application Number
CN202410064158.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing world models can usually only be processed for specific tasks in a single application scenario, cannot achieve universality and diversity, and cannot support the input and processing of multimodal data.

Method used

The multimodal world model is adopted to obtain multimodal data, encode each modal data using the encoder in the multimodal world model, mix features using the position encoder, and predict features without modality through the multimodal feature conversion model. Finally, the decoder is used for task processing, achieving the universality and ability diversity of multimodal data.

Benefits of technology

It realizes the universal processing capability of multimodal data, and can predict and decode multiple task results, including video prediction, three-dimensional generation, perceptual tasks, action prediction and speech prediction, which enhances the universality and diversity of world models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337112A_ABST
    Figure CN120337112A_ABST
Patent Text Reader

Abstract

The invention discloses a task processing method based on a multi-modal world model, a storage medium and equipment, and the method comprises the steps: obtaining multi-modal data which comprises data of at least one of the following modals: image data, voice data, text data, structured data and action data; encoding the various modal data in the multi-modal data by using encoders corresponding to the various modal data in the multi-modal world model to obtain features of the various modal data; mixing the features of the various modal data by using a position encoder in the multi-modal world model to obtain a first mixed feature; using a multi-modal feature conversion model in the multi-modal world model to predict a second mixed feature based on the first mixed feature, the second mixed feature including features of modal data not included in the multi-modal data; and decoding the features of the corresponding modes in the second mixed features by using a decoder of the corresponding task to obtain a task processing result of the corresponding task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to artificial intelligence technology and data processing technology, and in particular to a task processing method, a storage medium, and a device based on a multimodal world model. Background Art

[0002] A world model is an artificial intelligence system that predicts future events in an environment by constructing an internal representation of the environment.

[0003] In related technologies, data in the environment can be obtained through perception technologies such as computer vision and speech recognition, and then the data in the environment can be modeled, learned, and trained through technologies such as machine learning and deep learning, so as to form a world model that can recognize and understand the environment. Using this world model, the prediction and simulation of future events in the environment can be realized. However, existing world models are all trained using specific data of a specific modality (for example, driving actions) for a specific task in a single application scenario (for example, generating a future driving scenario), and the trained world model can only process the specific task for the data of the specific model and cannot be used in other application scenarios. Therefore, there are limitations in applications.

[0004] How to achieve the generality and diversity of capabilities of the world model has become a technical problem to be solved urgently. Summary of the Invention

[0005] To solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a task processing method, a storage medium, and a device based on a multimodal world model.

[0006] According to the first aspect of the embodiments of the present disclosure, a task processing method based on a multimodal world model is provided, including:

[0007] Obtain multimodal data, where the multimodal data includes at least one of the following types of data: image data, voice data, text data, structured data, action data;

[0008] Use the encoders corresponding to the various types of data in the multimodal world model to encode the various types of data in the multimodal data respectively to obtain the features of the various types of data;

[0009] Use the position encoder in the multimodal world model to mix the features of the various types of data to obtain a first mixed feature;

[0010] Use the multimodal feature conversion model in the multimodal world model to predict a second mixed feature based on the first mixed feature, where the second mixed feature includes the features of the types of data not included in the multimodal data;

[0011] Using the decoder corresponding to the task in the multi-modal world model, decode the features of the corresponding modality in the second hybrid feature to obtain the task processing result of the corresponding task, where the corresponding task includes at least one of the following: video prediction task, three-dimensional generation task, perception task, action prediction task, speech prediction task.

[0012] In some embodiments of the present disclosure, the encoding of various modality data in the multi-modal data using the encoders corresponding to the various modality data in the multi-modal world model to obtain the features of the multi-modal data includes:

[0013] In response to the multi-modal data including image data, input the image data into the visual encoder of the multi-modal world model for encoding to obtain visual features, where the image data includes at least one of the following: image, video;

[0014] In response to the multi-modal data including speech data, input the speech data into the speech encoder of the multi-modal world model for encoding to obtain audio features;

[0015] In response to the multi-modal data including text data, input the text data into the text encoder of the multi-modal world model for encoding to obtain text features;

[0016] In response to the multi-modal data including structured data or action data, input the structured data or action data into the structured encoder of the multi-modal world model for encoding to obtain structured features or action features.

[0017] In still other embodiments of the present disclosure, the position encoding mixing of the features of the multi-modal data using the position encoder in the multi-modal world model to obtain the first hybrid feature includes:

[0018] Obtain the position encoding parameters of the features of each modality data;

[0019] Sum the features of the modalities included in the multi-modal data with the corresponding position encoding parameters, and sum the features of the modalities not included in the multi-modal data after setting them to zero with the corresponding position encoding parameters to obtain the first hybrid feature.

[0020] In still other embodiments of the present disclosure, before the position encoding mixing of the features of the multi-modal data using the position encoder in the multi-modal world model, it further includes:

[0021] Unify the features of each modality data in the multi-modal data into features of the same dimension.

[0022] In some further embodiments of the present disclosure, the multi-modal world model is trained through the following steps:

[0023] Input various modal data in the training samples into the corresponding encoders, and respectively encode the various modal data in the training samples to obtain the features of at least one type of modal data;

[0024] Use the position encoder to perform position encoding mixing on the features of the at least one type of modal data to obtain a first training mixed feature;

[0025] Use the first training mixed feature to perform multiple rounds of self-supervised masked pre-training on the initial transformation network model until the training convergence condition is met, to obtain the multi-modal feature transformation model;

[0026] Combine the encoder, the multi-modal feature transformation model, the position encoder, and the decoder to form the multi-modal world model.

[0027] In some further embodiments of the present disclosure, the step of using the position encoder to perform position encoding mixing on the features of the at least one type of modal data to obtain a first training mixed feature includes:

[0028] Based on the network gradient, use the position encoder to generate the position encoding parameters of the features of the at least one type of modal data;

[0029] Sum the features of the modalities included in the training data and the corresponding position encoding parameters, and sum the features of the modalities not included in the training data after setting them to zero and the corresponding position encoding parameters, to obtain the first training mixed feature.

[0030] In some further embodiments of the present disclosure, the step of using the first training mixed feature to perform multiple rounds of self-supervised masked pre-training on the initial transformation network model until the training convergence condition is met, to obtain the multi-modal feature transformation model includes:

[0031] Input the first training mixed feature into the initial transformation network model. The initial transformation network model masks the first part of the features in the first training mixed feature, and predicts the masked features by using the method of masking and learning the second part of the features, where the second part of the features is the unmasked part of the first training mixed feature;

[0032] Based on the masked features and the first part of the features, calculate the self-supervised loss value;

[0033] Based on the self-supervised loss value, update the model parameters of the initial transformation network model;

[0034] Iteratively execute the above-mentioned masked pre-training process to implement the next round of masked pre-training until the predicted masked features and the first part of the features meet the training convergence condition, and obtain the multi-modal feature conversion model.

[0035] In some other embodiments of the present disclosure, the decoder for the corresponding task is obtained through the following operations:

[0036] Fine-tune the pre-trained video decoder using short video training samples to obtain a video decoder for the corresponding task;

[0037] Fine-tune the pre-trained 3D generation decoder using surround-view video training samples to obtain a 3D generation decoder for the corresponding task;

[0038] Fine-tune the pre-trained perception task decoder using image training samples to obtain a perception task decoder for the corresponding task;

[0039] Fine-tune the pre-trained action prediction decoder using action feature training samples to obtain an action prediction decoder for the corresponding task;

[0040] Fine-tune the pre-trained speech prediction decoder using speech feature training samples to obtain a speech prediction decoder for the corresponding task.

[0041] According to the second aspect of the present disclosure, there is provided a task processing device based on a multi-modal world model, including:

[0042] An acquisition module for acquiring multi-modal data, where the multi-modal data includes at least one of the following types of data: image data, voice data, text data, structured data, action data;

[0043] An encoding module for using the encoders corresponding to various types of modal data in the multi-modal world model to encode each type of modal data in the multi-modal data respectively to obtain the features of each type of modal data;

[0044] A feature mixing module for using the position encoder in the multi-modal world model to mix the features of each type of modal data to obtain a first mixed feature;

[0045] A feature prediction module for using the multi-modal feature conversion model in the multi-modal world model to predict a second mixed feature based on the first mixed feature, where the second mixed feature includes the features of the modal data not included in the multi-modal data;

[0046] A decoding module, configured to use a decoder corresponding to a task in the multimodal world model to decode features of a corresponding modality in the second hybrid feature, so as to obtain a task processing result of the corresponding task, where the corresponding task includes at least one of the following: video prediction task, three-dimensional generation task, perception task, action prediction task, speech prediction task.

[0047] In some other embodiments of the present disclosure, the encoding module includes:

[0048] A first encoding sub-module, configured to, in response to the multimodal data including image data, input the image data into a visual encoder of the multimodal world model for encoding to obtain visual features, where the image data includes at least one of the following: images, videos;

[0049] A second encoding sub-module, configured to, in response to the multimodal data including speech data, input the speech data into a speech encoder of the multimodal world model for encoding to obtain audio features;

[0050] A third encoding sub-module, configured to, in response to the multimodal data including text data, input the text data into a text encoder of the multimodal world model for encoding to obtain text features;

[0051] A fourth encoding sub-module, configured to, in response to the multimodal data including structured data or action data, input the structured data or action data into a structured encoder of the multimodal world model for encoding to obtain structured features or action features.

[0052] In some other embodiments of the present disclosure, the feature mixing module includes:

[0053] A parameter acquisition sub-module, configured to acquire position encoding parameters of features of each modality data;

[0054] A feature mixing sub-module, configured to sum the features of the modalities included in the multimodal data and the corresponding position encoding parameters, and sum the features of the modalities not included in the multimodal data after setting them to zero and the corresponding position encoding parameters, to obtain the first hybrid feature.

[0055] In some other embodiments of the present disclosure, the device further includes:

[0056] A feature dimension conversion module, configured to unify the features of each modality data in the multimodal data into features of the same dimension.

[0057] In some other embodiments of the present disclosure, the device further includes:

[0058] A model training module for training a multi-modal world model through the following steps;

[0059] The model training module includes:

[0060] An encoding sub-module for inputting various modal data in the training samples into corresponding encoders, respectively encoding the various modal data in the training samples, and obtaining features of at least one modal data;

[0061] A mixing sub-module for performing position encoding mixing on the features of the at least one modal data by using the position encoder to obtain a first training mixed feature;

[0062] A mask learning sub-module for performing multi-round self-supervised mask pre-training on the initial transformation network model by using the first training mixed feature until the training convergence condition is satisfied to obtain the multi-modal feature transformation model;

[0063] A merging sub-module for merging the encoder, the multi-modal feature transformation model, the position encoder, and the decoder to form the multi-modal world model.

[0064] In some other embodiments of the present disclosure, the mixing sub-module includes:

[0065] A parameter acquisition unit for generating position encoding parameters of the features of the at least one modal data by using the position encoder based on network gradients;

[0066] A position encoding mixing unit for summing the features of the modalities included in the training data with the corresponding position encoding parameters, and summing the features of the modalities not included in the training data after setting them to zero with the corresponding position encoding parameters to obtain the first training mixed feature.

[0067] In some other embodiments of the present disclosure, the mask learning sub-module includes:

[0068] A mask learning unit for inputting the first training mixed feature into the initial transformation network model, masking a first part of the features in the first training mixed feature by the initial transformation network model, and predicting the masked features by using a method of masking a second part of the features, where the second part of the features is the unmasked features in the first training mixed feature;

[0069] A loss calculation unit for calculating a self-supervised loss value based on the masked features and the first part of the features;

[0070] A parameter update unit for updating the model parameters of the initial transformation network model based on the self-supervised loss value;

[0071] An iterative unit is used to iteratively execute the above-mentioned masked pre-training process to implement the next round of masked pre-training until the predicted masked features satisfy the training convergence condition, and the multi-modal feature conversion model is obtained.

[0072] In some other embodiments of the present disclosure, the decoder for the corresponding task is obtained through the following operations:

[0073] Fine-tune the pre-trained video decoder with short video training samples to obtain a video decoder for the corresponding task;

[0074] Fine-tune the pre-trained 3D generation decoder with surround-view video training samples to obtain a 3D generation decoder for the corresponding task;

[0075] Fine-tune the pre-trained perception task decoder with image training samples to obtain a perception task decoder for the corresponding task;

[0076] Fine-tune the pre-trained action prediction decoder with action feature training samples to obtain an action prediction decoder for the corresponding task;

[0077] Fine-tune the pre-trained speech prediction decoder with speech feature training samples to obtain a speech prediction decoder for the corresponding task.

[0078] According to the third aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The storage medium stores computer program instructions, and when the computer program instructions are executed, the above-mentioned task processing method based on the multi-modal world model is implemented.

[0079] According to the fourth aspect of the embodiments of the present disclosure, an electronic device is provided. The electronic device includes:

[0080] A memory for storing a computer program product;

[0081] A processor for executing the computer program product stored in the memory, and when the computer program product is executed, the above-mentioned task processing method based on the multi-modal world model is implemented.

[0082] According to the fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including computer program instructions, and when the computer program instructions are executed by a processor, the above-mentioned task processing method based on the multi-modal world model is implemented.

[0083] A task processing method, storage medium, and device based on a multi-modal world model provided by the above embodiments of the present disclosure obtain multi-modal data, and then use the encoders corresponding to various modal data in the multi-modal world model to encode various modal data in the multi-modal data respectively to obtain the features of the multi-modal data. Then, use the position encoder in the multi-modal world model to perform position encoding mixing on the features of the multi-modal data to obtain a first mixed feature. Use the multi-modal feature conversion model in the multi-modal world model to predict a second mixed feature based on the first mixed feature. The second mixed feature includes the features of the modalities not included in the multi-modal data. Furthermore, use the decoder corresponding to the task in the multi-modal world model to perform decoding processing on the features of the corresponding modality in the second mixed feature to obtain the task processing result of the corresponding task. The corresponding task includes at least one of the following: video prediction task, three-dimensional generation task, perception task, action prediction task, and speech prediction task. Therefore, the technical solution of the present disclosure can predict the features of the modalities not included in the input multi-modal data by using the modality feature prediction function of the multi-modal feature conversion model in the multi-modal world model. Furthermore, it can realize obtaining the features of all modalities according to the input multi-modal data, and realize obtaining the task processing result of the corresponding task by decoding the features of the corresponding modality through the decoder corresponding to the task, which can ensure the versatility and diversity of capabilities of the multi-modal world model.

[0084] The technical solution of the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0085] By describing the embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation to the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.

[0086] Figure 1 It is a flowchart of an embodiment of the task processing method based on the multi-modal world model of the present disclosure;

[0087] Figure 2 It is a schematic diagram of the model training process of the multi-modal world model of the present disclosure;

[0088] Figure 3 It is a flowchart of an embodiment of the training of the multi-modal world model of the present disclosure;

[0089] Figure 4 It is a flowchart of the position encoding mixing in the training of the multi-modal world model of the present disclosure;

[0090] Figure 5 It is a training flow chart of a multi-modal feature conversion model in the training of the multi-modal world model of the present disclosure;

[0091] Figure 6 It is a schematic structural diagram of an embodiment of a task processing device based on a multi-modal world model of the present disclosure;

[0092] Figure 7 It is a schematic structural diagram of another embodiment of a task processing device based on a multi-modal world model of the present disclosure;

[0093] Figure 8 It is a structural diagram of an electronic device provided by an exemplary embodiment of the present application. Detailed implementation manners

[0094] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.

[0095] It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.

[0096] Those skilled in the art can understand that terms such as "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and do not represent any specific technical meaning nor indicate an inevitable logical order between them.

[0097] It should also be understood that in the embodiments of the present disclosure, "a plurality" may refer to two or more, and "at least one" may refer to one, two or more.

[0098] It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, without clear limitation or contrary indication in the context, it can generally be understood as one or more.

[0099] In addition, the term "and / or" in the present disclosure is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.

[0100] It should also be understood that the present disclosure emphasizes the differences between the various embodiments, and the same or similar parts can be referred to each other. For the sake of brevity, they will not be repeated one by one.

[0101] Meanwhile, it should be understood that, for the sake of description convenience, the dimensions of the various parts shown in the drawings are not drawn in accordance with actual proportional relationships.

[0102] The following description of at least one exemplary embodiment is in fact merely illustrative and in no way a limitation to the present disclosure, its application or use.

[0103] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods, and devices should be regarded as part of the specification.

[0104] It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, further discussion thereof is not required in subsequent drawings.

[0105] Embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, or servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, small computer systems, large computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.

[0106] Electronic devices such as terminal devices, computer systems, servers, etc. can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules may include routines, programs, object programs, components, logics, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment. In a distributed cloud computing environment, tasks can be executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.

[0107] Overview of the present disclosure

[0108] In the process of implementing the present disclosure, the inventors found that the world model usually only interacts with two types of modal inputs, namely visual signals and action signals, and does not support the input of modal data such as speech data, text data, and structured information (including 2D annotation boxes, 3D annotation boxes, etc.); moreover, the world model is usually only used to predict video and action data, and does not support downstream tasks such as three-dimensional object generation and perception tasks, which greatly limits the development and application of the world model.

[0109] The multi-modal world model provided by the present disclosure can receive the input of any type of modal data, and through the modal feature prediction function of the multi-modal feature conversion model in the multi-modal world model, predict the features of the modality that is not included in the input multi-modal data. Furthermore, it can realize obtaining the features of all modalities based on the input multi-modal data, and decode the features of the corresponding modality through the decoder of the corresponding task to obtain the task processing result of the corresponding task, ensuring the versatility and diversity of capabilities of the multi-modal world model.

[0110] Exemplary methods

[0111] Figure 1 It is a flowchart of an embodiment of the task processing method based on the multi-modal world model of the present disclosure. As Figure 1 shown, it includes step 101 - step 105. Each step will be described separately below.

[0112] In step 101, multi-modal data is obtained, and the multi-modal data includes at least one of the following types of data: image data, speech data, text data, structured data, and action data.

[0113] Among them, the multi-modal data can be data of one modality. In some optional embodiments, the multi-modal data can be text data. For example, the text "a little girl running" is input; in some other optional embodiments, the multi-modal data can also be image data. For example, an image of a little girl running is input; in some other optional embodiments, the multi-modal data can also be video data. For example, a video of a little girl running is input; in some other optional embodiments, the multi-modal data can also be structured data. For example, a 3D annotation box or a lane line is input; in some other optional embodiments, the multi-modal data can also be action data. For example, a driving action data is input.

[0114] Among them, the structured data and action data can be obtained from image or video data. For example, through image recognition, the structured data such as lane lines included in the image is recognized, or the lane lines and actions in the image are labeled by means of manual annotation.

[0115] Among them, the multi-modal data can also be data of more than two modalities, that is, it includes two or more of text data, image data, video data, motion data, structured data, and voice data.

[0116] As an example, the multi-modal data can include text data and image data at the same time. For example, input the text "A little girl running" and an image of a little girl running; or, the multi-modal data can also include text data, image data and video data at the same time. For example, input the text "A little girl running", an image of a little girl running, and a video of a little girl running, and so on.

[0117] Among them, image data can include image data, video data, etc.; voice data can include a piece of voice; structured data can include two-dimensional annotation boxes, three-dimensional annotation boxes, data in bird's eye view (Bird's Eye View, abbreviated as BEV), etc.; motion data can include information such as the angle, amplitude, and speed of the motion used to describe the motion, such as it can be a driving motion.

[0118] Among them, there are various ways to obtain multi-modal data. For example, receive multi-modal data input by the user, or obtain it from a database, or obtain it from a third-party application platform, etc.

[0119] In step 102, use the encoders corresponding to various modal data in the multi-modal world model to encode each modal data in the multi-modal data respectively to obtain the features of the various modal data.

[0120] Among them, the multi-modal world model can at least include components such as encoders corresponding to various modal data, position encoders, multi-modal feature conversion models, and decoders.

[0121] In some optional embodiments, the encoder is used to encode the input modal data to obtain the corresponding feature data, and can include a visual encoder, a text encoder, a structured encoder, a voice encoder, etc.

[0122] In some optional embodiments, in response to the multi-modal data including image data, the image data can be input into a visual encoder for encoding to obtain visual features, and the image data includes at least one of the following: image, video.

[0123] In specific implementation, VQGAN (Vector Quantized Generative Adversarial Network) can be used as a visual encoder, which can encode a single-frame image or single-frame video data into a latent space. For example, an image with a size of 256×256×3 can be encoded to obtain a two-dimensional feature vector with a size of 16×16×1024. Among them, the first 256 in the image size represents the height of the image as 256 pixels, the second 256 represents the length of the image as 256 pixels, 3 represents the RGB three-color channels, the first 16 represents the length of the two-dimensional feature vector, the second 16 represents the width of the two-dimensional feature vector, and 1024 represents the feature dimension of the two-dimensional feature vector. A video with a size of N×256×256×3 can also be encoded to obtain a temporal two-dimensional feature vector with a size of N×16×16×1024, where N indicates the number of frames of the video, 256×256×3 is the image size of each frame, and 16×16×1024 is the size of the two-dimensional feature vector of each frame.

[0124] In some alternative embodiments, in response to the multi-modal data including text data, the text data is input into a text encoder for encoding to obtain text features.

[0125] For text data, a pre-trained language model, such as the Text ToText Transfer Transformer (abbreviated as T5) model, can be used as the text encoder, which can encode the text data into text features, such as encoding the text data into a one-dimensional feature vector with a size of M×x1024, where M indicates the preset length of the text. If the actual length of the text is less than the preset length, it can be padded with 0. If the actual length of the text is greater than the preset length, the text part exceeding the preset length can be truncated. The specific value of M can be set in advance, such as it can be set to 256.

[0126] In some alternative embodiments, in response to the multi-modal data including structured data or action data, the structured data or action data is input into a pre-trained structured encoder for encoding to obtain structured features or action features.

[0127] For action data and structured data, a pre-trained structured encoder can be used for encoding. The pre-trained structured encoder can be trained using a Multi-Layer Perceptron (MLP for short). In specific implementation, a large number of training samples containing action data and structured data can be used to train the initial multi-layer perceptron. That is, by inputting a large number of training samples into the multi-layer perceptron, the multi-layer perceptron learns these action data and structured data to obtain corresponding action features and structured features. During the process of training these training samples with these training samples, the hyperparameters of the multi-layer perceptron can be fine-tuned. After the training is completed, the above-mentioned structured encoder is obtained from the above-mentioned multi-layer perceptron. When using the structured encoder to encode structured data or action data, corresponding structured features or action features can be obtained, such as a feature vector of size N×K×1024, where K is the preset length of each frame of structured data or action data. If the actual length is less than the preset length, it is padded with 0. If it is greater than the preset length, the text part exceeding the preset length can be truncated. The specific value of K can be set in advance, such as it can be set to 16 or other values, and N represents the number of frames of structured data or action data.

[0128] In some optional embodiments, in response to the multi-modal data including speech data, the speech data can be input into a speech encoder to obtain audio features.

[0129] For speech data, a pre-trained speech recognition model can be used as the speech encoder, and the speech encoder can encode the speech data into audio features.

[0130] In step 103, the position encoder in the multi-modal world model is used to mix the features of the various modal data to obtain a first mixed feature.

[0131] In some optional embodiments, the position encoder is a functional module for performing position encoding mixing on the features of each modal data, and can perform position encoding mixing on the visual features, text features, structured features, and action features obtained in step 102. By introducing position encoding parameters (position information) into the multi-modal features, it is possible to distinguish different modal features through the position encoding parameters, so that even when the feature order is disrupted, the category of the modal features can still be distinguished according to the position encoding parameters.

[0132] In some optional embodiments, the position encoding parameters of the features of each modal data can be obtained; then, the sum of the features of the modalities included in the multi-modal data and the corresponding position encoding parameters is calculated, and the sum of the features of the modalities not included in the multi-modal data after setting them to zero and the corresponding position encoding parameters is calculated to obtain the first mixed feature.

[0133] In some alternative embodiments, the position encoder in the multi-modal world model can set fixed position encoding parameters for various modalities. For example, the position encoding parameter for visual features is 0, the position encoding parameter for text features is 1, the position encoding parameter for structured features is 2, the position encoding parameter for action features is 3, and the position encoding parameter for audio features is 4.

[0134] In specific applications, the multi-modal data input into the multi-modal world model each time can include only one or more than two modalities of data. When performing position encoding, the features of the modalities not included in the multi-modal data can be set to zero, and then mixed with the features of the included modalities through position encoding to obtain the first mixed feature.

[0135] In specific implementation, before performing position encoding mixing, the multi-modal features can be unified into feature data of the same dimension. Usually, the visual features can be converted into one-dimensional feature vectors to align with text features, structured features, and action features. For example, the two-dimensional feature vector of 16×16×1024 of the image data is flattened into a one-dimensional feature vector of 1×256×1024. Thus, the visual features, text features, structured features, and action features all become one-dimensional feature vectors. The position encoder can generate corresponding position encoding parameters for these feature vectors, and then distinguish the features of different modalities by adding the features of each modality to the corresponding position encoding parameters.

[0136] In some alternative embodiments, the position encoding parameters can be added to the features of different modalities in sequence. For example, the feature order of each modality data is visual feature - text feature - structured feature - action feature. When adding the features of each modality to the corresponding position encoding parameters, the generated position encoding parameters can also be added to the corresponding modality features in the above feature order. After adding the position encoding parameters, different modality feature categories can be distinguished according to the position encoding parameters.

[0137] In some alternative embodiments, generating corresponding position encoding parameters for different modality features can distinguish the categories of modality features. For example, generating encoding 0 for visual features, generating encoding 1 for text features, generating encoding 2 for structured features, and generating encoding 3 for action features. For each vector in the same modality feature, a position encoding parameter can also be generated respectively. For example, the two-dimensional feature vector of an image is a feature vector of size 16×16×1024, and after flattening, it is a feature vector of size 1×256×1024. Among them, 256 features respectively correspond to pixels at different positions in the image. In order to distinguish different positions, a position encoding parameter can also be added to each of these 256 features for distinction.

[0138] After completing the positional encoding of each modal feature, the features of multiple modalities can be mixed to obtain the first mixed feature. For example, by concatenating the temporal two-dimensional feature vector N×16×16×1024, the text feature vector M×1024, and the structured feature vector N×K×1024, the first mixed feature with a size of (N×16×16 + M + N×K)×1024 can be obtained.

[0139] In some alternative embodiments, the features of multiple modalities can also be added to obtain the first mixed feature.

[0140] In step 104, using the multi-modal feature conversion model in the multi-modal world model, based on the first mixed feature, a second mixed feature is predicted, and the second mixed feature includes the features of the modal data not included in the multi-modal data.

[0141] In some alternative embodiments, the multi-modal feature conversion model is a model that can predict the features of other modalities based on a single modal feature. This multi-modal feature conversion model can be trained according to the initial conversion network model. For the specific training process, please refer to Figure 3 the embodiments shown, which will not be elaborated here for now.

[0142] In some alternative embodiments, after inputting the first mixed feature into the multi-modal feature conversion model, the second mixed feature can be predicted. For example, if the data input to the multi-modal world network only includes text data, the features obtained by the decoder only include text features. After mixing the text features through positional encoding, the first mixed feature is obtained, and the multi-modal feature conversion model can generate visual features, action features, and structured features from the text features.

[0143] In step 105, using the decoder corresponding to the task in the multi-modal world model, the features of the corresponding modality in the second mixed feature are decoded to obtain the task processing result of the corresponding task, and the corresponding task includes at least one of the following: video prediction task, three-dimensional generation task, perception task, action prediction task, speech prediction task, and multi-modal data.

[0144] In some alternative embodiments, the multi-modal world model can implement one or more than two downstream tasks according to any one of the input modal data. For example, one type of text data can be input for video prediction; or, one type of structured data can be input for action prediction, etc.

[0145] In some alternative embodiments, the multi-modal world model can implement one or more downstream tasks based on any two or more modalities of input data. When inputting data of two or more modalities, the content indicated by the data of the two or more modalities needs to be the same. For example, when inputting the text "a running little girl" and an image of a running little girl, although the text and the image are data of two modalities, the content indicated by the data of the two modalities is the same.

[0146] In some alternative embodiments, the video prediction task refers to the task of outputting predicted video data based on the input data; the 3D generation task refers to the task of outputting predicted 3D space objects based on the input data; the action prediction task refers to the task of outputting predicted actions based on the input data. For example, the action may include speed and steering angle; the perception task refers to the task of outputting information such as the position and classification of the target object based on the input data; the speech prediction task refers to the task of outputting predicted speech based on the input speech data.

[0147] Among them, which or which downstream tasks are specifically implemented based on the input data can be set according to the actual application scenario.

[0148] In some alternative embodiments, the decoder is used to implement different tasks. Among them, the video decoder is used to implement the video prediction task; the 3D generation decoder is used to implement the 3D generation task, such as outputting 3D objects; the perception task decoder is used to implement the perception task, such as outputting the classification and location information of the target object; the action prediction decoder is used to implement the action prediction task, such as outputting the predicted action; the speech prediction decoder is used to implement the speech prediction task, such as predicting the output speech, etc.

[0149] For the video decoder, in specific implementation, a diffusion model can be used as the model adopted by the video decoder. The diffusion model is a mathematical model used to describe and predict the propagation process of substances or information in space and time. By fine-tuning the initial video decoder using the diffusion model, the video decoder for the corresponding task can be obtained. A large number of short video datasets (WebVid) or Internet videos can be used to fine-tune the initial video decoder to obtain the video decoder for the corresponding task. The video decoder for the corresponding task can decode the input two-dimensional image features and output the predicted video.

[0150] As an example, the visual features output by the multi-modal feature conversion model can be restored to two-dimensional feature vectors of a two-dimensional image, and the restored two-dimensional feature vectors are input into the video decoder for the corresponding task, and this video decoder can output the predicted video. The resolution of the output video is upsampled by 16 times, and the effect is clear.

[0151] For the 3D generation decoder, in specific implementation, a diffusion model and a neural radiance field can be adopted as the models used by the 3D generation decoder. By fine-tuning the initial 3D generation decoder that uses the diffusion model and the neural radiance field with panoramic video training samples, the 3D generation decoder for the corresponding task can be obtained. A large amount of 3D panoramic data can be used as training samples to fine-tune the initial 3D generation decoder to obtain the 3D generation decoder for the corresponding task. The 3D generation decoder for the corresponding task can output the panoramic video of the 3D object for the input 3D panoramic data (multiple images from multiple perspectives).

[0152] Among them, the 3D panoramic training samples can be understood as multiple images from multiple perspectives or panoramic videos.

[0153] For the perception task decoder, the initial perception task decoder can be fine-tuned with image training samples to obtain the pre-trained perception task decoder. The perception task decoder in related technologies can be used as the initial perception task decoder, and then the initial perception task decoder can be fine-tuned with a large number of image training samples to obtain the perception task decoder for the corresponding task. The perception task decoder can adopt a deep convolutional neural network (Residual Network, ResNet) as the initial model or a target detection model (such as CenterNet) as the initial model. By inputting a large number of image training samples, the initial perception task decoder can be fine-tuned. The perception task decoder for the corresponding task can output information such as the classification and localization of the object for the input image features.

[0154] For the action prediction decoder, the initial action prediction decoder can be fine-tuned with action feature training samples to obtain the action prediction decoder for the corresponding task. The initial action prediction decoder can adopt a multi-layer perceptron as the model. In specific implementation, a large number of action feature training samples can be used to train the initial multi-layer perceptron, that is, by inputting a large number of action feature training samples into the multi-layer perceptron, the multi-layer perceptron can learn these action features to obtain the corresponding predicted actions. During the process of the multi-layer perceptron learning these training samples, the hyperparameters of the multi-layer perceptron can be fine-tuned. After training, the above-mentioned initial action decoder using the multi-layer perceptron can obtain the above-mentioned action prediction decoder for the corresponding task.

[0155] Using the average pooling operation, the action features (feature size: N×K×1024) output by the multi-modal feature conversion model can be pooled into (N, 1024). Then, using the pre-trained action prediction decoder to decode this part of the features, the predicted actions with a feature size of (N, 2) can be output. Here, 2 respectively represent the speed and the steering angle. According to the predicted actions, the robot can be controlled to execute corresponding actions or an autonomous vehicle can be controlled.

[0156] For the speech prediction decoder, the initial speech decoder can be fine-tuned using speech samples to obtain the speech decoder for the corresponding task. The speech prediction decoder for the corresponding task can output predicted speech information for the input audio features.

[0157] Through the above steps 101 - 105, by obtaining multi-modal data, and then using the encoders corresponding to various modal data in the multi-modal world model to encode various modal data in the multi-modal data respectively to obtain the features of the multi-modal data, and then using the position encoder in the multi-modal world model to perform position encoding mixing on the features of the multi-modal data to obtain the first mixed feature, and using the multi-modal feature conversion model in the multi-modal world model to predict the second mixed feature based on the first mixed feature, where the second mixed feature includes the features of the modalities not included in the multi-modal data. Furthermore, using the decoder for the corresponding task in the multi-modal world model to perform decoding processing on the features of the corresponding modality in the second mixed feature to obtain the task processing result for the corresponding task, and the corresponding task includes at least one of the following: video prediction task, 3D generation task, perception task, action prediction task, speech prediction task. Therefore, the technical solution of the present disclosure can predict the features of the modalities not included in the input multi-modal data by using the modality feature prediction function of the multi-modal feature conversion model in the multi-modal world model, and then can realize obtaining the features of all modalities according to the input multi-modal data, and realize obtaining the task processing result for the corresponding task by decoding the features of the corresponding modality through the decoder for the corresponding task, which can ensure the generality and diversity of capabilities of the multi-modal world model.

[0158] In some optional embodiments, the multi-modal world model can be pre-trained based on a large number of training samples, and the multi-modal world model can perform corresponding tasks for various modal data input.

[0159] Figure 3 The flowchart of an embodiment of the training of the multi-modal world model of the present disclosure is as Figure 3 shown, and the process of model training includes the following steps 301 - 304. Each step will be described separately below.

[0160] In step 301, various modal data in the training samples are input into the corresponding encoders, and various modal data in the training samples are encoded respectively to obtain the features of at least one modal data.

[0161] In some optional embodiments, a large number of training samples can be collected in advance, and each training sample can include at least one modal data. As Figure 2 shown, the training samples can include images, videos, texts, structured data, action data, and can also includeFigure 2 Speech data not shown therein. Among them, images and videos can be encoded by a visual encoder to obtain visual features, text can be encoded by a text encoder to obtain text features, while structured data and action data can be encoded by a shared structured encoder to obtain structured features and action features, and speech data can be encoded by a speech encoder to obtain audio features.

[0162] Among them, the encoders in the multimodal world model can include a visual encoder, a text encoder, a speech encoder, a structured encoder, a speech encoder, etc. Each encoder can be obtained in the manner described in step 102 of the illustrated embodiment, which will not be elaborated here. Figure 1 Shown in the embodiment described in step 102, which will not be elaborated here.

[0163] In step 302, the position encoder is used to perform position encoding mixing on the features of the at least one modality data to obtain a first training mixed feature.

[0164] In some optional embodiments, the position encoder is a functional module for performing position encoding mixing on the features of each modality data, and can perform position encoding mixing on the visual features, text features, structured features, and action features obtained in step 101. By introducing position encoding parameters (position information) into the multimodal features, it is possible to distinguish different modality features through the position encoding parameters, so that even when the feature order is disrupted, the category of the modality features can still be distinguished according to the position encoding parameters.

[0165] Specifically, before performing position encoding mixing, the multimodal features can be unified into feature data of the same dimension. Usually, the visual features can be converted into one-dimensional feature vectors to align with the text features, structured features, and action features. For example, the two-dimensional feature vector 16×16×1024 of the image data is flattened into a one-dimensional feature vector 1×16256×1024. Thus, the visual features, text features, structured features, and action features all become one-dimensional feature vectors. The position encoder can generate corresponding position encoding parameters for these feature vectors, and then by adding the features of each modality to the corresponding position encoding parameters, it is possible to distinguish the features of different modalities using position encoding.

[0166] Among them, during the training of the multimodal world model, the position encoding parameters generated by the position encoder are learnable position encodings. Learnable position encodings refer to treating the position encoding as a trainable parameter. For example, if the size of the input feature sequence is n×d, then a p∈R n×d matrix can be randomly initialized as the position encoding, and this matrix is updated along with the network gradient.

[0167] In some alternative embodiments, the position encoding parameters can be added to the features of different modalities in sequence. The feature order of each modality data can be visual features - text features - structured features - action features. When adding the features of each modality to the corresponding position encoding parameters, the generated position encoding parameters can also be added to the corresponding modality features in the above feature order. After adding the position encoding parameters, different modality feature categories can be distinguished according to the position encoding parameters.

[0168] In some alternative embodiments, generating corresponding position encoding parameters for different modality features can distinguish the modality feature categories. For example, generate position encoding 0 for visual features, position encoding 1 for text features, position encoding 2 for structured features, and position encoding 3 for action features. For each vector in the same modality feature, a position encoding parameter can also be generated respectively. For example, the two-dimensional feature vector of an image is a feature vector of size 16×16×1024. After flattening, it is a feature vector of size 1×256×1024. Among them, 256 features respectively correspond to pixels at different positions in the image. In order to distinguish different positions, a position encoding parameter can also be added to these 256 features respectively for distinction.

[0169] After completing the position encoding of each modality feature, the features of multiple modalities can be mixed. The first training mixed feature can be obtained by concatenating the features of various modalities, or by adding the features of various modalities.

[0170] In step 303, the initial transformation network model is pre-trained with self-supervised masking for multiple rounds using the first training mixed feature until the training convergence condition is met, and the multi-modal feature transformation model is obtained.

[0171] In some alternative embodiments, the initial transformation network model is a Transformer model. The first training mixed feature can be input into this initial transformation network model for masking learning. This model can use the method of masked pre-training to learn the features between multiple modalities. Specifically, during implementation, a mask can be set for a set proportion of the feature data in the first training mixed feature. For example, 50% of the feature data is set to 0, and then the remaining 50% of the feature data is used to predict the masked feature data to achieve self-supervised training.

[0172] See Figure 2 , by masking the feature data indicated by label 21 in the first training mixed feature after position encoding mixing of the position encoder, and then using the remaining unmasked feature data for the masked feature, self-supervised training is achieved.

[0173] In some alternative embodiments, the training convergence condition is used to indicate the condition for determining the end of the training process of the multi-modal feature conversion model.

[0174] In some alternative embodiments, the training convergence condition may include that the self-supervised loss value is less than a preset threshold. That is, when the self-supervised loss value is less than the preset threshold, it can be determined that the model training process ends, and the multi-modal feature conversion model is obtained from the initial conversion network model after updating the model parameters.

[0175] In some other alternative embodiments, the training convergence condition may further include that the current training round is equal to the set number of training times of the model. Thus, if the current training round is less than the set number of training rounds of the model, it can be determined that the model training process is not completed, and the next round of training process can be continued; if the current training round is equal to the set number of training rounds of the model, it can be determined that the model training process is completed.

[0176] In some alternative embodiments, the training convergence condition may further include that the current training round is not greater than the set number of training times of the model and the self-supervised loss value is less than the preset threshold. When the current training round is less than the set number of training rounds of the model, if the self-supervised loss value tends to be stable with respect to the self-supervised loss values obtained in the previous few rounds of training, such as no longer decreasing, it can also be determined that the model training process is completed and the model training is stopped.

[0177] Among them, the trained multi-modal feature conversion model can be used to predict the features of another modality from the features of one modality. For example, if the data input to the multi-modal world network only includes text data, the features obtained by the encoder only include text features, and the multi-modal feature conversion model can be used to predict the corresponding visual features, action features, and structured features from the text features.

[0178] In some alternative embodiments, by directly using the features after mixing the above position encodings to train the initial conversion network model, self-supervised training can be implemented to obtain the multi-modal feature conversion model.

[0179] In step 304, the encoder, the multi-modal feature conversion model, the position encoder, and the decoder are combined to form the multi-modal world model.

[0180] In some alternative embodiments, the pre-trained decoder is used to implement different downstream tasks. Among them, the video decoder is used to output the predicted video; the three-dimensional generation decoder is used to output three-dimensional objects, the perception task decoder is used to output the classification and localization information of the target object, and the action prediction decoder is used to output the predicted action.

[0181] In specific implementation, a diffusion model can be adopted as the model used by the video decoder. The diffusion model is a mathematical model for describing and predicting the propagation process of substances or information in space and time. By fine-tuning the initial video decoder using the diffusion model, a pre-trained video decoder can be obtained. In specific implementation, a large number of short video datasets (WebVid) or Internet videos can be used to fine-tune the initial video decoder to obtain a pre-trained video decoder. The pre-trained video decoder can decode the two-dimensional image features input and output a predicted video.

[0182] As an example, the visual features output by the multi-modal feature conversion model can be restored to two-dimensional feature vectors of a two-dimensional image, and the restored two-dimensional feature vectors are input into the video decoder for the corresponding task. This video decoder can then output a predicted video, and the resolution of the output video is upsampled by 16 times, with a clear effect.

[0183] In specific implementation, a diffusion model and a neural radiance field can be adopted as the models used by the three-dimensional generation decoder. By fine-tuning the initial three-dimensional generation decoder using the diffusion model and the neural radiance field with panoramic video training samples, a pre-trained three-dimensional generation decoder can be obtained. In specific implementation, a large number of three-dimensional panoramic data can be used as training samples to fine-tune the initial three-dimensional generation decoder to obtain a pre-trained three-dimensional generation decoder. The pre-trained three-dimensional generation decoder can input the three-dimensional panoramic data (multiple images from multiple perspectives) and output a panoramic video of a three-dimensional object.

[0184] Among them, the three-dimensional panoramic training samples can be understood as multiple images from multiple perspectives or a panoramic video.

[0185] For the perception task decoder, image training samples can be used to fine-tune the initial perception task decoder to obtain a pre-trained perception task decoder. The perception task decoder in related technologies can be used as the initial perception task decoder, and then a large number of image training samples are used to fine-tune the initial perception task decoder to obtain a pre-trained perception task decoder. The perception task decoder can adopt a deep convolutional neural network (Residual Network, ResNet) as the initial model or a target detection model (such as CenterNet) as the initial model. By inputting a large number of image training samples, the initial perception task decoder can be fine-tuned. The pre-trained perception task decoder can output information such as the classification and localization of objects for the input image features.

[0186] For the action prediction decoder, the initial action prediction decoder can be fine-tuned using action feature training samples to obtain a pre-trained action prediction decoder. The initial action prediction decoder can use a multi-layer perceptron as the model. In specific implementation, a large number of action feature training samples can be used to train the initial multi-layer perceptron. That is, by inputting a large number of action feature training samples into the multi-layer perceptron, the multi-layer perceptron learns these action features to obtain corresponding predicted actions. During the process of the multi-layer perceptron learning these training samples, the hyperparameters of the multi-layer perceptron can be fine-tuned. After training is completed, the above-mentioned pre-trained action prediction decoder is obtained from the initial action decoder using the multi-layer perceptron.

[0187] Using the average pooling operation, the action features (feature vector size is N×K×1024) output by the multi-modal feature conversion model can be pooled into (N, 1024). Then, the pre-trained action prediction decoder is used to decode this part of the features, and predicted actions with a feature size of (N, 2) can be output. Here, 2 represents speed and steering angle respectively. According to this predicted action, the robot can be controlled to execute corresponding actions or an autonomous vehicle.

[0188] By combining the multi-modal encoder, the multi-modal feature conversion model, the position encoder, and the pre-trained decoder, a pre-trained multi-modal world model can be constructed. This pre-trained multi-modal world model can implement corresponding downstream tasks based on the input multi-modal data.

[0189] Through the above steps 301 - step 304, the training method of the multi-modal world model is disclosed. By self-supervised learning the ability to predict various modal features from single-modal features during the training process, it can process the input of multi-modal data, and through multiple decoders, it can achieve the prediction of different downstream tasks, realizing the generality of the world model and the diversity of capabilities.

[0190] Figure 4 This is a flowchart of position encoding mixing in the training of the multi-modal world model of the present disclosure. As Figure 4 shown, the above step 302 can specifically include steps 321 - step 322. Each step will be introduced in detail below.

[0191] In step 321, based on the network gradient, the position encoding parameters of the features of the at least one modal data are generated using the position encoder.

[0192] Among them, the position encoding parameters generated by the position encoder are learnable position encodings. Learnable position encoding means treating the position encoding as a trainable parameter. For example, if the size of the input feature sequence is n×d, a p∈R n×d matrix can be randomly initialized as the position encoding, and this matrix is updated along with the network gradient.

[0193] In some alternative embodiments, the network gradient is used to indicate the model gradient during model training. The position encoder can generate corresponding position encoding parameters for the input feature sequence according to the network gradient.

[0194] In step 322, the features of the modalities included in the training data are summed with the corresponding position encoding parameters, and the features of the modalities not included in the training data are set to zero and then summed with the corresponding position encoding parameters to obtain the first training hybrid feature.

[0195] In some alternative embodiments, the position encoding parameters can be added to the features of different modalities in sequence. For example Figure 2 in which the feature order of each modality data is visual feature - text feature - structured feature - action feature. When adding the features of each modality to the corresponding position encoding parameters, the generated position encoding parameters can also be added to the corresponding modality features in the above feature order. After adding the position encoding parameters, different modality feature categories can be distinguished according to the position encoding parameters.

[0196] In some alternative embodiments, generating corresponding position encoding parameters for different modality features can distinguish the categories of modality features. For example, encoding 0 is generated for visual features, encoding 1 is generated for text features, encoding 2 is generated for structured features, and encoding 3 is generated for action features. For each vector in the same modality feature, a position encoding parameter can also be generated respectively. For example, the two - dimensional feature vector of an image is a feature vector of size 16×16×1024, and after flattening, it is a feature vector of size 1×256×1024, where 256 features respectively correspond to pixels at different positions in the image. In order to distinguish different positions, a position encoding parameter can also be added to these 256 features respectively for distinction.

[0197] After completing the position encoding of each modality feature, the features of multiple modalities can be mixed to obtain the first training hybrid feature. For example, by concatenating the temporal two - dimensional feature vector N×16×16×1024, the text feature vector M×1024, and the structured feature vector N×K×1024, the first training hybrid feature of size (N×16×16 + M + N×K)×1024 can be obtained.

[0198] In some alternative embodiments, the first training hybrid feature can also be obtained by adding the feature vectors of multiple modalities.

[0199] Through the above-mentioned Step 321 - Step 322, the implementation of position encoding mixing is disclosed. By generating learnable position encoding parameters during the training process, it helps to distinguish features of multiple modalities, and further helps to process multi-modal data input.

[0200] Figure 5 This is a training flow chart of the multi-modal feature conversion model in the training of the multi-modal world model of the present disclosure. As Figure 4 shown, the above-mentioned Step 303 can specifically include Step 331 - Step 334. Each step will be described below.

[0201] In Step 331, input the first training mixed feature into the initial conversion network model. The initial conversion network model masks the first part of the features in the first training mixed feature, and predicts the masked features in the way of using the mask to learn the second part of the features. The second part of the features is the features in the first training mixed feature that are not masked.

[0202] In some optional embodiments, the initial conversion network model is a Transformer model. The first training mixed feature can be input into this initial conversion network model, and this model can use the way of masked pre-training to learn the features between multiple modalities. Specifically, when implementing, a set proportion of the feature data in the first training mixed feature can be set as a mask. For example, 50% of the feature data is set to 0, and then the remaining 50% of the feature data is used to predict the masked feature data to achieve self-supervised training.

[0203] Among them, the second part of the features is the features in the first training mixed feature that are not masked.

[0204] In Step 332, based on the masked features and the first part of the features, calculate the self-supervised loss value.

[0205] In some optional embodiments, the self-supervised loss value can be calculated through various types of loss functions, such as the L1 loss function (absolute value loss function), the L2 loss function (calculate the square of the difference between each first part of the features and the predicted masked features, and then take the average), the cross-entropy loss function, etc.

[0206] In Step 333, based on the self-supervised loss value, update the model parameters of the initial conversion network model.

[0207] In some alternative embodiments, the model parameters can be updated through backpropagation of gradients. Backpropagation of gradients is an algorithm for updating model parameters by calculating the gradients of the loss function with respect to the model parameters. In a specific implementation, the loss function can be differentiated to obtain the gradients of each model parameter with respect to the loss function. These gradients can be used to represent the contribution magnitude and direction of the model parameters to the loss function in the current state, that is, the direction and magnitude of parameter updates. By adjusting the model parameters according to these gradients, the value of the loss function can be gradually reduced and the model performance can be gradually improved.

[0208] In step 334, the above-mentioned masked pre-training process is iteratively executed to implement the next round of masked pre-training until the predicted masked features satisfy the training convergence condition with respect to the first part of the features, thereby obtaining the multi-modal feature conversion model.

[0209] In some alternative embodiments, the training convergence condition is used to indicate the condition for determining the end of the training process of the multi-modal feature conversion model.

[0210] In some alternative embodiments, the training convergence condition may include that the self-supervised loss value is less than a preset threshold; in some other alternative embodiments, the training convergence condition may further include that the current training round is equal to the set number of training times of the model; in still some other alternative embodiments, the training convergence condition may further include that the current training round is not greater than the set number of training times of the model and the self-supervised loss value is less than a preset threshold.

[0211] Among them, the multi-modal feature conversion model can be used to predict the features of one modality from the features of another modality. For example, if the data input to this multi-modal world network only includes text data, then the features obtained by the decoder only include text features, and the multi-modal feature conversion model can be used to generate visual features, action features, and structured features from the text features.

[0212] In some alternative embodiments, by directly using the features after mixing the above-mentioned positional encodings to train the initial conversion network model, self-supervised training can be achieved to obtain the multi-modal feature conversion model.

[0213] Through the above steps 331 - 334, the training method of the multi-modal feature conversion model is disclosed. Through multiple rounds of self-supervised masked pre-training during the training process, it is possible to realize the feature prediction from the features of one modality of data to the features of other modalities of data, enabling the multi-modal world model to process data inputs of various modalities.

[0214] Exemplary apparatus

[0215] Figure 6Schematic structural diagram of an embodiment of a task processing apparatus based on a multi-modal world model according to the present disclosure. The apparatus in this embodiment can be used to implement the corresponding method embodiment of the present disclosure. As Figure 6 shown, the apparatus includes:

[0216] An acquisition module 61, configured to acquire multi-modal data, where the multi-modal data includes at least one of the following types of data: image data, voice data, text data, structured data, action data;

[0217] An encoding module 62, configured to use the encoders corresponding to various types of modal data in the multi-modal world model to respectively encode the various types of modal data in the multi-modal data to obtain the features of the various types of modal data;

[0218] A feature mixing module 63, configured to use the position encoder in the multi-modal world model to mix the features of the various types of modal data to obtain a first mixed feature;

[0219] A feature prediction module 64, configured to use the multi-modal feature conversion model in the multi-modal world model to predict a second mixed feature based on the first mixed feature, where the second mixed feature includes the features of the modal data not included in the multi-modal data;

[0220] A decoding module 65, configured to use the decoder corresponding to the task in the multi-modal world model to perform decoding processing on the features of the corresponding modality in the second mixed feature to obtain the task processing result of the corresponding task, where the corresponding task includes at least one of the following: video prediction task, three-dimensional generation task, perception task, action prediction task, voice prediction task.

[0221] Figure 7 Schematic structural diagram of another embodiment of a task processing apparatus based on a multi-modal world model according to the present disclosure. As Figure 7 shown, on the basis of the embodiment shown in Figure 6 In some optional embodiments, the encoding module 62 includes:

[0222] A first encoding sub-module 621, configured to, in response to the multi-modal data including image data, input the image data into the visual encoder of the multi-modal world model for encoding to obtain visual features, where the image data includes at least one of the following: image, video;

[0223] A second encoding sub-module 622, configured to, in response to the multi-modal data including voice data, input the voice data into the voice encoder of the multi-modal world model for encoding to obtain audio features;

[0224] The third encoding sub-module 623 is configured to, in response to the multimodal data including text data, input the text data into the text encoder of the multimodal world model for encoding to obtain text features;

[0225] The fourth encoding sub-module 624 is configured to, in response to the multimodal data including structured data or action data, input the structured data or action data into the structured encoder of the multimodal world model for encoding to obtain structured features or action features.

[0226] The feature mixing module 63 includes:

[0227] The parameter acquisition sub-module 631 is configured to acquire the position encoding parameters of the features of each modality data;

[0228] The feature mixing sub-module 632 is configured to sum the features of the modalities included in the multimodal data with the corresponding position encoding parameters, and sum the features of the modalities not included in the multimodal data after setting them to zero with the corresponding position encoding parameters to obtain the first mixed feature.

[0229] In some other alternative embodiments, the apparatus further includes:

[0230] The feature dimension conversion module 66 is configured to unify the features of each modality data in the multimodal data into features of the same dimension.

[0231] In some other alternative embodiments, the apparatus further includes:

[0232] The model training module 67 is configured to obtain the multimodal world model through the following steps;

[0233] The model training module 67 includes:

[0234] The encoding sub-module 671 is configured to input various modality data in the training samples into the corresponding encoders, and respectively encode the various modality data in the training samples to obtain the features of at least one modality data;

[0235] The mixing sub-module 672 is configured to perform position encoding mixing on the features of the at least one modality data by using the position encoder to obtain the first training mixed feature;

[0236] The mask learning sub-module 673 is configured to perform multi-round self-supervised mask pre-training on the initial transformation network model by using the first training mixed feature until the training convergence condition is satisfied to obtain the multimodal feature transformation model;

[0237] The merging sub-module 674 is used to merge the encoder, the multi-modal feature conversion model, the position encoder, and the decoder to form the multi-modal world model.

[0238] In some further alternative embodiments, the hybrid sub-module 672 includes:

[0239] The parameter acquisition unit 6721 is used to generate the position encoding parameters of the features of the at least one modal data based on the network gradient by using the position encoder;

[0240] The position encoding hybrid unit 6722 is used to sum the features of the modalities included in the training data and the corresponding position encoding parameters, and sum the features of the modalities not included in the training data after setting them to zero and the corresponding position encoding parameters to obtain the first training hybrid feature.

[0241] In some further alternative embodiments, the mask learning sub-module 673 includes:

[0242] The mask learning unit 6731 is used to input the first training hybrid feature into the initial conversion network model, mask the first part of the features in the first training hybrid feature by the initial conversion network model, and predict the masked features in a way of masking the second part of the features, where the second part of the features is the unmasked features in the first training hybrid feature;

[0243] The loss calculation unit 6732 is used to calculate the self-supervised loss value based on the masked features and the first part of the features;

[0244] The parameter update unit 6733 is used to update the model parameters of the initial conversion network model based on the self-supervised loss value;

[0245] The iteration unit 6734 is used to iteratively execute the above mask pre-training process to implement the next round of mask pre-training until the predicted masked features and the first part of the features meet the training convergence condition to obtain the multi-modal feature conversion model.

[0246] In some further alternative embodiments, the decoder for the corresponding task is obtained through the following operations:

[0247] Fine-tune the pre-trained video decoder with short video training samples to obtain the video decoder for the corresponding task;

[0248] Fine-tune the pre-trained 3D generation decoder with surround-view video training samples to obtain the 3D generation decoder for the corresponding task;

[0249] Fine-tune the pre-trained perception task decoder using image training samples to obtain a perception task decoder for the corresponding task;

[0250] Fine-tune the pre-trained action prediction decoder using action feature training samples to obtain an action prediction decoder for the corresponding task;

[0251] Fine-tune the pre-trained speech prediction decoder using speech feature training samples to obtain a speech prediction decoder for the corresponding task.

[0252] The device according to the embodiments of the present disclosure can be used to implement the methods of the above embodiments of the present disclosure. The specific implementations between the two correspond to each other, and the specific implementations of the relevant parts are referred to each other and will not be elaborated here.

[0253] Exemplary electronic devices, computer program products, and computer-readable storage media

[0254] The embodiments of the present disclosure further provide an electronic device, including: a memory for storing a computer program; a processor for executing the computer program stored in the memory, and when the computer program is executed, implementing the task processing method based on the multi-modal world model in any of the above embodiments of the present disclosure.

[0255] Next, with reference to Figure 8 to describe the electronic device according to the embodiments of the present disclosure, in which a device for implementing the method according to the embodiments of the present disclosure can be integrated. Figure 8 is a structural diagram of an electronic device provided by an illustrative embodiment of the present disclosure. As Figure 8 shown, the electronic device includes one or more processors 81, a memory 82 of one or more computer-readable storage media, and a computer program stored on the memory and executable on the processor. When executing the program of the memory 82, the above-mentioned task processing method based on the multi-modal world model can be implemented.

[0256] Specifically, in practical applications, the electronic device may further include components such as an input device 83 and an output device 84, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). Those skilled in the art can understand that Figure 8 the structure of the electronic device shown in

[0257] does not constitute a limitation to the electronic device, and may include more or fewer components than shown in the figure, or certain components, or different component arrangements. Among them: The processor 81 may be a central processing unit (CPU) or other forms of processing units with the ability to process tasks based on the multi-modal world model and / or instruction execution ability. By running or executing software programs and / or modules stored in the memory 82, and calling data stored in the memory 82, it executes various functions and processes data, thereby monitoring the electronic device as a whole.

[0258] The memory 82 can store one or more computer program products. The memory can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The above-mentioned volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The above-mentioned non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program products can be stored on the computer-readable storage media, and the processor 81 can run the computer program products to implement the task processing method based on the multi-modal world model of each embodiment of the present disclosure above and / or other desired functions.

[0259] The input device 83 can be used to receive input digital or character information. The input device 83 can include a keyboard, a mouse, a joystick, etc. related to user settings and function control.

[0260] The output device 84 can output various information to the outside, including the determined distance information, direction information, etc. The output device 84 can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0261] The electronic device can also include a power supply for powering each component, which can be logically connected to the processor 81 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator, etc.

[0262] Of course, for simplicity, Figure 8 only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device can also include any other appropriate components.

[0263] In addition to the above methods and devices, an embodiment of the present disclosure can also be a computer program product, which includes computer program instructions that cause the processor to execute the steps in the task processing method based on the multi-modal world model according to various embodiments of the present disclosure described in the "Exemplary Method" section above when the computer program instructions are run by the processor.

[0264] The computer program product may be written in any combination of one or more programming languages for performing the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0265] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the task processing method based on a multimodal world model according to various embodiments of the present disclosure described in the "Exemplary Method" section above of this specification.

[0266] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0267] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-disclosed specific details are only for the purposes of illustration and facilitating understanding, rather than limitations, and the above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0268] Each embodiment in this specification is described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference may be made to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and for the relevant parts, reference may be made to the partial description of the method embodiment.

[0269] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the foregoing storage medium includes various media that can store program codes, such as ROM, RAM, magnetic disks, or optical discs.

[0270] The methods and apparatuses of the present disclosure may be implemented in many ways. For example, the methods and apparatuses of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the methods is for illustrative purposes only. The steps of the methods of the present disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to the present disclosure. Thus, the present disclosure also covers a recording medium storing a program for executing the methods according to the present disclosure.

[0271] The description of the present disclosure is given for purposes of illustration and description, and is not intended to be exhaustive or to limit the present disclosure to the disclosed form. Many modifications and variations are obvious to those of ordinary skill in the art. The embodiments are chosen and described in order to better explain the principles and practical applications of the present disclosure, and to enable those of ordinary skill in the art to understand the present disclosure and design various embodiments with various modifications suitable for specific purposes.

Claims

1. A task processing method based on a multi-modal world model, characterized in that, Including: Obtain multimodal data, where the multimodal data includes data of at least one of the following modalities: image data, speech data, text data, structured data, action data; Use the encoders corresponding to various modality data in the multimodal world model to encode various modality data in the multimodal data respectively, and obtain the features of various modality data; Use the position encoder in the multimodal world model to mix the features of various modality data, and obtain a first mixed feature; Use the multimodal feature transformation model in the multimodal world model to predict a second mixed feature based on the first mixed feature, where the second mixed feature includes features of modality data not included in the multimodal data; Use the decoder corresponding to the task in the multimodal world model to decode the features of the corresponding modality in the second mixed feature, and obtain the task processing result of the corresponding task, where the corresponding task includes at least one of the following: video prediction task, 3D generation task, perception task, action prediction task, speech prediction task.

2. The method according to claim 1, characterized in that, The step of using the encoders corresponding to various modality data in the multimodal world model to encode various modality data in the multimodal data respectively, and obtain the features of the multimodal data, includes: In response to the multimodal data including image data, input the image data into the visual encoder of the multimodal world model for encoding, and obtain visual features, where the image data includes at least one of the following: image, video; In response to the multimodal data including speech data, input the speech data into the speech encoder of the multimodal world model for encoding, and obtain audio features; In response to the multimodal data including text data, input the text data into the text encoder of the multimodal world model for encoding, and obtain text features; In response to the multimodal data including structured data or action data, input the structured data or action data into the structured encoder of the multimodal world model for encoding, and obtain structured features or action features.

3. The method according to any one of claims 1-2, characterized in that, The step of using the position encoder in the multimodal world model to perform position encoding mixing on the features of the multimodal data, and obtain a first mixed feature, includes: Obtain the position encoding parameters of the features of each modality data; Sum the features of the modalities included in the multimodal data with the corresponding position encoding parameters, and sum the features of the modalities not included in the multimodal data after setting them to zero with the corresponding position encoding parameters, to obtain the first mixed feature.

4. The method according to any one of claims 1-3, characterized in that Before using the position encoder in the multimodal world model to perform position encoding mixing on the features of the multimodal data, it further includes: Unify the features of each modality data in the multimodal data into features of the same dimension.

5. The method according to any one of claims 1-4, characterized in that The multimodal world model is trained through the following steps: Input various modality data in the training samples into the corresponding encoders, and encode various modality data in the training samples respectively, to obtain the features of at least one modality data; Perform positional encoding mixing on the features of the at least one type of modality data using the positional encoder to obtain a first training mixed feature; Perform multiple rounds of self-supervised masked pre-training on the initial transformation network model using the first training mixed feature until the training convergence condition is met to obtain the multi-modal feature transformation model; Combine the encoder, the multi-modal feature transformation model, the positional encoder, and the decoder to form the multi-modal world model.

6. The method according to claim 5, wherein The step of performing positional encoding mixing on the features of the at least one type of modality data using the positional encoder to obtain a first training mixed feature includes: Generate positional encoding parameters for the features of the at least one type of modality data using the positional encoder based on the network gradient; Sum the features of the modalities included in the training data with the corresponding positional encoding parameters, and sum the features of the modalities not included in the training data after setting them to zero with the corresponding positional encoding parameters to obtain the first training mixed feature.

7. The method according to claim 5, characterized in that, The step of performing multiple rounds of self-supervised masked pre-training on the initial transformation network model using the first training mixed feature until the training convergence condition is met to obtain the multi-modal feature transformation model includes: Input the first training mixed feature into the initial transformation network model. The initial transformation network model masks a first part of the features in the first training mixed feature and predicts the masked features in a way of masking and learning a second part of the features, where the second part of the features is the unmasked features in the first training mixed feature; Calculate the self-supervised loss value based on the masked features and the first part of the features; Update the model parameters of the initial transformation network model based on the self-supervised loss value; Iteratively execute the above masked pre-training process to implement the next round of masked pre-training until the predicted masked features and the first part of the features meet the training convergence condition to obtain the multi-modal feature transformation model.

8. The method according to any one of claims 1-7, characterized in that, The decoder for the corresponding task is obtained through the following operations: Fine-tune the pre-trained video decoder using short video training samples to obtain the video decoder for the corresponding task; Fine-tune the pre-trained 3D generation decoder using surround-view video training samples to obtain the 3D generation decoder for the corresponding task; Fine-tune the pre-trained perception task decoder using image training samples to obtain the perception task decoder for the corresponding task; Fine-tune the pre-trained action prediction decoder using action feature training samples to obtain the action prediction decoder for the corresponding task; Fine-tune the pre-trained speech prediction decoder using speech feature training samples to obtain the speech prediction decoder for the corresponding task.

9. A task processing device based on a multi-modal world model, characterized in that, It includes: An acquisition module for acquiring multi-modal data, where the multi-modal data includes at least one of the following types of data: image data, speech data, text data, structured data, action data; An encoding module for encoding each type of modality data in the multi-modal data respectively using the encoders corresponding to the various modality data in the multi-modal world model to obtain the features of the various modality data; A feature mixing module for using the position encoder in the multimodal world model to mix the features of the various modal data to obtain a first mixed feature; A feature prediction module for using the multimodal feature transformation model in the multimodal world model to predict a second mixed feature based on the first mixed feature, where the second mixed feature includes the features of the modal data not included in the multimodal data; A decoding module for using the decoder corresponding to the task in the multimodal world model to perform decoding processing on the features of the corresponding modality in the second mixed feature to obtain the task processing result of the corresponding task, where the corresponding task includes at least one of the following: video prediction task, three-dimensional generation task, perception task, action prediction task, speech prediction task.

10. The device according to claim 9, characterized in that The encoding module includes: A first encoding sub-module for, in response to the multimodal data including image data, inputting the image data into the visual encoder of the multimodal world model for encoding to obtain visual features, where the image data includes at least one of the following: image, video; A second encoding sub-module for, in response to the multimodal data including speech data, inputting the speech data into the speech encoder of the multimodal world model for encoding to obtain audio features; A third encoding sub-module for, in response to the multimodal data including text data, inputting the text data into the text encoder of the multimodal world model for encoding to obtain text features; A fourth encoding sub-module for, in response to the multimodal data including structured data or action data, inputting the structured data or action data into the structured encoder of the multimodal world model for encoding to obtain structured features or action features.

11. The device according to any one of claims 9-10, characterized in that, The feature mixing module includes: A parameter acquisition sub-module for acquiring the position encoding parameters of the features of each modal data; A feature mixing sub-module for summing the features of the modalities included in the multimodal data with the corresponding position encoding parameters, and summing the features of the modalities not included in the multimodal data after setting them to zero with the corresponding position encoding parameters to obtain the first mixed feature.

12. The device according to any one of claims 9-11, characterized in that, The device further includes: A feature dimension conversion module for unifying the features of each modal data in the multimodal data into features of the same dimension.

13. The device according to any one of claims 9-12, characterized in that, The device further includes: A model training module for obtaining the multimodal world model through the following steps; The model training module includes: An encoding sub-module for inputting the various modal data in the training samples into the corresponding encoders to respectively encode the various modal data in the training samples to obtain the features of at least one modal data; A mixing sub-module for using the position encoder to perform position encoding mixing on the features of the at least one modal data to obtain a first training mixed feature; A mask learning sub-module for using the first training mixed feature to perform multi-round self-supervised mask pre-training on the initial transformation network model until the training convergence condition is met to obtain the multimodal feature transformation model; A merging sub-module, configured to merge the encoder, the multi-modal feature conversion model, the position encoder, and the decoder to form the multi-modal world model.

14. The device according to claim 13, wherein The hybrid sub-module includes: A parameter acquisition unit, configured to generate position encoding parameters of features of the at least one modal data by using the position encoder based on network gradients. A position encoding hybrid unit, configured to sum the features of the modalities included in the training data and the corresponding position encoding parameters, and sum the features of the modalities not included in the training data after setting them to zero and the corresponding position encoding parameters, to obtain the first training hybrid feature.

15. The device according to claim 13, characterized in that, The mask learning sub-module includes: A mask learning unit, configured to input the first training hybrid feature into the initial conversion network model, mask a first part of the features in the first training hybrid feature by the initial conversion network model, and predict the masked features in a way of masking and learning a second part of the features, where the second part of the features is the features in the first training hybrid feature that are not masked. A loss calculation unit, configured to calculate a self-supervised loss value based on the masked features and the first part of the features. A parameter update unit, configured to update the model parameters of the initial conversion network model based on the self-supervised loss value. An iteration unit, configured to iteratively execute the above mask pre-training process to implement the next round of mask pre-training until the predicted masked features and the first part of the features meet the training convergence condition, to obtain the multi-modal feature conversion model.

16. The device according to any one of claims 9-15, characterized in that, The decoder for the corresponding task is obtained through the following operations: Fine-tuning the pre-trained video decoder with short video training samples to obtain the video decoder for the corresponding task. Fine-tuning the pre-trained 3D generation decoder with surround-view video training samples to obtain the 3D generation decoder for the corresponding task. Fine-tuning the pre-trained perception task decoder with image training samples to obtain the perception task decoder for the corresponding task. Fine-tuning the pre-trained action prediction decoder with action feature training samples to obtain the action prediction decoder for the corresponding task. Fine-tuning the pre-trained speech prediction decoder with speech feature training samples to obtain the speech prediction decoder for the corresponding task.

17. A computer-readable storage medium storing computer program instructions, which when executed, implement the method according to any one of claims 1-8 above.

18. An electronic device, comprising: A memory, configured to store a computer program product; A processor, configured to execute the computer program product stored in the memory, and when the computer program product is executed, implement the method according to any one of claims 1-8 above.

19. A computer program product, comprising computer program instructions, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1-8 above is implemented.

Citation Information

Patent Citations

  • Training method, using method, device and equipment of multi-modal pre-training model

    CN116756574A

Cited By

  • Visual language action model training method and device, equipment and storage medium

    CN121259341A

  • Training method of visual language action model and mechanical arm operating device

    CN121267892A

  • Method for training visual language action model and mechanical arm operating device

    CN121267892B