Multi-modal automatic driving scene generation method based on autoregression closed-loop prediction
By adopting a multimodal autonomous driving scenario generation method based on autoregressive closed-loop prediction in the field of autonomous driving, integrating multimodal data and training a generative model, the shortcomings in the universality and efficiency of the existing world models are solved, and multitasking prediction and efficient closed-loop scenario prediction are achieved.
Patent Information
- Application Number
- CN202510098839.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-10
AI Technical Summary
The existing world models are inequality and efficiency in the field of autonomous driving, cannot be used directly for path planning or control decisions, and rely on high-precision occupancy representations, are inefficient and lack of advanced human-computer interaction.
A multimodal autonomous driving scenario generation method based on autoregressive closed-loop prediction is adopted. By obtaining multimodal multitasking autonomous driving data sets, integrating them into a unified discrete space, multimodal discrete encoding is generated, and a generation model is trained through the autoregressive paradigm, and fine-tuning is performed on scene understanding, scene prediction and trajectory planning tasks in turn to achieve closed-loop scenario prediction.
It significantly improves the universality and efficiency of the world model of autonomous driving, realizes multi-task prediction capabilities, including scene image generation, scene understanding Q&A and path planning, solves the problems of single and inefficient tasks of the existing model, and reduces costs, enhancing the model's adaptability to complex driving scenarios.
Smart Images

Figure CN120124737A_ABST
Abstract
Description
Technical Field
[0001] This application relates to autonomous driving and machine learning technologies, and particularly to a multi-modal autonomous driving scenario generation method based on autoregressive closed-loop prediction. Background Art
[0002] The proposed world model provides a new perspective for reinforcement learning and decision-making models. Especially when facing high-dimensional and dynamically changing environments, it can significantly improve performance by predicting environmental changes and the results of agent behavior. In the field of autonomous driving, the world model helps the system understand the decision-making process, enhance transparency and credibility by predicting future environmental states. At the same time, the world model effectively alleviates the problem of data scarcity in autonomous driving. By generating diverse driving data in a simulated environment, it reduces the dependence on manual annotation, thereby optimizing the perception, planning, and control modules and improving the generalization ability.
[0003] Existing research has applied the world model to driving video generation. For example, some scholars have designed a method combining a latent variable replacement mechanism and a dynamic loss enhancement strategy to capture the temporal changes of the environment and generate natural and consistent driving video sequences; other research has achieved cross-perspective and cross-frame modeling through a stable diffusion framework and an attention mechanism to generate high-quality multi-perspective driving videos. In addition, the world model has also been applied to the prediction of radar point clouds and 3D occupancy information, improving the interpretability of autonomous driving systems.
[0004] However, the generality and efficiency of existing world models are still insufficient. Most methods are only used as auxiliary tools for data generation or reinforcement learning reward functions and cannot be directly used for path planning or control decisions. For example, although some models can predict 3D semantic occupancy information and future safe trajectories, their applicable range is limited and they rely on high-precision occupancy representations. In addition, existing solutions mostly perform predictions based on 3D intermediate representations, with low efficiency and a lack of high-order human-machine interaction.
[0005] In recent years, the development of autoregressive multi-modal generation models has provided new ideas for solving the above problems. Some research has significantly improved the generation efficiency and multi-modal fusion ability compared with traditional methods by mapping images and text to a unified discrete space and combining a causal autoregressive mechanism to generate multi-modal information.
[0006] However, how to further improve the generality, efficiency, and deployment convenience of the world model remains a technical problem to be solved urgently. Summary of the Invention
[0007] This application aims to solve at least one of the technical problems in the related technologies to some extent.
[0008] To this end, the first objective of this application is to propose a multi-modal autonomous driving scenario generation method based on autoregressive closed-loop prediction.
[0009] The second object of the present application is to propose a multi-modal autonomous driving scenario generation device based on autoregressive closed-loop prediction.
[0010] The third object of the present application is to propose an electronic device.
[0011] The fourth object of the present application is to propose a computer-readable storage medium.
[0012] The fifth object of the present application is to propose a computer program product.
[0013] To achieve the above object, the first aspect embodiment of the present application proposes a multi-modal autonomous driving scenario generation method based on autoregressive closed-loop prediction, including:
[0014] Obtain a multi-modal multi-task autonomous driving data set, and integrate and map it to a unified discrete space through different discrete encoders to generate an integrated multi-modal discrete encoding;
[0015] Integrate the multi-modal discrete encodings into a unified encoding sequence, adopt a masking strategy to cover the trajectory encoding, use the original encoding after removing the mask as the supervision target, train a generation model through an autoregressive paradigm, and fine-tune the model on the scene understanding, scene prediction, and trajectory planning task data sets in turn;
[0016] Take the initial scene image collected by the vehicle camera as the input, combine with the prompt words designed by the user, and according to the trained model, output a future scene discrete encoding sequence. By replacing the trajectory information in the discrete encoding sequence with a mask encoding and re-inputting it into the model, closed-loop scene prediction is completed;
[0017] Use a decoder symmetric to the corresponding modal encoder to decode the generated discrete encoding into the image, text Q&A, and vehicle trajectory information of the future scene.
[0018] Optionally, the obtaining a multi-modal multi-task autonomous driving data set, and integrating and mapping it to a unified discrete space through different discrete encoders to generate an integrated multi-modal discrete encoding includes:
[0019] Collect and integrate a multi-modal autonomous driving data set including autonomous driving scene videos, scene understanding prediction Q&A, and trajectory planning data;
[0020] Use a pre-trained VQGAN, BPE tokenizer, and positional embedding encoder to encode the autonomous driving scene video, scene understanding prediction Q&A, and trajectory planning data respectively, map the image modality, text modality, and trajectory modality data to a unified discrete codebook space, and generate an integrated multi-modal discrete encoding as the subsequent training data.
[0021] Optionally, to integrate the multimodal discrete encodings into a unified encoding sequence, a masking strategy is adopted to cover the trajectory encoding, and the original encoding after removing the mask is used as the supervision target. The generative model is trained through an autoregressive paradigm, and the model is fine-tuned on the task datasets of scene understanding, scene prediction, and trajectory planning in sequence, including:
[0022] Integrate the discrete encodings of the image modality, the text modality, and the trajectory modality in the training data into a unified encoding sequence. Among them, a masking strategy is designed for the trajectory encoding part to cover the target trajectory information;
[0023] Use the original encoding after removing the mask as the training supervision target, and train the generative model through an autoregressive paradigm to maximize the probability corresponding to the target encoding in the probability distribution output by the model;
[0024] Conduct basic training on the model, and based on the data sampled from a large-scale driving video dataset, learn the basic ability of the text-to-image generation of the training model;
[0025] After the model converges initially, fine-tune the model on the task datasets of scene understanding, scene prediction, and trajectory planning in sequence, and set weights for the loss function of the trajectory planning part during supervision;
[0026] Integrate the multimodal dataset containing scene understanding, scene prediction, and trajectory planning tasks into a unified dataset, and conduct complete fine-tuning on the model.
[0027] Optionally, taking the initial scene image collected by the vehicle camera as the input, combining with the prompt words designed by the user, according to the trained model to output the future scene discrete encoding sequence, and complete the closed-loop scene prediction by replacing the trajectory information in the discrete encoding sequence with the mask encoding and re-inputting it into the model, including:
[0028] Obtain the initial scene forward-view image through the vehicle camera, and encode the image into the unified discrete codebook space to obtain the initial scene information;
[0029] Input the prompt words designed by the user according to the task requirements, and input them into the trained model together with the initial scene information to generate the probability distribution of the future encoding;
[0030] Perform specific sampling on the probability distribution to generate a future scene discrete encoding sequence containing future scene image encoding, scene information understanding encoding, and vehicle trajectory encoding;
[0031] Replace the trajectory encoding in the future scene discrete encoding sequence with the mask encoding, and integrate the replaced encoding sequence into the original input;
[0032] Re-input the integrated encoded sequence into the model to generate a discrete encoded sequence of a farther future scene without providing the real camera future frame information of the vehicle, thus realizing closed-loop scene prediction.
[0033] To achieve the above object, an embodiment of the second aspect of the present application provides a multi-modal autonomous driving scene generation device based on autoregressive closed-loop prediction, including:
[0034] An encoding module, configured to obtain a multi-modal multi-task autonomous driving data set, and integrate and map it to a unified discrete space through different discrete encoders to generate an integrated multi-modal discrete encoding;
[0035] A training module, configured to integrate the multi-modal discrete encoding into a unified encoding sequence, adopt a masking strategy to cover the trajectory encoding, use the original encoding after removing the mask as a supervision target, train a generation model through an autoregressive paradigm, and fine-tune the model on the data sets of scene understanding, scene prediction, and trajectory planning tasks in sequence;
[0036] An inference module, configured to use the initial scene image collected by the vehicle camera as an input, combine with the prompt words designed by the user, output a future scene discrete encoding sequence according to the trained model, and complete closed-loop scene prediction by replacing the trajectory information in the discrete encoding sequence with a mask encoding and re-inputting it into the model;
[0037] A decoding module, configured to use a decoder symmetric to the corresponding modal encoder to decode the generated discrete encoding into the image, text Q&A, and vehicle trajectory information of the future scene.
[0038] Optionally, the encoding module is configured to:
[0039] Collect and integrate a multi-modal autonomous driving data set including autonomous driving scene videos, scene understanding prediction Q&A, and trajectory planning data;
[0040] Use the pre-trained VQGAN, BPE tokenizer, and positional embedding encoder to encode the autonomous driving scene video, scene understanding prediction Q&A, and trajectory planning data respectively, map the image modality, text modality, and trajectory modality data to a unified discrete codebook space, and generate an integrated multi-modal discrete encoding as subsequent training data.
[0041] Optionally, the training module is configured to:
[0042] Integrate the discrete encodings of the image modality, the text modality, and the trajectory modality in the training data into a unified encoding sequence, where a masking strategy is designed for the trajectory encoding part to cover the target trajectory information;
[0043] Use the original encoded data after removing the mask as the training supervision target, train the generative model through an autoregressive paradigm, and maximize the probability corresponding to the target encoding in the probability distribution of the model output;
[0044] Conduct basic training on the model, and based on the data sampled from a large-scale driving video dataset, learn the basic text-to-image generation ability of the training model;
[0045] After the model is initially converged, fine-tune the model successively on the scene understanding task dataset, the scene prediction task dataset, and the trajectory planning task dataset, and set weights for the loss function of the trajectory planning part during supervision;
[0046] Integrate the multi-modal dataset including scene understanding, scene prediction, and trajectory planning tasks into a unified dataset, and conduct complete fine-tuning on the model.
[0047] Optionally, the inference module is used for:
[0048] Obtain the initial scene forward-view image through the vehicle camera, and encode the image into a unified discrete codebook space to obtain the initial scene information;
[0049] According to the task requirements, input the prompt words designed by the user, and input them together with the initial scene information into the trained model to generate the probability distribution of future encodings;
[0050] Perform specific sampling on the probability distribution to generate a future scene discrete encoding sequence including future scene image encoding, scene information understanding encoding, and vehicle trajectory encoding;
[0051] Replace the trajectory encoding in the future scene discrete encoding sequence with a mask encoding, and integrate the replaced encoding sequence into the original input;
[0052] Re-input the integrated encoding sequence into the model, and generate a further future scene discrete encoding sequence in the case of not providing the future frame information of the vehicle's real camera to achieve closed-loop scene prediction.
[0053] To achieve the above object, an embodiment of the third aspect of the present application proposes an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0054] The memory stores computer execution instructions;
[0055] The processor executes the computer execution instructions stored in the memory to implement the method described in any one of the first aspect.
[0056] To achieve the above object, an embodiment of the fourth aspect of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in any one of the first aspect.
[0057] To achieve the above object, an embodiment of the fifth aspect of the present application provides a computer program product, and when the computer program is executed by a processor, it implements the method described in any one of the first aspect.
[0058] The technical solutions provided by the embodiments of the present application at least bring the following beneficial effects:
[0059] Through the combination of multi-modal discrete coding and the autoregressive mechanism, the present application significantly improves the generality and efficiency of the autonomous driving world model; encodes historical scene images, text descriptions, and driving control parameters using a unified discrete space, and through a multi-modal generation model with staged fine-tuning, realizes multi-task prediction capabilities, including tasks such as scene picture generation, scene understanding Q&A, and path planning, solving the problems of single task and low efficiency of existing models; in addition, the present application uses probability sampling to generate future scene encodings and integrates the prediction information into the input to form a fully closed-loop driving scene simulation, reducing costs and enhancing the model's adaptability to complex driving scenes.
[0060] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0062] Figure 1 is a schematic flowchart of a multi-modal autonomous driving scene generation method based on autoregressive closed-loop prediction provided by an embodiment of the present application;
[0063] Figure 2 is a schematic diagram of a model of a multi-modal autonomous driving scene generation method based on autoregressive closed-loop prediction provided by an embodiment of the present application;
[0064] Figure 3 is a schematic structural diagram of a multi-modal autonomous driving scene generation device based on autoregressive closed-loop prediction provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, but should not be construed as limiting the present application.
[0066] In view of the technical problems existing in the prior art, an embodiment of the present application provides a multi-modal autonomous driving scenario generation method based on autoregressive closed-loop prediction. Figure 1 and Figure 2 are respectively a flowchart and a model diagram of a multi-modal autonomous driving scenario generation method provided by an embodiment of the present application. Referring to Figure 1 and Figure 2 , the method includes the following steps:
[0067] Step 101, obtain a multi-modal multi-task autonomous driving data set, and integrate and map it to a unified discrete space through different discrete encoders to generate an integrated multi-modal discrete encoding.
[0068] To avoid information loss and inductive bias that may be caused by artificially designed intermediate scene representations, an embodiment of the present application realizes the efficient integration and unified representation of multi-modal data by constructing a unified discrete codebook space and combining the encoding methods of multi-modal data.
[0069] Specifically, the embodiment of the present application includes the following steps:
[0070] First, collect and integrate a multi-modal autonomous driving data set including scene videos, scene understanding prediction Q&A, and trajectory planning data. These data are derived from historical information in autonomous driving scenarios and can provide comprehensive input support for subsequent model training.
[0071] To achieve a standardized representation of different modal data, the embodiment of the present application designs corresponding encoding methods for each modal data:
[0072] (1) Processing of the image modality: To efficiently extract visual and spatial feature information in scene videos, the embodiment of the present application uses a pre-trained VQGAN model to encode video data. Through this model, the original video frames are converted into discrete and efficient feature representations, ensuring that the data of the image modality can be represented in the same space as other modalities.
[0073] (2) Processing of the text modality: For scene understanding prediction Q&A data, the embodiment of the present application uses a BPE tokenizer to encode text data. By discretizing the semantic information in the text and mapping it to a unified discrete codebook space, a standardized representation of the text modality is achieved, thereby ensuring that the model can accurately understand and process text inputs.
[0074] (3) Processing of the trajectory modality: To address the problem of high processing complexity for continuous floating-point number sequences in trajectory planning data, the embodiments of the present application use a position embedding encoder to encode the trajectory data. This method can convert the continuous information of the trajectory into a discrete trajectory feature representation while retaining the spatial and temporal information in the trajectory, laying a foundation for subsequent prediction tasks.
[0075] Through the above encoding process, the embodiments of the present application successfully integrate the data of the image modality, text modality, and trajectory modality into a unified discrete codebook space, generating an integrated multi-modal discrete encoding. This unified representation form not only avoids the feature expression differences between different modalities but also provides a standardized input with strong adaptability and high expressive ability for subsequent autoregressive model training.
[0076] Step 102: Integrate the multi-modal discrete encoding into a unified encoding sequence, adopt a masking strategy to cover the trajectory encoding, use the original encoding after removing the mask as the supervision target, and train the generation model through an autoregressive paradigm. Then, fine-tune the model on the task datasets of scene understanding, scene prediction, and trajectory planning in sequence.
[0077] To make full use of the information in the multi-modal discrete encoding, the embodiments of the present application integrate the multi-modal encoding into a unified encoding sequence and adopt a masking strategy to process the trajectory encoding part, thereby achieving efficient training and fine-tuning of the model.
[0078] Specifically, it includes the following steps:
[0079] First, the embodiments of the present application integrate the discrete encodings of the image modality, text modality, and trajectory modality in the training data according to the time series to generate a unified encoding sequence. During this process, to enable the model to learn the trajectory prediction task, a masking strategy is designed for the trajectory encoding part in the unified encoding sequence to cover the target trajectory information, and the original encoding after removing the mask is used as the training supervision target to guide the training of the model.
[0080] To achieve efficient training, the embodiments of the present application train the generation model through an autoregressive paradigm and optimize the model by maximizing the probability of the target encoding in the probability distribution of the model output. In this way, the model can gradually learn the associations between multi-modal tasks and adapt to the requirements of trajectory prediction.
[0081] To ensure the generality and multi-task performance of the model, in the embodiments of the present application, the model is first trained based on the data sampled from a large-scale driving video dataset to learn its basic ability of generating images from text. After the model training converges initially, the embodiments of the present application further fine-tune the model on the scene understanding task dataset, the scene prediction task dataset, and the trajectory planning task dataset respectively to enhance the performance ability of the model on each task.
[0082] In addition, to improve the accuracy of trajectory prediction, in the supervision process, the embodiments of the present application set weights for the loss function of the trajectory planning task part to strengthen the training effect of the trajectory part.
[0083] Finally, the embodiments of the present application integrate the datasets of the scene understanding, scene prediction, and trajectory planning tasks into a unified dataset, and perform complete fine-tuning on the model, thereby further improving the generality and performance ability of the model among multiple tasks.
[0084] Through the above steps, the embodiments of the present application achieve the effective integration and processing of multi-modal data. The generated model obtained by training has higher prediction accuracy and multi-task adaptation ability, laying a solid foundation for the subsequent inference stage.
[0085] Step 103: Take the initial scene image collected by the vehicle camera as the input, combine with the prompt words designed by the user, and according to the trained model, output the future scene discrete coding sequence. By replacing the trajectory information in the discrete coding sequence with a mask coding and re-inputting it into the model, the closed-loop scene prediction is completed.
[0086] To achieve efficient prediction and closed-loop simulation of the autonomous driving scene, the embodiments of the present application are based on the initial scene information collected by the vehicle camera and the prompt words designed by the user, combine with the trained multi-modal autoregressive generation model, gradually generate the discrete coding sequence of the future scene, and complete the closed-loop scene prediction through the mask strategy of the trajectory information.
[0087] Specifically, it includes the following steps:
[0088] First, obtain the forward-view image of the vehicle's current environment through the vehicle camera. This image contains the spatial structure and visual feature information of the scene, providing the basic data for subsequent scene prediction. The embodiments of the present application further encode the obtained forward-view image into a unified discrete codebook space to obtain the initial scene information. Through this unified discrete representation, the information loss caused by modal differences is avoided, providing a standardized input form for multi-modal fusion.
[0089] After obtaining the initial scenario information, the embodiments of the present application input the prompt words designed by the user according to the task requirements. The prompt words may include specific scenario target descriptions, path planning requirements, or restrictions on environmental conditions, etc., aiming to clarify the prediction task of the model. Subsequently, the prompt words and the initial scenario information are input into the trained multi-modal autoregressive generation model together. Based on the input information, the model generates the discrete coding probability distribution of the future scenario. This probability distribution describes various possibilities of image features, scenario understanding information, and vehicle trajectories in the future scenario.
[0090] Next, the embodiments of the present application perform specific sampling on the generated discrete coding probability distribution of the future scenario. The purpose of sampling is to select the coding information that best conforms to the prompt words and the initial scenario conditions from the possible scenario codings, so as to generate the discrete coding sequence of the future scenario. The coding sequence includes the image coding, scenario understanding coding, and vehicle trajectory coding of the future scenario, and can comprehensively represent the visual information, semantic information, and vehicle driving state of the future scenario.
[0091] To achieve closed-loop scenario prediction, the embodiments of the present application also perform masking processing on the trajectory coding part in the discrete coding sequence of the future scenario. Specifically, the future trajectory information is replaced with a mask coding, so as to hide the true future trajectory state. The masked coding sequence is integrated into the original input as the input condition for the next round of model inference. This method not only eliminates the dependence on future real sensor data, but also ensures the continuity and consistency of closed-loop prediction.
[0092] Finally, the integrated coding sequence is re-input into the trained generation model. In the case of not providing the future frame information of the vehicle's real camera, the model can generate the discrete coding sequence of the farther future scenario and complete the multi-step iterative prediction of the future scenario. Through this closed-loop prediction mechanism, the embodiments of the present application achieve an efficient simulation of the dynamic driving scenario. In each round of prediction, the model can adjust the output result according to the updated input conditions, thereby enhancing the dynamic adaptation ability to complex driving environments.
[0093] The embodiments of the present application have successfully constructed a closed-loop autonomous driving scenario prediction system with both generality and efficiency through the above method. Compared with the existing prediction methods that rely on real sensor data, this method not only reduces the data acquisition cost, but also enhances the prediction ability of the model, providing strong support for the application of autonomous driving systems in complex scenarios.
[0094] Step 104: Use a decoder symmetric to the corresponding modal encoder to decode the generated discrete coding into the image, text Q&A, and vehicle trajectory information of the future scenario.
[0095] In the embodiments of the present application, through the decoding process, the discrete encoding of the future scenarios generated by the model is converted into visual or actionable output information, including future scenario images, scenario understanding Q&A, and vehicle trajectory information, so as to comprehensively present the multi-modal scenario prediction results.
[0096] Specifically, after the model completes the generation of the discrete encoding of the future scenarios, the embodiments of the present application design and use a symmetric decoder according to the encoders corresponding to each modality. The role of the decoder is to restore the discretized multi-modal encoding to the original data form while maintaining the integrity of the information and the accuracy of the expression.
[0097] (1) For the encoding of future scenario images, the embodiments of the present application use a decoder symmetric to the pre-trained VQGAN encoder to restore the discrete image encoding to a high-resolution image output. The decoding process can restore the spatial layout and visual features in the scenario, thereby generating a future scenario image with strong realism and providing visual support for the perception module of the autonomous driving system.
[0098] (2) For the text encoding of scenario understanding Q&A, the embodiments of the present application adopt a decoder symmetric to the BPE tokenizer to restore the discrete text encoding to a natural language description or Q&A form. Through decoding, the system can output the semantic understanding of the future scenario, such as a description of the predicted environmental changes or a verbalized explanation of specific task objectives.
[0099] (3) For the encoding of vehicle trajectories, the embodiments of the present application use a decoder symmetric to the position embedding encoder to restore the discrete trajectory encoding to the future trajectory planning information of the vehicle. The decoded trajectory data is presented in the form of a continuous floating-point number sequence, which can accurately describe the future driving path of the vehicle and provide support for the driving planning and control module.
[0100] In the embodiments of the present application, through the above decoding process, the generated multi-modal discrete encoding is converted into scenario information in the form of original data. This method not only realizes the multi-dimensional and multi-modal expression of future scenarios, but also provides a comprehensive solution for perception, semantic understanding, and trajectory planning of the autonomous driving system.
[0101] Through the decoder symmetric to the encoder, the embodiments of the present application can ensure the accuracy and consistency of the decoding results, so that the generated future scenarios highly match the actual driving requirements in terms of visual, semantic, and motion features, further enhancing the adaptability and practicality of the autonomous driving system.
[0102] To implement the above embodiments, the present application also proposes a multi-modal autonomous driving scenario generation device based on autoregressive closed-loop prediction. Figure 3 This is a schematic structural diagram of a multi-modal autonomous driving scenario generation device 10 provided for the embodiments of the present application. AsFigure 3 As shown in the figure, the device includes:
[0103] An encoding module 100, configured to obtain a multi-modal multi-task autonomous driving data set, and integrate and map it to a unified discrete space through different discrete encoders to generate an integrated multi-modal discrete encoding;
[0104] A training module 200, configured to integrate the multi-modal discrete encodings into a unified encoding sequence, adopt a masking strategy to cover the trajectory encoding, use the original encoding after removing the mask as the supervision target, train a generation model through an autoregressive paradigm, and fine-tune the model on the data sets of scene understanding, scene prediction, and trajectory planning tasks in sequence;
[0105] An inference module 300, configured to use the initial scene image collected by the vehicle camera as input, combine the prompt words designed by the user, output a future scene discrete encoding sequence according to the trained model, and complete the closed-loop scene prediction by replacing the trajectory information in the discrete encoding sequence with a mask encoding and re-inputting it into the model;
[0106] A decoding module 400, configured to use a decoder symmetric to the corresponding modal encoder to decode the generated discrete encoding into the image, text Q&A, and vehicle trajectory information of the future scene.
[0107] Optionally, the encoding module 100 is configured to:
[0108] Collect and integrate a multi-modal autonomous driving data set including autonomous driving scene videos, scene understanding prediction Q&A, and trajectory planning data;
[0109] Use the pre-trained VQGAN, BPE tokenizer, and positional embedding encoder to encode the autonomous driving scene video, scene understanding prediction Q&A, and trajectory planning data respectively, map the image modality, text modality, and trajectory modality data to a unified discrete codebook space, and generate an integrated multi-modal discrete encoding as the subsequent training data.
[0110] Optionally, the training module 200 is configured to:
[0111] Integrate the discrete encodings of the image modality, the text modality, and the trajectory modality in the training data into a unified encoding sequence, where a masking strategy is designed for the trajectory encoding part to cover the target trajectory information;
[0112] Use the original encoding after removing the mask as the training supervision target, train a generation model through an autoregressive paradigm, and maximize the probability corresponding to the target encoding in the probability distribution output by the model;
[0113] Conduct basic training on the model, and learn the text-to-image basic ability of the training model based on the data sampled from a large-scale driving video data set;
[0114] After the model is initially converged, the model is fine-tuned on the scene understanding task dataset, the scene prediction task dataset, and the trajectory planning task dataset in sequence, and the loss function of the trajectory planning part is weighted during supervision;
[0115] Integrate the multi-modal dataset including scene understanding, scene prediction, and trajectory planning tasks into a unified dataset, and perform complete fine-tuning on the model.
[0116] Optionally, the inference module 300 is used for:
[0117] Obtain the initial scene forward view image through the vehicle camera, and encode the image into the unified discrete codebook space to obtain the initial scene information;
[0118] According to the task requirements, input the prompt words designed by the user, and input them into the trained model together with the initial scene information to generate the probability distribution of the future encoding;
[0119] Perform specific sampling on the probability distribution to generate a future scene discrete encoding sequence including future scene image encoding, scene information understanding encoding, and vehicle trajectory encoding;
[0120] Replace the trajectory encoding in the future scene discrete encoding sequence with a mask encoding, and integrate the replaced encoding sequence into the original input;
[0121] Re-input the integrated encoding sequence into the model, and generate a scene discrete encoding sequence in the farther future without providing the future frame information of the vehicle real camera, so as to realize closed-loop scene prediction.
[0122] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0123] To implement the above embodiments, the present application also provides an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.
[0124] To implement the above embodiments, the present application also provides a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the method provided in the foregoing embodiments.
[0125] To implement the above embodiments, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the method provided in the foregoing embodiments.
[0126] The collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in this application all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0127] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of these legal uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization including authorizing relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others with access to the personal information data comply with their privacy policies and procedures.
[0128] The present application anticipates providing embodiments for users to selectively block the use or access of personal information data. That is, the present disclosure anticipates providing hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of the user.
[0129] In the description of the foregoing embodiments, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0130] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0131] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing custom logical functions or processes, and the scope of the preferred embodiments of the present application includes additional implementations where functions may be executed not in the order shown or discussed, including in a substantially simultaneous manner or in the reverse order according to the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0132] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing a logical function and can be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection having one or more wires (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or otherwise appropriate processing if necessary, and then storing it in a computer memory.
[0133] It should be understood that the various parts of the present application can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0134] Those of ordinary skill in the art can understand that all or part of the steps carried out in implementing the above-described embodiment methods can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0135] In addition, in each of the embodiments of the present application, the functional units can be integrated into one processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0136] The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
[0137] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present application can be achieved. This is not limited herein.
[0138] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. A multimodal autonomous driving scene generation method based on autoregressive closed-loop prediction, characterized in that: The following steps are involved: Obtain a multimodal and multitask autonomous driving dataset, integrate and map it into a unified discrete space through different discrete encoders, and generate an integrated multimodal discrete encoding; The multimodal discrete codes are integrated into a unified code sequence, and the trajectory code is covered by a masking strategy, and the original code without the mask is used as the supervision target. The model is generated by training the autoregressive paradigm, and the model is fine-tuned on the scene understanding, scene prediction and trajectory planning task datasets in turn; The initial scene image captured by the vehicle camera is used as input, combined with the prompt words designed by the user, and the future scene discrete code sequence is output according to the trained model. The trajectory information in the discrete code sequence is replaced with the mask code and re-input into the model to complete the closed-loop scene prediction; Using a decoder that is symmetric to the corresponding modality encoder, the generated discrete codes are decoded into images of future scenes, text questions and answers, and vehicle trajectory information.
2. The method according to claim 1, characterized in that The method of obtaining a multimodal multitask autonomous driving dataset and integrating and mapping it into a unified discrete space through different discrete encoders to generate an integrated multimodal discrete encoding includes: Collect and integrate multimodal autonomous driving datasets including autonomous driving scene videos, scene understanding prediction question answering, and trajectory planning data; Use pre-trained VQGAN, BPE word segmenter and position embedding encoder to encode autonomous driving scene videos, scene understanding prediction question and answer, and trajectory planning data respectively, map the image modality, text modality and trajectory modality data to a unified discrete codebook space, and generate integrated multimodal discrete coding as subsequent training data.
3. The method according to claim 2, characterized in that The multimodal discrete codes are integrated into a unified code sequence, a mask strategy is used to cover the trajectory code, the original code without the mask is used as the supervision target, the model is generated through autoregressive paradigm training, and the model is fine-tuned on scene understanding, scene prediction and trajectory planning task data sets in turn, including: Integrate the discrete codes of the image modality, the text modality and the trajectory modality in the training data into a unified coding sequence, wherein a masking strategy is designed for the trajectory coding part to cover the target trajectory information; The original code without mask is used as the training supervision target, and the generative model is trained through the autoregressive paradigm to maximize the probability corresponding to the target code in the probability distribution of the model output; Performing basic training on the model, based on data sampled from a large-scale driving video dataset, to learn the basic capabilities of the training model; After the model is initially converged, the model is fine-tuned on a scene understanding task dataset, a scene prediction task dataset, and a trajectory planning task dataset in turn, and a weighted setting is performed on the loss function of the trajectory planning part during supervision; Multimodal datasets covering scene understanding, scene prediction, and trajectory planning tasks are integrated into a unified dataset and the model is fully fine-tuned.
4. The method according to claim 3, characterized in that The initial scene image captured by the vehicle camera is used as input, combined with the prompt words designed by the user, and the future scene discrete code sequence is output according to the trained model. The trajectory information in the discrete code sequence is replaced with the mask code and re-input into the model to complete the closed-loop scene prediction, including: The forward perspective image of the initial scene is obtained through the vehicle camera, and the image is encoded into a unified discrete codebook space to obtain the initial scene information; Input the prompt words designed by the user according to the task requirements, and input them into the trained model together with the initial scene information to generate the probability distribution of future encoding; Performing specific sampling on the probability distribution to generate a future scene discrete code sequence including future scene image code, scene information understanding code and vehicle trajectory code; Replacing the trajectory code in the future scene discrete code sequence with the mask code, and integrating the replaced code sequence into the original input; The integrated coding sequence is re-input into the model to generate a discrete coding sequence of scenes in the more distant future without providing future frame information of the vehicle's real camera, thereby achieving closed-loop scene prediction.
5. A multimodal autonomous driving scene generation device based on autoregressive closed-loop prediction, characterized in that: include: The encoding module is used to obtain a multi-modal and multi-task autonomous driving dataset, integrate and map it into a unified discrete space through different discrete encoders, and generate an integrated multi-modal discrete encoding; A training module is used to integrate the multimodal discrete codes into a unified code sequence, use a masking strategy to cover the trajectory code, use the original code without the mask as the supervision target, train the model through an autoregressive paradigm, and fine-tune the model on scene understanding, scene prediction, and trajectory planning task datasets in turn; The inference module is used to take the initial scene image captured by the vehicle camera as input, combine it with the prompt words designed by the user, output the discrete code sequence of the future scene according to the trained model, and complete the closed-loop scene prediction by replacing the trajectory information in the discrete code sequence with the mask code and re-inputting it into the model; The decoding module is used to decode the generated discrete code into images of future scenes, text questions and answers, and vehicle trajectory information using a decoder symmetrical to the corresponding modal encoder.
6. The device according to claim 5, characterized in that The encoding module is used for: Collect and integrate multimodal autonomous driving datasets including autonomous driving scene videos, scene understanding prediction question answering, and trajectory planning data; Use pre-trained VQGAN, BPE word segmenter and position embedding encoder to encode autonomous driving scene videos, scene understanding prediction question and answer, and trajectory planning data respectively, map the image modality, text modality and trajectory modality data to a unified discrete codebook space, and generate integrated multimodal discrete coding as subsequent training data.
7. The device according to claim 6, characterized in that The training module is used to: Integrate the discrete codes of the image modality, the text modality and the trajectory modality in the training data into a unified coding sequence, wherein a masking strategy is designed for the trajectory coding part to cover the target trajectory information; The original code without mask is used as the training supervision target, and the generative model is trained through the autoregressive paradigm to maximize the probability corresponding to the target code in the probability distribution of the model output; Performing basic training on the model, based on data sampled from a large-scale driving video dataset, to learn the basic capabilities of the training model; After the model is initially converged, the model is fine-tuned on a scene understanding task dataset, a scene prediction task dataset, and a trajectory planning task dataset in turn, and a weighted setting is performed on the loss function of the trajectory planning part during supervision; Multimodal datasets covering scene understanding, scene prediction, and trajectory planning tasks are integrated into a unified dataset and the model is fully fine-tuned.
8. The device according to claim 7, characterized in that The reasoning module is used to: The forward perspective image of the initial scene is obtained through the vehicle camera, and the image is encoded into a unified discrete codebook space to obtain the initial scene information; Input the prompt words designed by the user according to the task requirements, and input them into the trained model together with the initial scene information to generate the probability distribution of future encoding; Performing specific sampling on the probability distribution to generate a future scene discrete code sequence including future scene image code, scene information understanding code and vehicle trajectory code; Replacing the trajectory code in the future scene discrete code sequence with the mask code, and integrating the replaced code sequence into the original input; The integrated coding sequence is re-input into the model to generate a discrete coding sequence of scenes in the more distant future without providing future frame information of the vehicle's real camera, thereby achieving closed-loop scene prediction.
9. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 4 when executed by a processor.
Citation Information
Cited By
Track scene prediction method and device, equipment and medium
CN120766086A
Rail scene prediction method, device, equipment and medium
CN120766086B
Unmanned aerial vehicle visual navigation method and system based on world modeling
CN121053211A
Driving video generation method based on space-time factorization architecture and hybrid modulation
CN122093518A
Driving video generation method based on spatio-temporal factorization architecture and hybrid modulation
CN122093518B