World model understanding-based end-to-end control method for computing base platform

Through the end-to-end control method based on the world model understanding, vehicle control signals are directly generated from perceived data, solving the problem that fully end-to-end autonomous driving from perceived to control in the prior art is difficult to achieve, reducing the control cost and supporting a customized driving style, and improving the efficiency of autonomous driving.

CN120057020APending Publication Date: 2025-05-30TSINGHUA UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510333632.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing autonomous driving technology is difficult to achieve fully end-to-end autonomous driving from perceived to control, and the control is expensive and the driving style is uncustomized.

Method used

The end-to-end control method based on world model understanding is adopted, and multi-view video data sets, candidate input instructions, map information and text information are obtained, multi-view images and optimal control instructions are generated, and end-to-end control is performed through a unified interface of multi-modal information.

Benefits of technology

It realizes fully end-to-end autonomous driving from perceived to control, reduces the control cost, and supports a customized driving style, improving the efficiency of autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120057020A_ABST
    Figure CN120057020A_ABST
Patent Text Reader

Abstract

The invention relates to an end-to-end control method based on world model understanding for a computing basic platform, and the method comprises the steps: obtaining a multi-view video data set, a candidate input instruction, map information and text information of a current vehicle; generating a multi-view image according to the multi-view video data set, and generating an optimal control instruction according to a preset reward function based on the multi-view image, the candidate input instruction, the map information and the text information; and generating a multi-modal information uniform interface according to the multi-view image, the map information, the text information and the optimal control instruction, and performing end-to-end control on the current vehicle based on the multi-modal information uniform interface. Therefore, by constructing the end-to-end control model, the vehicle control signal is directly generated from the sensing data, the problems that in the prior art, complete end-to-end automatic driving from sensing to control is difficult to achieve, the control cost is high, and the driving style cannot be customized are solved, and the efficiency of automatic driving is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving technology, and particularly to an end-to-end control method based on world model understanding for a computing base platform. Background Art

[0002] The Computing Brain Development System (CBDES) consists of two parts: the Computing Base Brain (CBB) and the Graphical ADAS-AD Software Developer (GAASD). The CBB is composed of computing platform hardware, a real-time kernel, middleware, and functional software. The functional software is the core of this product, aiming to provide basic algorithm components and frameworks for various intelligent driving systems. CBDES transfers the application algorithm development ability to the host manufacturers, supporting the host factory engineers to quickly build their own defined intelligent driving systems and perform function adaptation and parameter tuning. Most of the current autonomous driving technologies adopt a multi-level architecture, including multiple links such as perception, positioning, planning, decision-making, and control, which helps to clarify the responsibilities of each functional module.

[0003] In related technologies, the structure of autonomous driving systems mostly adopts a traditional multi-level architecture, which is divided into links such as perception, positioning, planning, decision-making, and control; existing end-to-end autonomous driving attempts to integrate the entire process from perception to control into one model to simplify the system structure and improve the overall performance.

[0004] However, the multi-level architecture in related technologies requires a large amount of manual development and maintenance, and it is difficult to achieve collaborative optimization between each link; although the existing end-to-end autonomous driving has improved the global environment perception, it still needs to design a path tracking controller, and most of them can only achieve a direct mapping from perception to planning, and have not yet achieved a complete end-to-end automation from perception to control; in addition, the existing vehicle control solutions have problems of high control cost and non-customizable driving styles, which need to be solved urgently. Summary of the Invention

[0005] This application provides an end-to-end control method based on world model understanding for a computing base platform to solve problems such as the difficult-to-achieve complete end-to-end autonomous driving from perception to control, high control cost, and non-customizable driving styles in the prior art, and improves the efficiency of autonomous driving.

[0006] The first aspect embodiment of this application provides an end-to-end control method based on world model understanding for a computing base platform, including the following steps:

[0007] Obtain the multi-view video dataset, candidate input instructions, map information, and text information of the current vehicle;

[0008] Generate multi-view images based on the multi-view video dataset, and generate optimal control instructions based on the multi-view images, the candidate input instructions, the map information, and the text information according to a preset reward function;

[0009] Generate a multi-modal information unified interface based on the multi-view images, the map information, the text information, and the optimal control instructions, and perform end-to-end control of the current vehicle based on the multi-modal information unified interface.

[0010] Optionally, the generating multi-view images based on the multi-view video dataset includes:

[0011] Use the encoder of a preset variational auto-encoder network to extract features from the original image data in the multi-view video dataset to obtain feature data;

[0012] Obtain a joint modeling image based on the feature data, and generate the multi-view images according to a preset joint image modeling method based on a preset factorized multi-view view strategy.

[0013] Optionally, the obtaining a joint modeling image based on the feature data includes:

[0014] Use a preset diffusion model to introduce random noise into the feature data to obtain preprocessed data;

[0015] Denoise the preprocessed data to obtain denoised data, and perform VAE decoding on the denoised data to obtain the joint modeling image.

[0016] Optionally, the map information includes 3DBoxes, high-precision maps, and segmentation images in BEV view.

[0017] Optionally, the generating a multi-modal information unified interface based on the multi-view images, the map information, the text information, and the optimal control instructions includes:

[0018] Based on preset image conditions, encode the multi-view images using the ConvNeXt method, and flatten according to the first encoding result to obtain a first multi-dimensional embedding sequence;

[0019] Encode the map information using the ConvNeXt method, and flatten according to the second encoding result to obtain a second multi-dimensional embedding sequence.

[0020] Optionally, the preset reward function = first weight coefficient * (target reward * map reward) + second weight coefficient * control cost reward.

[0021] In the second aspect of the present application, an embodiment provides an end-to-end control device for a computing infrastructure platform based on world model understanding, including:

[0022] An acquisition module, configured to acquire a multi-view video data set, candidate input instructions, map information, and text information of the current vehicle;

[0023] A generation module, configured to generate multi-view images according to the multi-view video data set, and generate optimal control instructions according to a preset reward function based on the multi-view images, the candidate input instructions, the map information, and the text information;

[0024] A control module, configured to generate a multi-modal information unified interface according to the multi-view images, the map information, the text information, and the optimal control instructions, and perform end-to-end control on the current vehicle based on the multi-modal information unified interface.

[0025] Optionally, the generation module is specifically configured to:

[0026] Use an encoder of a preset variational auto-encoder network to extract features from the original image data in the multi-view video data set to obtain feature data;

[0027] Obtain a jointly modeled image according to the feature data, and generate the multi-view image according to the jointly modeled image based on a preset factorized multi-view view strategy.

[0028] Optionally, the generation module is specifically configured to:

[0029] Introduce random noise into the feature data using a preset diffusion model to obtain preprocessed data;

[0030] Perform denoising processing on the preprocessed data to obtain denoised data, and perform VAE decoding on the denoised data to obtain the jointly modeled image.

[0031] Optionally, the map information includes 3DBoxes, a high-precision map, and a segmented image in BEV view.

[0032] Optionally, the control module is specifically configured to:

[0033] Based on a preset image condition, use the ConvNeXt method to encode the multi-view image, and flatten the first encoding result to obtain a first multi-dimensional embedding sequence;

[0034] Use the ConvNeXt method to encode the map information, and flatten the second encoding result to obtain a second multi-dimensional embedding sequence.

[0035] Optionally, the preset reward function = the first weight coefficient * (target reward * map reward) + the second weight coefficient * control cost reward.

[0036] An embodiment of the third aspect of the present application provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method for calculating the end-to-end control based on world model understanding of the basic platform as described in the above embodiments.

[0037] An embodiment of the fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the method for calculating the end-to-end control based on world model understanding of the basic platform as described in the above embodiments.

[0038] Thus, after obtaining the multi-view video data set, candidate input instructions, map information, and text information of the current vehicle, generate multi-view images based on the multi-view video data set, and based on the multi-view images, candidate input instructions, map information, and text information, generate an optimal control instruction according to the preset reward function, and then generate a multi-modal information unified interface based on the multi-view images, map information, text information, and the optimal control instruction, and perform end-to-end control on the current vehicle based on the multi-modal information unified interface. Thus, by constructing an end-to-end control model, directly generating vehicle control signals from perception data, it solves the problems of difficult-to-implement fully end-to-end autonomous driving from perception to control, high control cost, and non-customizable driving styles in the prior art, and improves the efficiency of autonomous driving.

[0039] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Description of the Drawings

[0040] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0041] Figure 1 It is a flowchart of a method for calculating an end-to-end control based on world model understanding of a basic platform according to an embodiment of the present application;

[0042] Figure 2 It is a schematic diagram of the components of an end-to-end control algorithm network of a method for calculating an end-to-end control based on world model understanding of a basic platform according to an embodiment of the present application;

[0043] Figure 3Schematic diagram of a multi-perspective prediction video generation network structure for an end-to-end control method based on world model understanding for a computing infrastructure platform according to an embodiment of the present application;

[0044] Figure 4 Schematic diagram of the design of a video evaluation performance metric reward function for an end-to-end control method based on world model understanding for a computing infrastructure platform according to an embodiment of the present application;

[0045] Figure 5 Schematic diagram of an end-to-end control device based on world model understanding for a computing infrastructure platform according to an embodiment of the present application;

[0046] Figure 6 Schematic diagram of a structure of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0047] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.

[0048] An end-to-end control method based on world model understanding for a computing infrastructure platform according to an embodiment of the present application will be described below with reference to the accompanying drawings. In view of the problems in the prior art mentioned in the above background art, namely, the difficulty in realizing full end-to-end autonomous driving from perception to control, high control cost, and non-customizable driving styles, the present application provides an end-to-end control method based on world model understanding for a computing infrastructure platform. In this method, a multi-perspective video dataset, candidate input instructions, map information, and text information of the current vehicle are obtained; multi-perspective images are generated based on the multi-perspective video dataset, and an optimal control instruction is generated based on the multi-perspective images, candidate input instructions, map information, and text information according to a preset reward function; a multi-modal information unified interface is generated based on the multi-perspective images, map information, text information, and the optimal control instruction, and the current vehicle is subjected to end-to-end control based on the multi-modal information unified interface. Thus, by constructing an end-to-end control model to directly generate vehicle control signals from perception data, the problems in the prior art, such as the difficulty in realizing full end-to-end autonomous driving from perception to control, high control cost, and non-customizable driving styles, are solved, and the efficiency of autonomous driving is improved.

[0049] Specifically, Figure 1 Schematic diagram of a flow of an end-to-end control method based on world model understanding for a computing infrastructure platform provided by an embodiment of the present application.

[0050] As Figure 1As shown in the figure, the end-to-end control method based on world model understanding for the computing infrastructure platform includes the following steps:

[0051] In step S101, a multi-view video dataset of the current vehicle, candidate input instructions, map information, and text information are obtained.

[0052] Among them, the multi-view video dataset contains video clips taken from multiple perspectives around the vehicle, which helps the system better understand and predict changes in the surrounding environment; the candidate input instructions are selected from multiple possible control options and are used to guide the vehicle's next action.

[0053] Specifically, in the embodiments of the present application, multi-view images of the surrounding environment can be captured by on-vehicle cameras installed on the vehicle. These cameras are usually distributed at different positions of the vehicle, such as in the front, rear, both sides, etc., so as to comprehensively cover the field of view around the vehicle; the data acquisition system is responsible for collecting real-time video streams from each camera, converting them into a data format suitable for subsequent processing, and storing them for use by the model. An operation interface provided for passengers or drivers is used to input information such as destinations and driving preferences, and these information will be part of the candidate input instructions; based on the current location information, destination, and traffic rules, a series of possible driving routes and corresponding control instructions are automatically generated as candidate input instructions. By connecting to the high-precision map API, detailed map information near the current location is obtained, including road types, traffic signs, speed limits, etc.; the exact current position of the vehicle is determined to extract relevant map data from the map service. If the system supports voice commands, the user's voice input needs to be received through the built-in microphone and converted into text information through speech recognition technology; in addition to voice commands, the user's text input, such as address input and special requirements, can also be received through a touch screen or other input devices.

[0054] Optionally, in some embodiments, the map information includes 3D Boxes, high-precision maps, and segmentation images from the BEV perspective.

[0055] Among them, 3D Boxes refers to a data structure representing object bounding boxes in three-dimensional space. These bounding boxes are usually used to label the positions and sizes of traffic participants such as vehicles, pedestrians, and bicycles; a high-precision map (High-Definition Map, HD Map) is a detailed digital map that contains rich geographical information such as road geometries, lane lines, traffic signs, and traffic lights; the BEV perspective refers to a perspective looking down from above, similar to a bird's-eye view. The segmentation image from the BEV perspective generates an annotation map by performing pixel-level classification on different objects in the image to distinguish roads, lane lines, obstacles, etc.

[0056] It can be understood that the processing and use of map information involve fusing multiple data sources such as 3D Boxes, high-precision maps, and segmented images from the BEV perspective to generate comprehensive map information; ensuring the consistency of the time and space coordinates of different data sources for subsequent processing and analysis; using a convolutional neural network (such as ConvNeXt) to encode the segmented images from the BEV perspective to generate a multi-dimensional embedding sequence; projecting 3D Boxes and high-precision maps into a 2D perspective view and also using a convolutional neural network for encoding to generate a multi-dimensional embedding sequence; mapping data of different modalities such as the encoded image information, map information, and text information into the same feature space to form a unified multi-modal information interface; and enabling the conditional embedding to interact with the latent features through a cross-attention mechanism to enhance the model's understanding and prediction ability of the environment.

[0057] In step S102, multi-view images are generated based on the multi-view video dataset, and an optimal control instruction is generated based on the multi-view images, candidate input instructions, map information, and text information according to a preset reward function.

[0058] Among them, the preset reward function is used to evaluate the effects of different control instructions, so as to select the optimal control instruction.

[0059] Specifically, in the embodiments of the present application, the encoder of a variational autoencoder (VAE) can be used to extract features from the original image data in the multi-view video dataset to obtain feature data; a diffusion model is used to introduce random noise into the feature data to obtain preprocessed data; the preprocessed data is denoised to obtain denoised data; the denoised data is decoded by VAE to obtain a jointly modeled image; based on the jointly modeled image, a factored multi-view view strategy is used to generate multi-view images; the multi-view images and map information are encoded using the ConvNeXt method, and the encoded results are flattened to obtain a multi-dimensional embedding sequence; according to the preset reward function, the performance of candidate control instructions is evaluated, and the control instruction with the maximum reward is selected as the optimal control instruction.

[0060] Optionally, in some embodiments, generating multi-view images based on the multi-view video dataset includes: using the encoder of a preset variational autoencoder to extract features from the original image data in the multi-view video dataset to obtain feature data; obtaining a jointly modeled image based on the feature data, and generating multi-view images according to a preset jointly image modeling method based on a preset factored multi-view view strategy.

[0061] Among them, the encoder of the preset variational autoencoder network is one of the key components for generating multi-view images. The encoder is used to extract useful features from the multi-view video dataset; the feature data refers to the data obtained by extracting features from the original image data in the multi-view video dataset through the encoder of the preset variational autoencoder network, which contains the key information of the original image and can be used for subsequent joint modeling image generation and multi-view image generation.

[0062] It can be understood that feature extraction is achieved through a variational autoencoder network, which can effectively extract useful feature information from the original image data; introducing random noise is to expand the dataset and improve the generalization performance of the network; denoising processing is to obtain clearer image data; the factorized multi-view view strategy ensures the consistency of images between different views and expands the visual blank between multiple views. The multi-view images generated through this process provide richer environmental perception information for the autonomous driving system, which helps to improve the decision-making and control performance of the system; enables the system to directly generate control instructions from the perception data without traditional control algorithms; adopts a data-driven method, through pre-training and fine-tuning, can greatly reduce the number of code lines, improve efficiency, and save resources and costs.

[0063] Optionally, in some embodiments, obtaining the joint modeling image according to the feature data includes: introducing random noise into the feature data using a preset diffusion model to obtain preprocessed data; performing denoising processing on the preprocessed data to obtain denoised data, and performing VAE decoding on the denoised data to obtain the joint modeling image.

[0064] Among them, the preset diffusion model is a generative model that destroys data by gradually adding noise and then learns the inverse process to recover the original data from the noise.

[0065] It can be understood that by introducing random noise into the feature data, more data samples can be generated, thereby enhancing the generalization ability of the model; the introduction of noise is equivalent to regularizing the model, which helps to prevent the model from overfitting; the denoising processing step is responsible for extracting key information, removing redundancy and noise, and providing high-quality input for VAE decoding; finally, through the action of the VAE decoder, the denoised data is reconstructed into a high-quality joint modeling image. The generated joint modeling image retains the key features of the original image while removing noise and redundant information; improves the consistency and coordination between multi-view images, providing a solid foundation for subsequent multi-view image generation.

[0066] Optionally, in some embodiments, the preset reward function = the first weight coefficient * (the target reward * the map reward) + the second weight coefficient * the control cost reward.

[0067] Among them, the target reward refers to the distances from other road users in the longitudinal and lateral directions. This reward is set to avoid collisions between the autonomous vehicle and other road users and ensure driving safety. The map reward includes two factors: the distance from the curb and the centerline consistency. The distance-from-curb reward encourages the vehicle to stay in the correct drivable area, while the centerline-consistency reward prevents the vehicle from frequently changing lanes and deviating from the lane laterally, thus maintaining driving stability. The control-cost reward is a negative index of the absolute value of the vehicle's acceleration. This reward is set to avoid frequent acceleration or deceleration of the vehicle, minimize the absolute value of the acceleration, increase the comfort of the ride, and also reflect the cost of the control link and ensure the stability of the control link.

[0068] It can be understood that by adjusting the first weight coefficient and the second weight coefficient, the relationship among safety, stability, and comfort can be balanced, thereby obtaining the optimal control command. The preset reward function provides a basis for the autonomous vehicle to select the optimal control command through the comprehensive evaluation of future scenarios.

[0069] In step S103, a multi-modal information unified interface is generated based on the multi-view images, map information, text information, and the optimal control command, and the current vehicle is controlled end-to-end based on the multi-modal information unified interface.

[0070] Specifically, define the format of the multi-modal information unified interface, including the dimension, data type, etc. of the input feature vector. The interface should also support receiving the optimal control command and outputting a control signal to the vehicle actuator. Use a large amount of labeled multi-modal data and the corresponding optimal control commands to train an end-to-end deep learning model. The model should be able to receive the feature vector input by the multi-modal information unified interface and output a control command. During the vehicle's driving process, multi-view images, map information, and text information are collected in real time. Through preprocessing and feature fusion, a unified feature vector is formed and input into the trained end-to-end model. The model outputs the optimal control command, such as the steering angle, acceleration, etc., and converts the control command into a control signal executable by the vehicle, such as the steering motor signal, throttle signal, etc. According to the actual driving situation and feedback data of the vehicle, the model is adjusted and optimized online. The reinforcement learning method can be used to continuously optimize the control strategy through trial and error and the reward mechanism.

[0071] Optionally, in some embodiments, generating a multi-modal information unified interface based on the multi-view images, map information, text information, and the optimal control command includes: encoding the multi-view images using the ConvNeXt method based on the preset image conditions, and flattening the first encoding result to obtain the first multi-dimensional embedding sequence; encoding the map information using the ConvNeXt method, and flattening the second encoding result to obtain the second multi-dimensional embedding sequence.

[0072] Among them, the ConvNeXt model is an image classification model based on the Transformer architecture. It adopts more concise and efficient convolutional operations and can extract deep features in images. The first encoding result is the output obtained after encoding multi-view image data using a convolutional neural network (ConvNeXt); the second encoding result is the output obtained after encoding map information using a convolutional neural network; the first multi-dimensional embedding sequence is obtained by further processing the first encoding result, usually by flattening the encoding result or converting it into a high-dimensional vector; the second multi-dimensional embedding sequence is obtained by further processing the second encoding result, by flattening the encoding result or converting it into a high-dimensional vector.

[0073] It can be understood that the preset image conditions may include requirements such as the resolution, brightness, contrast, color space, etc. of the image to ensure that the image data input to the ConvNeXt model meets certain quality and format standards, and may also involve the consistency requirements for multi-view images to ensure that images from different perspectives can accurately reflect the environment around the vehicle; inputting the multi-view image data into the ConvNeXt model, through the encoding process of the ConvNeXt model, the multi-view images are converted into high-dimensional feature representations, and these feature representations contain key information in the images, such as the shape, texture, position, etc. of objects; performing a flattening operation on the output (i.e., the encoded feature representation) of the ConvNeXt model to obtain the first multi-dimensional embedding sequence; inputting the map information into the ConvNeXt model and performing encoding processing in the same way, and the encoded map information is flattened into the second multi-dimensional embedding sequence for subsequent fusion with the features of the multi-view images.

[0074] Thus, by using the ConvNeXt model to encode and process multi-view image and map information, it is possible to efficiently extract the key features in this information and generate multi-dimensional embedding sequences, providing strong support for subsequent algorithm processing, enabling the autonomous driving system to more accurately understand the environment around the vehicle and make more intelligent decisions. In addition, due to the efficient and concise characteristics of the ConvNeXt model, it can also effectively reduce the use of computing resources and improve the real-time performance and stability of the autonomous driving system.

[0075] Next, a specific embodiment of the present application will be used to detail the end-to-end control method based on world model understanding for the computing infrastructure platform.

[0076] Specifically, as Figure 2 shown, the schematic diagram of the composition of the end-to-end control algorithm network of the end-to-end control method based on world model understanding for the computing infrastructure platform includes the following steps:

[0077] This architecture is an improvement based on an end-to-end world model. Compared with the original perception-to-planning end control model, the world model proposed in the embodiments of this application realizes an end-to-end control architecture design that directly reaches vehicle control signals based on perception data. This architecture can be mainly divided into three parts, namely, a multi-view video generation network, a control tree generation based on input instructions, and a reward function based on images and inputs.

[0078] Thus, by constructing core components such as a multi-view video prediction and generation network, a control tree generation based on input instructions, a reward function based on images and inputs, and a multi-modal information unified interface, end-to-end real-time control from the perception end to the control end is achieved, which has significant technical innovation points and beneficial effects.

[0079] Furthermore, as Figure 3 shown, the following steps are for the schematic diagram of the multi-view prediction video generation network structure of the end-to-end control method based on world model understanding for the computing basic platform:

[0080] In order to jointly model multi-viewpoint time data, taking the more mature image diffusion model in existing research as the baseline, and making it adapt to the multi-viewpoint time scenario by introducing additional time layers and multi-viewpoint layers. The joint modeling network of images, the factorized multi-view generation process, and the fusion input mechanism of multi-modal information are mainly introduced therein.

[0081] The joint modeling network of images: In the Figure 3 pipeline shown, taking the image data of T frames and K viewpoints in the dataset as an example for introduction, first, the encoder of the Variational Autoencoder (VAE) is used to extract features; subsequently, in order to expand the dataset and improve the generalization performance of the network, a relatively mature diffusion model is used to introduce random noise into the feature data; then the preprocessed data is denoised; then the denoised result is decoded by VAE to obtain the jointly modeled image. According to the established practice in VideoLDM, a time encoding layer is appended after the 2D spatial layer of each block. The spatial layer encodes the latent features frame by frame and view by view. Finally, the latent features are rearranged to enhance the temporal dependence.

[0082] Finally, after the pre-training is completed, the multi-viewpoint time dimension parameters are fine-tuned. During the training process, first, the standard image diffusion model is pre-trained using single-view image data, then the parameters of the diffusion model are fixed, and then the parameters of the additional time layer and multi-viewpoint layer are fine-tuned using video data. In summary, multi-viewpoint joint modeling of the original input image can be achieved.

[0083] Factorized Multi-View Generation: Although the images obtained above are jointly modeled and their joint probability distribution can generate similar patterns among different views, it is difficult to ensure strict consistency in their overlapping regions, nor can the visual gaps between multiple views be extended. Therefore, joint model decomposition is needed to enhance the consistency of multi-views.

[0084] Therefore, a factorized multi-view generation process is constructed, generating different views in an autoregressive manner, that is, the new view is conditioned on the existing views, and its joint conditional distribution p( 1,…,K ) can be expressed as:

[0085] p(x 1,…,K ) = p(x 1 )p(x 2 |x 1 )…p(x K |x 1 ,…,x K-1 ) (1);

[0086] In the above formula, x i represents the image of the i-th view. However, such an autoregressive generation method is inefficient, making such a full decomposition infeasible in practice. To simplify the modeling in formula (1), all views need to be divided into two types: reference views x r and stitching views x s . Views belonging to the same type do not overlap with each other, while views of different types may overlap. Therefore, it is necessary to first model the joint distribution of the reference views. Here, joint modeling is effective for those non-overlapping reference views that do not require strict consistency. The distribution of x s is modeled as a conditional distribution conditioned on x r , and formula (1) can be simplified to:

[0087] p(x) = p(x s ,x r ) = p(x r )p(x s |x r ) (2)

[0088] Considering temporal coherence and incorporating the previous framework as an additional condition, formula (2) can be rewritten as:

[0089] p(x) = p(x s ,x r |x pre ) = p(x r |x pre )p(x s |x r ,x pre ) (3)

[0090] where x pre are adjacent frames of previously generated video clips (e.g., the last two frames). The distribution p(x r |x pre ) is implemented by the structure in (1). For p(x s , x r |x pre ), a similar structure is adopted, but adjacent reference views are merged as additional conditions.

[0091] Generation of the multi-modal information unified interface: Due to the great complexity of the real world, the world model needs to utilize multi-modal conditions as input information, such as the initial context framework, text description, ego vehicle behavior, 3D Boxes, BEV images, and information of each reference view, etc. More conditions are beneficial for achieving better controllability, but developing a dedicated data format for each condition is too time-consuming and inflexible. Therefore, a unified data format is defined to simply and effectively integrate the input information of multiple modalities. Next, first introduce how to encode each input condition, and then describe the interface for unifying this information.

[0092] Multi-modal information: ① Image information: Use the initial context frame and reference views as image conditions, and encode the given image conditions using the ConvNeXt method and flatten them into a d-dimensional embedding sequence

[0093] ② Map information: Map information refers to 3D Boxes, high-precision maps, and segmented images from the BEV perspective. Project 3D Boxes and high-precision maps into a 2D perspective view, and use the same strategy as encoding image information to encode map information, generating an embedding sequence

[0094] ③ Text information: Following the convention of the diffusion model, use the pre-trained CLIP as the text encoder. Specifically, combine view information, weather, and light to derive the text description, labeled as

[0095] Control Information: The control condition is an indispensable condition for the world model to generate the future. To be compatible with existing planning methods, the original world model defines the control instruction within a time step as (Δx, Δy), which represents the predicted lateral and longitudinal displacements of the vehicle in the future. However, the drawback of this approach is that the model can only learn the mapping relationship between the input image information and the ideal displacement (Δx, Δy) of the vehicle. To obtain the relationship between the final input image information and the actual control signal of the vehicle, it is necessary to introduce control signals h and e in the multi-modal input of UNet, which represent the throttle opening and the actual steering angle of the steering wheel of the vehicle, respectively. Similar to the practice in the original world model, an MLP is used to map the control action instruction signal to a d-dimensional embedding. In this way, the end-to-end network can learn the mapping relationship between the multi-view images of in-vehicle sensors and the actual control signal of the vehicle, thus realizing end-to-end real-time control from the perception end to the control end. One significant difference from end-to-end planning is that the end-to-end control algorithm in the embodiments of this application replaces the self-behavior information of the vehicle with the control information of the vehicle, directly connecting to the underlying control signal of the vehicle, eliminating the design and parameter adjustment links of the control system.

[0096] In summary, all conditions are mapped to a d-dimensional feature space. This combination of different conditions provides a unified data format and can be adjusted according to requirements. Finally, the condition embedding and the latent feature interact with each other frame by frame through the cross-attention mechanism.

[0097] (1) Generation of the control tree based on the input instruction.

[0098] In this section, planning and control using the world model will be described. At each time step, the world model is used to generate a predicted future scenario for the candidate control instructions sampled from the controller, evaluate the future using an image-based reward function, and select the optimal control instruction to expand the control tree. The control tree is defined as a series of predicted self-control execution results evolving over time. The pre-trained planner takes the real multi-view images as input and samples possible execution trajectories. For a given control instruction, conditional combinations are used for video generation. After generation, the image-based reward function is used to select the optimal control execution result as the final decision.

[0099] (2) Image-based reward function.

[0100] After generating a future video for the execution result after control, a reward function is needed to evaluate the performance of multiple future results, which specifically includes four aspects of settings. First, obtain rewards from the perception results, and use an image-based 3D object detector and an online HDMap predictor to obtain the perception results of the generated video. Second, the map reward includes two factors, namely the distance from the curb and the consistency with the center line. The former encourages the vehicle to stay in the correct drivable area, and the latter prevents frequent lane changes and lane deviations in the lateral direction. Third, the target reward refers to the distance from other road users in the longitudinal and lateral directions, avoiding collisions between the ego vehicle and other road users. Fourth, a control cost reward is also defined, that is, the absolute value of the vehicle acceleration, so as to avoid frequent acceleration or deceleration of the vehicle, and at the same time make the absolute value of the acceleration as small as possible to increase comfort. If it is necessary to change the driving style or requirements, this reward can also be changed according to the specific situation. This reward can reflect the cost and feedback of the control link, ensuring the stability of the control link and the comfort of the occupants.

[0101] The total reward is defined as the weighted sum of the product of the target reward and the map reward and the control cost reward. Finally, select the ego vehicle control quantity prediction with the maximum reward. Then the planning tree is forwarded to the next timestamp, iteratively generating subsequent control instructions and the reward mechanism design of the results, as Figure 4 shown.

[0102] Thus, the schematic diagram of the network generation structure shows an end-to-end process from multi-modal conditional input to final view generation, including key steps such as feature extraction, data fusion, generation of control signal sequences, denoising and enhancement, VAE decoding and view generation. This process provides an end-to-end control method based on world model understanding for the computing infrastructure platform, which can simulate and predict the dynamic behavior of the real world.

[0103] According to the end-to-end control method based on world model understanding for the computing infrastructure platform proposed in the embodiments of the present application, obtain the multi-view video dataset, candidate input instructions, map information and text information of the current vehicle; generate multi-view images based on the multi-view video dataset, and generate optimal control instructions based on the multi-view images, candidate input instructions, map information and text information according to a preset reward function; generate a multi-modal information unified interface based on the multi-view images, map information, text information and optimal control instructions, and perform end-to-end control on the current vehicle based on the multi-modal information unified interface. Thus, by constructing an end-to-end control model, directly generating vehicle control signals from perception data, it solves the problems of full end-to-end autonomous driving from perception to control, high control cost and non-customizable driving style that are difficult to achieve in the prior art, and improves the efficiency of autonomous driving.

[0104] Next, a description will be given of an end-to-end control device for a computing infrastructure based on world model understanding according to an embodiment of the present application with reference to the accompanying drawings.

[0105] Figure 5 It is a block diagram of an end-to-end control device for a computing infrastructure based on world model understanding according to an embodiment of the present application.

[0106] As Figure 5 shown, the end-to-end control device 10 for a computing infrastructure based on world model understanding includes: an acquisition module 100, a generation module 200, and a control module 300.

[0107] Among them, the acquisition module 100 is configured to acquire a multi-view video data set, a candidate input instruction, map information, and text information of the current vehicle;

[0108] The generation module 200 is configured to generate multi-view images according to the multi-view video data set, and generate an optimal control instruction based on the multi-view images, the candidate input instruction, the map information, and the text information according to a preset reward function;

[0109] The control module 300 is configured to generate a multi-modal information unified interface according to the multi-view images, the map information, the text information, and the optimal control instruction, and perform end-to-end control on the current vehicle based on the multi-modal information unified interface.

[0110] Optionally, in some embodiments, the generation module 200 is specifically configured to: extract features from the original image data in the multi-view video data set by using an encoder of a preset variational auto-encoding network to obtain feature data; obtain a jointly modeled image according to the feature data, and generate multi-view images based on the jointly modeled image according to a preset factorized multi-view view strategy.

[0111] Optionally, in some embodiments, the generation module 200 is specifically configured to: introduce random noise into the feature data by using a preset diffusion model to obtain preprocessed data; perform denoising processing on the preprocessed data to obtain denoised data, and perform VAE decoding on the denoised data to obtain a jointly modeled image.

[0112] Optionally, in some embodiments, the map information includes 3DBoxes, a high-precision map, and a segmented image in a BEV view.

[0113] Optionally, in some embodiments, the control module 300 is specifically configured to: encode the multi-view images by using the ConvNeXt method based on a preset image condition, and flatten the first encoding result to obtain a first multi-dimensional embedding sequence; encode the map information by using the ConvNeXt method, and flatten the second encoding result to obtain a second multi-dimensional embedding sequence.

[0114] Optionally, in some embodiments, the preset reward function = the first weight coefficient * (the target reward * the map reward) + the second weight coefficient * the control cost reward.

[0115] It should be noted that the foregoing explanation of the embodiments of the end-to-end control method for calculating the basic platform based on world model understanding also applies to the end-to-end control device for calculating the basic platform based on world model understanding in this embodiment, and will not be elaborated here.

[0116] The end-to-end control device for calculating the basic platform based on world model understanding proposed according to the embodiments of the present application acquires a multi-view video data set, candidate input instructions, map information, and text information of the current vehicle; generates multi-view images based on the multi-view video data set, and generates optimal control instructions based on the multi-view images, candidate input instructions, map information, and text information according to the preset reward function; generates a multi-modal information unified interface based on the multi-view images, map information, text information, and optimal control instructions, and performs end-to-end control on the current vehicle based on the multi-modal information unified interface. Thus, by constructing an end-to-end control model, vehicle control signals are directly generated from perception data, solving problems such as the difficult-to-implement fully end-to-end autonomous driving from perception to control, high control cost, and non-customizable driving styles in the prior art, and improving the efficiency of autonomous driving.

[0117] Figure 6 It is a schematic structural diagram of an electronic device provided in an embodiment of the present application. The electronic device may include:

[0118] A memory 601, a processor 602, and a computer program stored on the memory 601 and executable on the processor 602.

[0119] When the processor 602 executes the program, it implements the end-to-end control method for calculating the basic platform based on world model understanding provided in the above embodiments.

[0120] Further, the electronic device further includes:

[0121] A communication interface 603 for communication between the memory 601 and the processor 602.

[0122] The memory 601 is used to store a computer program executable on the processor 602.

[0123] The memory 601 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0124] If the memory 601, the processor 602, and the communication interface 603 are implemented independently, the communication interface 603, the memory 601, and the processor 602 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 only a thick line is used in Figure 6 , but it does not mean that there is only one bus or one type of bus.

[0125] Optionally, in a specific implementation, if the memory 601, the processor 602, and the communication interface 603 are integrated on a single chip, the memory 601, the processor 602, and the communication interface 603 can communicate with each other through an internal interface.

[0126] The processor 602 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0127] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, it implements the above-mentioned end-to-end control method for a computing infrastructure platform based on world model understanding.

[0128] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0129] In addition, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present application, "N" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0130] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment, or portion of code including one or more N executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present application includes additional implementations where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0131] It should be understood that each part of the present application may be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods may be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art may be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0132] Those of ordinary skill in the art of the present technology may understand that all or part of the steps carried by the methods of the above embodiments may be completed by instructing relevant hardware through a program. The program may be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

Claims

1. An end-to-end control method based on world model understanding for a computing infrastructure platform, characterized in that: The following steps are involved: Obtain a multi-view video dataset, candidate input instructions, map information, and text information of the current vehicle; Generate a multi-view image according to the multi-view video data set, and generate an optimal control instruction based on the multi-view image, the candidate input instruction, the map information and the text information according to a preset reward function; A multimodal information unified interface is generated according to the multi-view images, the map information, the text information and the optimal control instruction, and end-to-end control of the current vehicle is performed based on the multimodal information unified interface.

2. The method according to claim 1, characterized in that The generating a multi-view image according to the multi-view video data set comprises: Using a preset encoder of a variational autoencoder network, extracting features from the original image data in the multi-view video data set to obtain feature data; A joint modeling image is obtained according to the feature data, and based on a preset factorized multi-view strategy, the multi-view image is generated according to a preset joint image modeling method.

3. The method according to claim 2, characterized in that The step of obtaining a joint modeling image according to the feature data includes: Using a preset diffusion model to introduce random noise into the characteristic data to obtain preprocessed data; The preprocessed data is denoised to obtain denoised data, and the denoised data is VAE decoded to obtain the joint modeling image.

4. The method according to claim 1, characterized in that: The map information includes 3DBoxes, high-precision maps and segmented images from the BEV perspective.

5. The method according to claim 4, characterized in that The generating of a unified multimodal information interface according to the multi-view image, the map information, the text information and the optimal control instruction includes: Based on a preset image condition, encoding the multi-view image using a ConvNeXt method, and flattening the first encoding result to obtain a first multi-dimensional embedding sequence; The map information is encoded using the ConvNeXt method, and a second multi-dimensional embedding sequence is obtained by flattening the second encoding result.

6. The method according to claim 1, characterized in that The preset reward function=first weight coefficient*(target reward*map reward)+second weight coefficient*control cost reward.

7. An end-to-end control device based on world model understanding for a computing infrastructure platform, characterized in that: The following steps are involved: An acquisition module is used to acquire a multi-view video dataset, candidate input instructions, map information, and text information of the current vehicle; A generating module, configured to generate a multi-view image according to the multi-view video data set, and generate an optimal control instruction based on the multi-view image, the candidate input instruction, the map information and the text information according to a preset reward function; A control module is used to generate a multimodal information unified interface according to the multi-view images, the map information, the text information and the optimal control instruction, and to perform end-to-end control on the current vehicle based on the multimodal information unified interface.

8. The device according to claim 7, characterized in that The generation module is specifically used for: Using a preset encoder of a variational autoencoder network, extracting features from the original image data in the multi-view video data set to obtain feature data; A joint modeling image is obtained according to the feature data, and the multi-view image is generated according to the joint modeling image based on a preset factorized multi-view strategy.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement an end-to-end control method based on world model understanding for a computing base platform as described in any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement an end-to-end control method based on world model understanding for a computing base platform as described in any one of claims 1-6.

Citation Information

Cited By

  • Automatic driving end-to-end model self-correction method and device and medium

    CN120564157A

  • Self-correction method and device for end-to-end model of automatic driving and medium

    CN120564157B