End-to-end automatic driving method, and automatic driving model training method and device
By introducing lightweight language models and pruning technology into the end-to-end autonomous driving model, the problems of insufficient visual understanding and command understanding are solved, more efficient autonomous driving performance and lower deployment complexity are achieved, and the safety of autonomous driving and the human-vehicle interaction experience are improved.
Patent Information
- Application Number
- CN202510805659.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-10-03
AI Technical Summary
Existing end-to-end autonomous driving models suffer from insufficient visual understanding (visual misunderstanding) and insufficient command understanding (command misunderstanding), and their high network complexity and large number of parameters lead to deployment challenges.
A lightweight language model and pruning technology are used to prune visual features through the CCDP mechanism, retaining key visual features and aligning them to the language space. The lightweight language model is then used to process navigation instructions and generate driving instructions.
It reduces visual misunderstandings, improves recognition accuracy and safety of autonomous driving, reduces model complexity, facilitates deployment, and enhances the convenience of human-vehicle interaction.
Smart Images

Figure CN120747896A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to computer vision, deep learning, large models and other technical fields, and can be applied to scenarios such as autonomous driving. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, autonomous vehicle systems are undergoing a significant evolution from modular design to end-to-end integration. In this process, the end-to-end architecture, with its efficient and direct input-output mapping, enables vehicles to generate corresponding control commands based on raw input data, such as the road environment, further enhancing autonomous driving capabilities. Summary of the Invention
[0003] The present disclosure provides an end-to-end autonomous driving method, an autonomous driving model training method, and an apparatus.
[0004] According to one aspect of the present disclosure, an end-to-end autonomous driving method is provided, including:
[0005] Extracting first visual features from visual data collected by the autonomous vehicle;
[0006] Prune the first visual features to obtain key visual features;
[0007] Align the key visual features to the language space to obtain the second visual features corresponding to the key visual features;
[0008] Processing the second visual features and navigation instructions based on the language model to obtain the future trajectory of the autonomous vehicle;
[0009] Generate driving instructions for autonomous vehicles based on future trajectories.
[0010] According to another aspect of the present disclosure, an end-to-end autonomous driving device is provided, comprising:
[0011] a first extraction unit, configured to extract a first visual feature from visual data collected by the autonomous driving vehicle;
[0012] A pruning unit, configured to prune the first visual feature to obtain a key visual feature;
[0013] an alignment unit, configured to align the key visual feature to the language space to obtain a second visual feature corresponding to the key visual feature;
[0014] A prediction unit, configured to process the second visual features and navigation instructions based on the language model to obtain the future trajectory of the autonomous vehicle;
[0015] A control unit is used to generate driving instructions for the autonomous vehicle based on the future trajectory.
[0016] According to one aspect of the present disclosure, an end-to-end autonomous driving model training method is provided, comprising:
[0017] The visual encoder based on the end-to-end autonomous driving model extracts third-party visual features from the sample vision;
[0018] The connector based on the end-to-end autonomous driving model prunes the third visual feature to obtain the core visual feature; aligns the core visual feature to the language space to obtain the fourth visual feature corresponding to the core visual feature;
[0019] The language model based on the end-to-end autonomous driving model processes the fourth visual feature and the command sample to obtain the expected trajectory of the autonomous driving vehicle;
[0020] Based on the first loss between the expected trajectory and the corresponding data ground truth, the connector and language model are optimized.
[0021] According to another aspect of the present disclosure, an end-to-end autonomous driving model training device is provided, comprising:
[0022] A second extraction unit is configured to extract a third visual feature from the sample vision based on a visual encoder of an end-to-end autonomous driving model;
[0023] A sparse unit is used to prune the third visual feature based on the connector of the end-to-end autonomous driving model to obtain a core visual feature; and align the core visual feature to the language space to obtain a fourth visual feature corresponding to the core visual feature;
[0024] an estimation unit, configured to process the fourth visual feature and the command sample based on a language model of the end-to-end autonomous driving model to obtain an expected trajectory of the autonomous driving vehicle;
[0025] An optimization unit for optimizing the connector and the language model based on a first loss between the expected trajectory and the corresponding data ground truth.
[0026] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0027] at least one processor; and
[0028] a memory communicatively connected to the at least one processor; wherein,
[0029] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.
[0030] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.
[0031] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.
[0032] According to another aspect of the present disclosure, an autonomous driving vehicle is provided, comprising the aforementioned electronic device.
[0033] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0035] Figure 1 is a schematic flow chart of an end-to-end autonomous driving method according to an embodiment of the present disclosure;
[0036] Figure 2 is a schematic structural diagram of different functional modules of an end-to-end autonomous driving model according to an embodiment of the present disclosure;
[0037] Figure 3 is a schematic diagram of a process for obtaining key visual features according to an embodiment of the present disclosure;
[0038] Figure 4 is a schematic diagram of a process for obtaining a second visual feature corresponding to a key visual feature according to an embodiment of the present disclosure;
[0039] Figure 5 1 is a schematic diagram of the internal implementation structure of a connector of an end-to-end autonomous driving model according to an embodiment of the present disclosure;
[0040] Figure 6 is a schematic diagram of a process for generating a future trajectory of a vehicle according to an embodiment of the present disclosure;
[0041] Figure 7 2 is a schematic diagram of feature attention calculation of an attention module of a language model according to an embodiment of the present disclosure;
[0042] Figure 8 is a flowchart of an end-to-end autonomous driving model training method according to an embodiment of the present disclosure;
[0043] Figure 9is a schematic diagram of a process for obtaining core visual features according to an embodiment of the present disclosure;
[0044] Figure 10 is a schematic diagram of a process for obtaining a fourth visual feature corresponding to a core visual feature according to an embodiment of the present disclosure;
[0045] Figure 11 is a schematic diagram of a process for obtaining an expected trajectory for controlling an autonomous vehicle according to an embodiment of the present disclosure;
[0046] Figure 12 is a schematic diagram of a process for obtaining a fifth visual feature according to an embodiment of the present disclosure;
[0047] Figure 13 is a schematic diagram of the overall structure of an end-to-end autonomous driving model according to an embodiment of the present disclosure;
[0048] Figure 14 1 is a schematic diagram showing the effect of an end-to-end autonomous driving method according to an embodiment of the present disclosure;
[0049] Figure 15 is a schematic structural diagram of an end-to-end autonomous driving device according to an embodiment of the present disclosure;
[0050] Figure 16 is a schematic structural diagram of an end-to-end autonomous driving model device according to an embodiment of the present disclosure;
[0051] Figure 17 This is a block diagram of an electronic device used to implement the end-to-end autonomous driving method and / or end-to-end autonomous driving model training method of the embodiments of the present disclosure. DETAILED DESCRIPTION
[0052] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0053] The terms "first," "second," and the like in this disclosure are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. Furthermore, the terms "including," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions, such as, for example, inclusion of a series of steps or elements. A method, system, product, or apparatus is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.
[0054] It should be noted that, unless it is explicitly stated that there is a sequence of execution between different operations shown in the flowchart in the embodiments of the present disclosure, or there is a sequence of execution between different operations in technical implementation, otherwise, the execution order between multiple operations may not be prioritized, and multiple operations may also be executed simultaneously.
[0055] With the rapid development of artificial intelligence (AI), autonomous driving technology has evolved from a modular design to an integrated end-to-end network. This network supports end-to-end autonomous driving. The input of the end-to-end network is raw sensor data (such as camera images and lidar point clouds). After processing this input, the end-to-end network can output specific driving commands (such as steering angle and acceleration).
[0056] Early end-to-end autonomous driving approaches were primarily categorized into two categories: imitation learning and reinforcement learning. Subsequently, benefiting from the powerful feature modeling capabilities of the Transformer (a deep learning network architecture based on an attention mechanism), researchers introduced additional modalities to provide complementary information for autonomous driving. For example, Transfuse (a network model) utilizes the Transformer to fuse features from RGB (Red, Green, Blue) and LiDAR (Lidar) BEV (Bird's-Eye View) images to enhance global understanding of 3D (Three-Dimensional) scenes. InterFuse (Interpretable Sensor Fusion Transformer) improves model safety and interpretability by directly presenting intermediate features and constraining actions within a specified safety range. In recent years, modular end-to-end planning methods, in which all components are connected and jointly optimized toward the final goal, have gained increasing attention, further improving driving performance.
[0057] However, these end-to-end networks still rely on specific inputs such as target points or action commands to guide driving behavior, which greatly limits their interaction with humans and their application in real-world scenarios.
[0058] The emergence of large language models (LLMs) has revolutionized the field of autonomous driving. A growing number of researchers are incorporating them into autonomous driving systems, expanding their scope and capabilities. One key research direction is leveraging LLMs for driving-related visual question answering, providing enhanced explainability for driving decisions. Another research direction aims to leverage the powerful logical reasoning capabilities of LLMs to fundamentally shift the research paradigm for autonomous driving tasks. Specifically, GPT-Driver (Generative Pre-trained Transformer Driver), a generative large model for autonomous driving, inputs observations and self-states into the LLMs as language cues and generates planned trajectories and corresponding decision-making processes in natural language format. Agent-Driver constructs an agent driven by a large language model, leveraging a tool library, cognitive memory, and reasoning engine for human-like autonomous driving. Recently, LLMs have been successfully extended to the visual modality, enabling them to understand visual information for autonomous driving. Researchers have introduced multimodal large language models to process vectorized numerical modality data and generate answers to driving-related questions and action decisions. Based on the CARLA simulation platform (an open-source autonomous driving simulator), many studies have explored and performed closed-loop evaluations directly driven by LLMs. For example, LMDrive (Closed-Loop End-to-End Driving with Large Language Models) takes multimodal data and navigation instructions as input, enabling language-guided driving and enhancing interaction with humans.
[0059] Human-machine friendly language-guided autonomous driving has become a research hotspot. The vehicle's driving behavior can be driven by natural language instructions, which will significantly enhance the human-vehicle interaction experience and improve the intelligence level of autonomous driving.
[0060] Human-machine friendly, language-guided, end-to-end autonomous driving still faces some challenges, such as poor driving performance. The inventors have discovered that poor driving performance is due in part to insufficient visual comprehension (hereinafter referred to as visual misunderstanding) and in part to insufficient command comprehension (hereinafter referred to as command misunderstanding).
[0061] In addition, the human-machine friendly language-guided end-to-end autonomous driving model in related technologies has a high network complexity and the number of network parameters is as high as 7 billion. Such a large number of parameters brings huge deployment challenges.
[0062] In order to solve at least one of the above problems, the embodiments of the present disclosure provide a language-guided end-to-end autonomous driving model and an end-to-end autonomous driving method based on the end-to-end autonomous driving model.
[0063] like Figure 1 FIG. 1 is a flow chart of an end-to-end autonomous driving method provided by the present disclosure, including the following contents:
[0064] S101, extracting a first visual feature from visual data collected by the autonomous driving vehicle.
[0065] Visual data includes image data obtained by the surrounding environment continuously collected by image sensors with multiple perspectives equipped on autonomous vehicles.
[0066] During implementation, the collected visual data can be input into the visual encoder of the end-to-end autonomous driving model to generate a unified visual coding feature as the first visual feature.
[0067] S102: Prune the first visual features to obtain key visual features.
[0068] The end-to-end autonomous driving model provided by this disclosure introduces CCDP (Channel-wise Dilated Pruning) to address visual misinterpretation. CCDP uses pruning to preserve key features, reducing the number of parameters required for subsequent processing while also removing noise and mitigating visual misinterpretation.
[0069] Among them, pruning is a feature selection or feature dimensionality reduction method. Its core is to eliminate features that have low contribution to target recognition or decision-making tasks through specific rules or algorithms. In the embodiment of the present disclosure, a first visual feature is extracted from the visual data collected by the autonomous vehicle. This first visual feature may contain redundant or non-critical information. By pruning the first visual feature, irrelevant visual features can be effectively eliminated, thereby obtaining the key visual features that are critical to the vehicle's autonomous driving.
[0070] S103: align the key visual features to the language space to obtain a second visual feature corresponding to the key visual features.
[0071] In the autonomous driving scenario, combining visual information with navigation instructions based on natural language expression can make further driving decisions. Among them, key visual features are extracted from visual data, usually represented in the form of high-dimensional vectors. The key visual features extracted by pruning in the disclosed embodiment can identify key information that is conducive to autonomous driving in a high-dimensional space. Among them, the language space is a vector space built based on the language model to represent the semantics of text. The two are essentially different, so the key visual features can be aligned to the language space to obtain the second visual features corresponding to the key visual features, thereby realizing the collaborative work of visual information and navigation instructions, so as to utilize the powerful logical reasoning ability of the language model, comprehensively consider the semantics of visual information and navigation instructions, and be able to output an interpretable future trajectory of the autonomous driving vehicle.
[0072] S104: Process the second visual feature and the navigation instruction based on the language model to obtain a future trajectory of the autonomous driving vehicle.
[0073] S105, generating driving instructions for the autonomous driving vehicle based on the future trajectory.
[0074] It should be understood that the core capabilities of the above-mentioned end-to-end autonomous driving method are realized by the collaboration of different functional modules of the end-to-end autonomous driving model provided by the embodiments of the present disclosure, such as Figure 2 As shown:
[0075] During implementation, the autonomous vehicle continuously collects visual data in the form of visual data sequences, such as Figure 2 Medium X1, X i 、X T , are one frame of visual data respectively. Each frame of visual data is processed in the same way. Figure 1 The relevant process is explained using a frame of visual data as an example.
[0076] The visual encoder 201 in the end-to-end autonomous driving model serves as the entry point for visual data processing and is responsible for extracting primary visual features from the visual data. Through multiple neural network layers, it abstracts the features of the autonomous vehicle's visual data, converting the visual data input into primary visual features, which are high-dimensional semantic features.
[0077] Connector 202 in the end-to-end autonomous driving model serves as a bridge between the visual and language spaces. It prunes first visual features to further understand and reason about them. This pruning process removes redundant features that contribute little to autonomous driving decisions, preserving key visual features. By aligning key visual features to the language space, these features are converted into second visual features that the language model can understand, allowing them to be integrated with navigation instructions in the same semantic space. These second visual features include multiple second visual tokens.
[0078] The language model 203 of the end-to-end autonomous driving model receives the second visual features and navigation instructions, captures the correlation logic between the second visual features and the navigation instructions through a self-attention mechanism, and infers and generates the future trajectory of the autonomous driving vehicle. The control conversion module 204 of the end-to-end autonomous driving model then processes the future trajectory into driving instructions. The driving instructions may include, for example, steering and acceleration information.
[0079] In summary, in the disclosed embodiments, first visual features are extracted from visual data collected by an autonomous vehicle, transforming the raw, complex visual data into feature vectors in a high-dimensional semantic space to describe the environmental information surrounding the autonomous vehicle. Pruning the first visual features, by removing redundant and non-critical visual features, further focuses on visual features that are critical to autonomous driving decision-making. This reduces computational complexity while avoiding interference from irrelevant first visual features, improving recognition accuracy for key autonomous driving scenarios and reducing visual misinterpretation. Aligning key visual features to the language space allows the visually perceived environmental information to be combined with the user's driving intent and navigation instructions. Processing the second visual features and navigation instructions based on a language model generates an interpretable and understandable future trajectory, which is then used to generate driving instructions to control the autonomous vehicle. The language model understands the semantics of the navigation instructions and, in combination with factors such as the environmental conditions in the visual information, comprehensively determines the optimal autonomous driving control strategy, thereby improving the safety and reliability of the autonomous vehicle. Thus, the disclosed embodiments provide a language-guided, end-to-end autonomous driving method that reduces visual misinterpretation and improves the convenience of human-vehicle interaction.
[0080] Although large language models have demonstrated significant effectiveness in autonomous driving systems due to their advanced cognitive and logical reasoning capabilities, the inventors' research has found that driving performance is not linearly correlated with the number of model parameters.
[0081] Therefore, during implementation, the language model used in the end-to-end autonomous driving model in the embodiment of the present disclosure can be an LLM (Lightweight Language Model), so as to process the second visual features and navigation instructions based on the lightweight language model to obtain the future trajectory of the autonomous driving vehicle.
[0082] A lightweight language model is essentially a language processing model that can be designed through technical means such as architecture optimization, parameter compression, and computational acceleration. It is significantly lower than traditional large language models in terms of model scale, computing power consumption, and storage requirements, while maintaining core semantic understanding and generation capabilities. For example, the number of parameters of the lightweight language model used in the embodiment of the present disclosure can be reduced to less than 1B (Billion), so that the number of parameters of the entire end-to-end driving model remains at around 1B, that is, the gap with 1B is within the preset gap. In this way, reducing the number of model parameters from 7B to 1B is conducive to the deployment of end-to-end autonomous driving models.
[0083] In the disclosed embodiments, a lightweight language model is used to process the second visual features and navigation instructions to determine the future trajectory of the autonomous vehicle. This reduces the complexity of the end-to-end autonomous driving model by reducing the number of model parameters, while maintaining comparable autonomous driving performance. This facilitates model deployment, for example, by deploying the end-to-end autonomous driving model to the autonomous vehicle.
[0084] The above describes the core modules of the end-to-end autonomous driving model provided by the embodiments of the present disclosure. To facilitate understanding of how pruning is performed and key visual features are obtained in the embodiments of the present disclosure, a detailed description is provided below.
[0085] In the embodiment of the present disclosure, the implementation method of pruning the first visual feature to obtain the key visual feature can be as follows: Figure 3 Shown, including:
[0086] S301 : Predicting probabilities of multiple first visual word-grams in a first visual feature being key visual word-grams.
[0087] Based on the above, the first visual feature is generated by the visual encoder, which contains multiple first visual tokens. However, not all first visual tokens are equally valuable for autonomous driving decisions, so it is necessary to obtain key visual tokens through pruning. In the embodiment of the present disclosure, pruning is achieved through probabilistic prediction. Through probabilistic prediction, key visual tokens that play a key role in autonomous driving decisions are screened from the multiple first visual tokens of the first visual feature.
[0088] During implementation, the probability of predicting the first visual word-grams in the first visual feature as key visual word-grams can be implemented as follows: Figure 3Shown, including:
[0089] S3011: Extract a first global feature and a first local feature from the first visual feature.
[0090] The first global feature is an overall description of the vehicle's autonomous driving scene, focusing on macro information such as the overall environment layout of the autonomous driving scene. The first local feature focuses on the description of details in the visual scene and provides detailed information.
[0091] During implementation, extracting the first global feature and the first local feature from the first visual feature can be achieved based on the following steps:
[0092] Step A1, performing feature transformation on the first visual feature to obtain a first local feature of the first visual feature;
[0093] The feature transformation in this step is a computational method that maps original features to a new dimensional space. Its core purpose is to extract more valuable information from the original data, such as converting the first visual features into a high-level semantic space.
[0094] For example, the first visual feature is transformed to obtain the first local feature of the first visual feature, which can be described by formula (1):
[0095]
[0096] In formula (1), F i represents the first visual feature of the i-th frame visual data, and has a shape of N×C, where N is the number of first visual words in the first visual feature and C is the dimension of the first visual feature; MLP() in formula (1) is based on MLP (Multi-Layer Perceptron), which is a neural network module used to perform nonlinear transformation on the first visual feature, that is, to realize the feature transformation in step A1; L i Represents the first layout feature, with dimensions of This shows that MLP reduces the dimension of the first visual feature from C to To capture more abstract and advanced detail information in the first visual features.
[0097] Step A2: performing a pooling operation on the first local feature to obtain a first global feature of the first visual feature.
[0098] Pooling is a downsampling technique that is often used to reduce the spatial dimension of features while retaining the most important global information.
[0099] In one possible implementation, the first global feature can be described by formula (2):
[0100]
[0101] In formula (2), L i is the first local feature; Avg() is the average pooling operation, which is used to calculate the average value of each column feature dimension of the first local feature; G i is the first global feature, which is obtained by aggregating the first local features through the average pooling operation, and its dimension is
[0102] In some embodiments, a maximum pooling operation may be used to obtain the first global feature. Alternatively, all first local feature vectors may be processed based on a convolutional neural network, and then linear transformation and nonlinear activation may be performed through a fully connected layer to obtain the first global feature. The present disclosure does not limit the specific method for obtaining the first global feature. As long as an accurate first global feature can be obtained based on a relatively small number of parameters, it will be sufficient.
[0103] In the disclosed embodiment, a feature transformation is performed on the first visual feature to obtain a first local feature of the first visual feature, thereby enhancing the detail expression capability of the first visual feature. A pooling operation is performed on the first local feature to obtain a first global feature of the first visual feature, thereby extracting the global expression capability of the first visual feature. The pooling operation can effectively reduce the feature dimension and computational complexity, thereby providing efficient and reliable feature input for subsequent visual tasks while reducing model complexity.
[0104] S3012 : Predicting probabilities of the plurality of first visual word-grams being key visual word-grams based on the first global feature and the first local feature.
[0105] In summary, the disclosed embodiments predict the probabilities of multiple first visual tokens being key visual tokens based on the first global feature and the first local feature, and can screen out key visual tokens that are crucial for the autonomous driving task based on their probabilities. This helps focus on key information during subsequent processing, reduces visual misunderstandings, and improves autonomous driving performance.
[0106] In implementation, the prediction of the probability of the plurality of first visual word-grams being key visual word-grams may be achieved based on the following steps:
[0107] Step B1, splicing the first global feature and the first local feature to obtain a first spliced feature;
[0108] Step B2: performing feature transformation on the first concatenated features to obtain probabilities of the plurality of first visual word-grams being key visual word-grams.
[0109] For example, the probability of predicting multiple first visual words as key visual words can be described by formula (3):
[0110]
[0111] In formula (3), (;) represents a connection with a broadcast mechanism, which is used to splice the first global feature and the first local feature along the specified dimension to obtain the first spliced feature; MLP() represents a multi-layer perceptron, which is used to perform nonlinear feature transformation on the first spliced feature; Softmax() represents a normalization function, which is used to convert the output of MLP into a probability distribution; S i represents the probability of each of the first visual word-grams in the i-th frame of visual data being a key visual word-gram.
[0112] In the disclosed embodiment, the first global feature and the first local feature are spliced together through a connection with a broadcast mechanism, which automatically aligns the feature dimensions, allowing the global feature and the local feature to be spliced together along a specified dimension. The resulting first spliced feature can integrate global and detailed information, providing effective information for subsequent processing. After performing feature transformation on the spliced first spliced features, the features can be converted into a probabilistic form, which can more intuitively reflect the importance of each first visual word in describing the image content or semantics, thereby facilitating a reasonable prediction of the probability of each first visual word being a key visual word.
[0113] S302 : Based on the probabilities that the multiple first visual word-grams are key visual word-grams, the key visual word-grams are selected from the multiple first visual word-grams to obtain key visual features.
[0114] In summary, in the disclosed embodiments, by predicting the probabilities of multiple first visual words in the first visual features being key visual words, it is possible to clearly identify which first visual words are important for autonomous driving decisions. Based on the probabilities of multiple first visual words being key visual words, non-key first visual words can be removed, and key visual words can be filtered out from the multiple first visual words to obtain key visual features. In subsequent processing tasks, focused processing can be performed based on these key visual features, thereby reducing visual misunderstandings and improving autonomous driving performance.
[0115] In a possible implementation, the first visual word-gram with a probability greater than a threshold may be selected as the key visual word-gram.
[0116] In another possible implementation, in order to avoid the problem of setting the probability threshold and improve the flexibility of key visual word recognition, the key visual features can also be obtained based on the following steps:
[0117] Step C1, processing the probabilities of the plurality of first visual words as key visual words based on a continuous classification distribution relaxation method to obtain a binary mask;
[0118] Among them, the continuous classification distribution relaxation method can adopt Gumbel-Softmax distribution (also known as Concrete distribution).
[0119] Step C2: obtaining key visual word-grams from the plurality of first visual word-grams based on the binary mask to obtain key visual features.
[0120] During implementation, the key visual features can be described by formula (4):
[0121] M i =Gumbel-Softmax(S i ) *,1 ∈{0, 1} N (4)
[0122] In formula (4), S i represents the probability of the obtained multiple first visual word units as key visual word units; Gumbel-Softmax() represents the continuous classification distribution relaxation method, which is used to convert the continuous probability distribution into a discretized probability while maintaining the transferability of the gradient, so as to facilitate the optimization using the back propagation algorithm during the training process; M i The resulting binary mask is a binary vector of length N, where each element has a value of 0 or 1.
[0123] In the disclosed embodiments, a continuous classification distribution relaxation method is used to process the probabilities of multiple first visual words as key visual words and convert them into discrete masks in the form of 0 or 1. This avoids the blunt processing caused by directly using continuous probability values, improves the flexibility of identifying key visual words, and reduces visual misunderstandings. In addition, the continuous classification distribution relaxation method can also better optimize the end-to-end autonomous driving model during the model training phase, improving the overall performance of the end-to-end autonomous driving model for actual pruning.
[0124] In human driving scenarios, for example, drivers typically analyze the recent behavior of surrounding vehicles to infer their future movements or intentions and adjust their driving behavior accordingly. This mechanism can be understood as a temporal reasoning mechanism.
[0125] In autonomous driving scenarios, visual understanding is very important. In order to further reduce visual misunderstandings, the disclosed embodiment introduces a MEFA (Memory Enhanced Feature Aggregation) mechanism based on the temporal reasoning mechanism to optimize the visual features that are ultimately input to the language model. During implementation, in order to achieve autonomous driving that matches the driver's driving safety and reliability, the key visual features can be aligned to the language space through the MEFA mechanism to obtain the second visual features corresponding to the key visual features, thereby enhancing the temporal reasoning capability. The specific implementation method of this process is as follows: Figure 4 Shown, including:
[0126] S401 , based on first visual features of multiple frames of reference data within a specified neighborhood of the visual data, feature enhancement is performed on the first visual features of the visual data to obtain memory enhancement features of the visual data.
[0127] The autonomous driving vehicle can continuously collect data to obtain a visual data sequence, wherein the aforementioned visual data is any frame of visual data in the visual data sequence. The first visual features corresponding to each frame of visual data are extracted based on the same method.
[0128] To introduce a temporal reasoning mechanism, the MEFA mechanism in the disclosed embodiment may store first visual features of a preset number of frames of visual data (i.e., reference data within a specified neighborhood of the current frame) to enhance the visual features of the visual data of the current frame.
[0129] For example, assuming that a specified neighborhood is represented by a Z frame, the memory enhancement features of the visual data can be expressed by formulas (5) and (6):
[0130]
[0131] In formula (5), B i A memory bank representing the i-th frame visual data stores the first visual features of the current frame and the first visual features of the reference data of its adjacent Z frames. i-Z , F i-Z+1 ,...,F i-1 ] indicates that F i is the first visual feature of the i-th frame visual data (i.e., the visual data of the current frame), F i-Z to F i-1 It is the first visual feature of the Z frame in the specified neighborhood. The purpose of this memory bank is to capture the historical visual feature information in the time series and help the model understand the changes in the surrounding environment; Avg() represents the average pooling operation, which is used for the memory bank B i The first visual features of multiple frames of reference data in the specified neighborhood are averaged to obtain the visual feature time coding Its shape is 1×C, where C represents the dimension.
[0132] During implementation, the feature sequence of the first visual feature of the multi-frame reference data can also be modeled by the gating mechanism of LSTM (Long Short-Term Memory Network) to obtain the visual feature time coding output by LSTM.
[0133] After obtaining the visual feature temporal code, the first visual feature of the visual data of the current frame is enhanced by using the temporal code to obtain a memory enhancement feature of the visual data. For example, the feature enhancement operation can be described by formula (6):
[0134]
[0135] In formula (6), F i is the first visual feature of the visual data of the current frame; Represents temporal coding of visual features; Indicates that the first visual feature F i and splicing; MLP() represents a multi-layer perceptron, which is used to learn the complex relationship between the first visual feature and the average data of the historical first visual feature, and generate memory-enhanced features of the visual data; TE i Memory-enhanced features for representing visual data provide feature representations enhanced with temporal context.
[0136] S402: Based on the memory enhancement feature, align the key visual feature to the language space to obtain a second visual feature corresponding to the key visual feature.
[0137] In the disclosed embodiments, feature enhancement based on the first visual features of multi-frame reference data enables subsequent language models to better capture the dynamic changes in autonomous driving scenarios. Aligning key visual features to the language space based on memory-enhanced features enables subsequent language models to make more reasonable decisions based on semantic information, thereby supporting more natural human-computer interaction and improving autonomous driving performance.
[0138] In implementation, a query transformer can be used to achieve alignment from visual space to language space. The query transformer, also known as the Querying Transformer (Q-Former), is a lightweight Transformer architecture designed for vision-language alignment. It aims to bridge the modality gap between visual encoders and LLMs, enabling multimodal pre-trained models to achieve more accurate cross-modal understanding and generation.
[0139] In one possible implementation, a set of learnable query vectors can be used, which interact with the visual model and language model through self-attention layers and cross-attention layers to extract the most relevant visual representation of the text from key visual features and pass it to the language model.
[0140] In another possible implementation, in order to enhance the ability of the connector of the end-to-end autonomous driving model to focus on temporally significant moving objects, obtaining the second visual feature corresponding to the key visual feature may also be achieved based on the following steps:
[0141] Step D1, fusing the memory enhancement feature and the first visual feature to obtain a fused feature;
[0142] In step D2, the fused feature is input into the query transformer, and the key visual feature is input into the query transformer as a query value to obtain a second visual feature corresponding to the key visual feature output by the query transformer.
[0143] During implementation, the second visual feature corresponding to the key visual feature output by the query transformer is obtained, which can be described by formula (7):
[0144]
[0145] In formula (7), F i +TE i Represents the fusion feature, which is obtained by fusing the memory enhancement feature and the first visual feature; Represents the key visual features, which are used to perform query operations in the query transformer; Q-Former() represents the query transformer, which processes the input visual features with an attention mechanism and aggregates key information through key visual feature queries; The second visual feature corresponding to the key visual feature output by the query transformer.
[0146] In this disclosed embodiment, the memory-enhanced features are fused with the primary visual features to generate fused features. These features utilize both the primary visual features and the additional memory information, enabling subsequent language models to better understand and interpret the current visual content, leading to a more comprehensive understanding of the driving scene. The fused features are then fed into a query transformer, which interacts with key visual features to further optimize their extraction. This improves the model's ability to focus on temporally significant moving objects, thereby reducing visual misinterpretations.
[0147] In summary, the connector in the multimodal large language model is the first visual feature (in Figure 5The abcd sequence is used to perform pruning to retain the key visual features. By aligning the key visual features to the language space, the key visual features are converted into second visual features that can be understood by the language model. Figure 5 As shown, a method for implementing visual pruning and memory enhancement within the connector of an end-to-end autonomous driving model is provided. To reduce visual misunderstandings, the connector includes a dynamic pruning module (CCDP) 501 and a memory feature enhancement module (MEFA) 502:
[0148] In the dynamic pruning module 501, the probability of multiple first visual words in the first visual feature being a key visual word can be predicted, and then the probability of multiple first visual words being a key visual word can be processed based on the continuous classification distribution relaxation method to obtain a binary mask. Based on the binary mask, the key visual word can be obtained from the multiple first visual words to obtain the key visual feature.
[0149] In the memory feature enhancement module 502, a memory bank is introduced to capture historical visual feature data in a time series. The feature data of all frames in the memory bank are averaged through an average pooling operation (Avg) to obtain a visual feature temporal encoding. A perceptron (MLP) is then used to learn the complex relationship between the first visual feature and the visual feature temporal encoding, generating a memory-enhanced feature for the visual data. The memory-enhanced feature is fused with the first visual feature to obtain a fused feature. The fused feature is input into a query transformer, and the key visual feature is used as the query of the query transformer to align the key visual feature with the language space, thereby obtaining a second visual feature corresponding to the key visual feature.
[0150] During implementation, the visual data sequence is input into an end-to-end autonomous driving model, where the model's visual encoder and connector processes the data to obtain the second visual features of each frame. The second visual features of each frame in the visual data sequence, along with the navigation instructions, are then fed into a language model for processing. The language model then infers and generates a reasonable future trajectory based on the environmental information described by the multiple second visual features and the navigation instructions.
[0151] When language models process multimodal features, they need to interact with features from different modalities to extract cross-modal information. This cross-modal information can be used to understand navigation instructions and the environment described by the second visual feature.
[0152] In autonomous driving scenarios, position encoding is considered to be a key component for explicitly injecting position information into input data and enhancing the context modeling capabilities of language models. Its essence is to add position information to the input data of autonomous driving models through mathematical methods, so that the language model can understand the temporal order and contextual association of the input data, thereby more accurately processing the sequence information in autonomous driving scenarios and providing a basis for trajectory prediction and decision-making control.
[0153] However, in long-distance driving scenarios, positional encoding can lead to increasing distances between the second visual word in the visual data sequence and the first text word in the navigation instruction as historical visual data accumulates. The inventors have discovered that the feature differences introduced by positional encoding can negatively impact the attentional sensitivity of the second visual word in the visual features to the first text word in the navigation instruction, causing trajectory predictions to deviate from the specified navigation instruction.
[0154] Therefore, in order to reduce instruction misunderstandings, the attention module in the language model in the embodiment of the present disclosure adds DDIA (Distance-Decoupled Instruction Attention). The principle is as follows: by retaining the distance-related attention within the single-modal text or visual word, the cross-modal attention between the visual word and the text word is freed from the constraints of position encoding. On this basis, the second visual feature and navigation instructions can be processed based on the optimized language model to generate control information for controlling the autonomous vehicle. Its implementation is as follows: Figure 6 Shown, including:
[0155] S601, an attention module based on a language model, performs the following operations: determining a first attention feature between multiple first text words based on the position embedding of multiple first text words in the navigation instruction; determining a second attention feature between multiple second visual words based on the position embedding of multiple second visual words in the second visual feature; determining a third attention feature between the navigation instruction and the second visual feature based on the multiple first text words and the multiple second visual words.
[0156] During implementation, for each frame of visual data, the first attention feature, the second attention feature, and the third attention feature can be described based on formula (8):
[0157]
[0158] In formula (8), I represents the set of first text words corresponding to the navigation instruction; V represents the set of second visual words corresponding to the second visual feature; the set of first text words of the navigation instruction and the set of second visual words of the second visual feature of multiple frames of visual data in the visual data sequence constitute the matrix to be processed; represents the first attention feature, R() represents ROPE (Rotary Position Embedding, a position encoding technology) position embedding, q j is the jth word in the query matrix of the matrix to be processed (including the first text word and the second visual word), k i is the i-th word in the key matrix K of the matrix to be processed; represents the second attention feature, and {v|<j} represents the causal subset in the second visual word; represents the third attention feature, sim(q j , k i ) represents the similarity between the jth word in the query matrix Q and the ith word in the key matrix K, which is calculated as follows: Among them, qj T q j The transpose of .
[0159] In order to better understand the calculation of the first attention feature, the second attention feature and the third attention feature, it can be based on Figure 7 Elucidate formula (8):
[0160] like Figure 7 As shown, the vertical direction is the query matrix q j , the horizontal direction is the bond matrix k i For example, where q j The corresponding part includes the first text word of the navigation instruction and the second visual word of the second visual feature; similarly, k i The corresponding part also includes the first text word and the second visual word of the second visual feature. The processing of the second visual feature and the navigation instruction based on the optimized language model (i.e., the language model introduced with the DDIA mechanism) can be described based on the following three parts:
[0161] (1) In qj, the first text word belongs to the navigation instruction, k i In the case of the first text word element that also belongs to the navigation instruction, based on the position embedding of multiple first text words in the navigation instruction, the first attention feature between the multiple first text words is determined, that is, Figure 7 The first text word corresponding to the part of region 1 in the figure is calculated based on the attention within the unimodal text. Moreover, this part retains the self-attention within the instruction segment, allowing each text token to access sufficient context information. That is, Figure 7 As shown, in region 1, attention features are calculated between each first text word and all other text words.
[0162] (2) In q jThe second visual token belonging to the second visual feature, k i In the case of the second visual token that also belongs to the second visual feature, based on the position embedding of multiple second visual tokens in the second visual feature, determine the second attention feature between the multiple second visual tokens, that is Figure 7 The relevant attention calculation of the second visual token corresponding to the part of region 3 in the unimodal visual token. During implementation, as Figure 7 As shown in and Expression 8, for simplicity of calculation, each second visual token can calculate the attention feature with the second visual token before it, and the second visual token after it is represented by a mask.
[0163] (3) When q j The second visual token belonging to the second visual feature, k i The first text token belonging to the navigation instruction, and in the case of i < j, based on the multiple first text tokens and the multiple second visual tokens, determine the third attention feature between the navigation instruction and the second visual feature, that is Figure 7 The calculation based on (DDIA) between the multiple first text tokens and the multiple second visual tokens corresponding to the part of region 2 in. Among them, when calculating the third attention feature, the influence of the position encoding can be removed, and the weight is calculated only based on the semantic relevance of the content itself.
[0164] S602, the feature processing module based on the language model processes the first attention feature, the second attention feature and the third attention feature to obtain the future trajectory.
[0165] [[ID=2M]]That is, the feature processing module in the language model will further integrate the received different attention features to generate control information for representing the future trajectory of the predicted autonomous vehicle.
[0166] During implementation, the future trajectory can be converted into a lateral steering action and a longitudinal acceleration action through a PID controller, so as to determine the driving instruction of the autonomous vehicle. Taking the future trajectory as a dynamic reference can improve the overall performance of autonomous driving from dimensions such as trajectory tracking accuracy and environmental adaptability.
[0167] In the embodiment of the present disclosure, based on the position embedding of multiple first text words in the navigation instruction, the first attention feature between the multiple first text words is determined, which can capture the logical dependency within the navigation instruction and ensure the integrity and order rationality of the navigation instruction semantics. Based on the position embedding of multiple second visual words in the second visual feature, the second attention feature between the multiple second visual words is determined, which second visual features are most important to the current decision. By determining the third attention feature between the navigation instruction and the second visual feature based on multiple first text words and multiple second visual words, the association between text information and visual information can be determined. The above-mentioned calculation method of retaining the distance-related attention within the monomodal text or visual word, while freeing the cross-modal attention between the visual word and the text word from the constraints of position encoding, can avoid the feature differences introduced by position encoding, which will have a negative impact on the attention sensitivity of the visual word in the visual feature to the text word in the navigation instruction, help improve the understanding accuracy of the navigation instruction, reduce instruction misunderstandings, and further improve the performance of autonomous driving.
[0168] In summary, the end-to-end autonomous driving method provided in the embodiments of this disclosure addresses driving failures caused by insufficient visual understanding by enhancing visual features through visual pruning to retain key visual features, thereby reducing visual misunderstandings. By optimizing the internal structure of the language model and decoupling the positional encoding of navigation commands and visual features (DDIA), command misunderstandings can be reduced. The use of a lightweight language model simultaneously reduces computational overhead and improves driving performance, facilitating practical deployment.
[0169] Based on the same technical concept, the disclosed embodiments provide an end-to-end autonomous driving model training method. The data processing process of the end-to-end autonomous driving model is consistent with the previous end-to-end autonomous driving method, with the difference being that the model parameters of the end-to-end autonomous driving model here require optimization.
[0170] like Figure 8 FIG. 1 is a flow chart of an end-to-end autonomous driving model training method provided by an embodiment of the present disclosure, including the following contents:
[0171] S801, a visual encoder based on an end-to-end autonomous driving model extracts a third visual feature from the sample vision.
[0172] During implementation, the visual encoder of the end-to-end autonomous driving model performs feature abstraction on the sample visual image data, converts the pixel-level visual data input into high-dimensional semantic features, and obtains the third visual feature.
[0173] S802, pruning the third visual feature based on the connector of the end-to-end autonomous driving model to obtain a core visual feature; aligning the core visual feature to the language space to obtain a fourth visual feature corresponding to the core visual feature.
[0174] Specifically, the third visual feature is pruned through the connector of the end-to-end autonomous driving model to reduce the dimensionality of the visual features and remove non-critical visual features to obtain the core visual features. This core visual feature is then aligned to the language space through a nonlinear transformation, allowing the visual features to interact with navigation instructions in the language space.
[0175] S803 , processing the fourth visual feature and the instruction sample based on the language model of the end-to-end autonomous driving model to obtain an expected trajectory of the autonomous driving vehicle.
[0176] The command samples are samples of navigation commands generated based on natural language, and are used to implement a human-vehicle interaction mode based on language guidance.
[0177] Among them, the expected trajectory is the predicted trajectory that the autonomous driving vehicle should travel in the future.
[0178] S804: Optimize the connector and the language model based on the first loss between the expected trajectory and the corresponding data ground truth.
[0179] In other words, the visual encoder of the end-to-end autonomous driving model can use an already mature visual encoder and freeze it during the model training phase. During the model training phase, the model parameters of the language model and connector need to be optimized.
[0180] Among them, the first loss required to optimize the model parameters is the difference between the expected trajectory output by the model and the true value of the trajectory. This difference is used to ensure that the expected trajectory output by the model meets the actual driving requirements.
[0181] In the disclosed embodiment, a third visual feature is extracted from the sample vision by a visual encoder, which can preliminarily convert a complex visual scene into a vector form with certain semantics and feature expressions, and can provide basic visual information for subsequent analysis and decision-making. Pruning the third visual feature with a connector to obtain a core visual feature helps to remove redundant information and extract visual features that are critical to the autonomous driving task, thereby reducing visual misunderstandings. Aligning the core visual feature to the language space to obtain the fourth visual feature achieves the fusion of visual information and language information in the same semantic space, enabling the model to combine the visual scene with the language instruction and better understand the relationship between the two. The language model based on the end-to-end autonomous driving model processes the fourth visual feature and the instruction sample, so that the model can comprehensively judge and output the predicted expected trajectory of the autonomous driving vehicle based on the current visual scene and language instruction. The first loss directly reflects the gap between the expected trajectory output by the model and the true value of the data. Through optimization, this gap can be continuously narrowed, the accuracy and reliability of the model in generating the expected trajectory can be improved, and the driving performance of the model can be improved.
[0182] As described above, the language model of the end-to-end autonomous driving model in the embodiment of the present disclosure is a lightweight language model.
[0183] A lightweight language model is a deep learning model whose number of parameters does not exceed a specific scale and has strong generalization and complex semantic understanding capabilities.
[0184] In the disclosed embodiment, the language model of the end-to-end autonomous driving model is a lightweight language model, which can facilitate model deployment by reducing model parameters while obtaining comparable autonomous driving performance.
[0185] As described above, in the embodiment of the present disclosure, the CCDP mechanism is introduced to prune the third visual feature to obtain the core visual feature, such as Figure 9 As shown, it can be implemented as:
[0186] S901 , predicting probabilities of multiple third visual word-grams in the third visual feature being core visual word-grams.
[0187] The multiple third visual words in the third visual feature do not all have equal value for autonomous driving decisions. Through probabilistic prediction, key visual words that play a key role in autonomous driving decisions are screened out from the multiple third visual words in the third visual feature.
[0188] During implementation, the probability of predicting multiple third visual words in the third visual feature as core visual words can be as follows: Figure 9 As shown, the implementation is:
[0189] S9011, extracting a second global feature and a second local feature from the third visual feature.
[0190] The second global feature is an overall description of the vehicle's autonomous driving scene, covering macro information such as the overall environmental layout of the autonomous driving scene. The second local feature is a description of key details in the visual scene.
[0191] During implementation, extracting the second global feature and the second local feature from the third visual feature can be achieved based on the following steps:
[0192] Step E1, performing feature transformation on the third visual feature to obtain a second local feature;
[0193] During implementation, the third visual feature can be transformed by a multi-layer perceptron to obtain the second local feature.
[0194] Step E2: performing a pooling operation on the second local feature to obtain a second global feature.
[0195] During implementation, the average pooling operation can be used to calculate the average value of each column feature dimension of the second local feature to obtain the second global feature.
[0196] In the disclosed embodiment, a feature transformation is performed on the third visual feature to obtain a second local feature of the third visual feature, and the representation capability of the third visual feature can be enhanced through nonlinear transformation. A pooling operation is performed on the second local feature to obtain a second global feature of the third visual feature, and the second local feature can be smoothly aggregated to obtain global cognitive knowledge of the surrounding environment. The pooling operation can also effectively reduce feature dimensionality and computational complexity, thereby providing efficient and reliable feature input for subsequent visual tasks.
[0197] S9012: Predict probabilities of the plurality of third visual word-grams being core visual word-grams based on the second global feature and the second local feature.
[0198] In the disclosed embodiment, the probability of multiple third visual words being core visual words is predicted based on the second global feature and the second local feature, and the core visual words that have an impact on the autonomous driving task can be screened out according to the level of probability, thereby reducing visual misunderstanding.
[0199] In implementation, the probability of predicting multiple third visual word-grams as core visual word-grams can be achieved based on the following steps:
[0200] Step F1, splicing the second global feature and the second local feature to obtain a second spliced feature;
[0201] Step F2: performing feature transformation on the second concatenated features to obtain probabilities of the plurality of third visual word-grams being core visual word-grams.
[0202] During implementation, the second global feature and the second local feature can be concatenated using a connection with a broadcast mechanism to obtain a second concatenated feature. Subsequently, the second concatenated feature, which has undergone nonlinear feature transformation using a multilayer perceptron, is processed based on a normalization function to obtain the probability that multiple third visual words are core visual words. The specific processing flow is the same as that of the previous formula (3) and will not be repeated here.
[0203] In the disclosed embodiment, the second global feature and the second local feature are spliced together to automatically align the feature dimensions, allowing the global and local features to be spliced together along a specified dimension. Transforming the second spliced feature to obtain the probabilities of multiple third visual terms being core visual terms can more intuitively reflect the importance of each third visual term in describing image content or semantics, thereby reducing visual misunderstandings.
[0204] S902 : Based on the probabilities that the plurality of third visual words are core visual words, core visual words are selected from the plurality of third visual words to obtain core visual features.
[0205] In the disclosed embodiments, by predicting the probabilities of multiple third visual words within the third visual feature being key visual words, it is possible to clearly identify which third visual words are important for autonomous driving decision-making. Based on the probabilities of multiple third visual words being key visual words, non-key third visual words can be removed, and key visual words can be filtered out from the multiple third visual words to obtain key visual features. In subsequent processing tasks, focused processing can be performed based on these key visual features, reducing visual misunderstandings.
[0206] During implementation, the core visual features can be obtained based on the following steps:
[0207] Step G1, processing the probabilities of multiple third visual word-grams as core visual word-grams based on a continuous classification distribution relaxation method to obtain mask information;
[0208] Step G2: obtaining a core visual word-gram from multiple third visual word-grams based on the mask information to obtain a core visual feature.
[0209] Exemplarily, the expressions of step G1 and step G2 are the same as the above formula (4), where S i It represents the probability of the obtained multiple third visual word-grams being the core visual word-grams, and the rest will not be described here.
[0210] In the disclosed embodiment, the probability of multiple third visual words being key visual words is processed based on the continuous classification distribution relaxation method, which can better perform gradient optimization on the model, avoid the rigid processing caused by directly using continuous probability values, and can be easily combined with gradients, which is beneficial to optimizing the model parameters of the end-to-end autonomous driving model.
[0211] In order to further reduce visual misunderstandings, the MEFA mechanism is introduced in the embodiment of the present disclosure to align the core visual features to the language space and obtain the fourth visual features corresponding to the core visual features, such as Figure 10 Shown, including:
[0212] S1001, based on the third visual features of multiple frames of sample data within a preset neighborhood of the sample vision, feature enhancement is performed on the third visual features of the sample vision to obtain a time coding feature of the sample vision.
[0213] The autonomous driving vehicle can continuously collect data to obtain a sample data sequence, and the aforementioned sample vision is the data of any frame of sample vision in the sample data sequence. Each frame of sample vision is extracted based on the same method to obtain the corresponding third visual feature.
[0214] To introduce a temporal reasoning mechanism, the MEFA mechanism in the disclosed embodiment can store third visual features of a preset number of frames of sample data (i.e., visual data within a preset range of the current frame's sample data as training samples) to enhance the visual features of the current frame's sample data.
[0215] For example, assuming that the preset neighborhood is represented by Z frames, the temporal coding features of the sample vision (ie, the memory enhancement features of the sample vision) can be expressed with reference to the aforementioned formulas (5) and (6).
[0216] In the above formula (5), B i represents a memory bank that stores the third visual feature of the sample visual of the current frame and the third visual feature of the sample data of its adjacent Z frames. For example, the memory bank can be obtained by [F i-Z , F i-Z+1 ,...,F i-1 ] indicates that F i is the third visual feature of the i-th frame sample vision (i.e. the current frame sample vision), F i-Z to F i-1 It is the third visual feature of the Z frame in the preset neighborhood. The purpose of this memory bank is to capture the historical visual feature information in the time series to help the model understand the surrounding environment; Avg() represents the average pooling operation, which is used for the memory bank B i The third visual feature of multiple frames of sample data in the preset neighborhood is averaged to obtain the visual feature time coding Its shape is 1×C, where C represents the dimension.
[0217] During implementation, the feature sequence of the third visual feature of multiple frames of sample data can also be temporally modeled through the LSTM (Long Short-Term Memory Network) gating mechanism to obtain the visual feature time coding.
[0218] After obtaining the temporal encoding of the visual features, the third visual feature of the sample visual image of the current frame is used to perform feature enhancement to obtain a memory enhancement feature of the visual data. Exemplarily, this feature enhancement operation can be performed using formula (6). In formula (6), the first visual feature is replaced by the third visual feature.
[0219] S1002: Based on the time coding feature, align the core visual feature to the language space to obtain a fourth visual feature corresponding to the core visual feature.
[0220] In the disclosed embodiment, time coding is introduced, the third visual feature is enhanced, and then the core visual feature is aligned to the language space. This allows the subsequent language model to make more reasonable decisions based on semantic information, thereby supporting more natural human-computer interaction and improving the performance of autonomous driving.
[0221] In implementation, a query transformer can be used to project the visual space into the language space. Specifically, it can be implemented as follows:
[0222] Step H1, fusing the temporal coding feature and the third visual feature to obtain a joint feature;
[0223] In step H2, the joint feature is input into the query transformer of the connector, and the core visual feature is input into the query transformer as a query value to obtain a fourth visual feature corresponding to the core visual feature output by the query transformer.
[0224] In this disclosed embodiment, the temporal encoding feature and the third visual feature are fused to generate a joint feature, enabling the multimodal language model to better understand and interpret the current visual content. The joint feature is then fed into a query transformer, which interacts with the core visual features to further optimize their extraction. This helps the model focus on objects that significantly move over time, thereby reducing visual misinterpretation.
[0225] In the embodiment of the present disclosure, based on the content described above, in order to reduce command misunderstanding, the language model based on the end-to-end autonomous driving model processes the fourth visual feature and the command sample to obtain the expected trajectory of the autonomous driving vehicle, such as Figure 11 As shown, it can be implemented as:
[0226] S1101, the attention module based on the language model performs the following operations: based on the position embedding of multiple second text words in the instruction sample, determining the fourth attention feature between the multiple second text words; based on the position embedding of multiple fourth visual words in the fourth visual feature, determining the fifth attention feature between the multiple fourth visual words; based on the multiple second text words and the multiple fourth visual words, determining the sixth attention feature between the instruction sample and the fourth visual feature.
[0227] This processing step is the same as the implementation of the end-to-end autonomous driving control method described above and will not be repeated here. Here, the first text word in formula (8) is replaced by the second text word, and the second visual word in formula (8) is replaced by the third visual word.
[0228] S1102: Process the fourth attention feature, the fifth attention feature, and the sixth attention feature by a feature processing module based on the language model to obtain an expected trajectory.
[0229] That is, the feature processing module in the language model will further integrate the different attention features received to predict the expected trajectory of the autonomous driving vehicle.
[0230] In the disclosed embodiment, a calculation method that retains distance-related attention within a single-modal text or visual word, while freeing the cross-modal attention between visual words and text words from the constraints of position encoding, can avoid feature differences introduced by position encoding, which would negatively affect the attention sensitivity of visual words in visual features to text words in navigation instructions, help improve the accuracy of understanding navigation instructions, reduce instruction misunderstandings, and further enhance autonomous driving performance.
[0231] Furthermore, in the disclosed embodiments, the end-to-end autonomous driving model also includes a control conversion module, which is used to process the expected trajectory and obtain control instructions for the autonomous vehicle to travel along the expected trajectory. The model parameters of the control conversion module are optimized. That is, during implementation, the model parameters of the control conversion model need to be optimized based on the calculated losses (such as the first loss in the disclosed embodiments and the second loss described below).
[0232] The expected trajectory generated in the disclosed embodiments represents the future trajectory of the autonomous vehicle, enabling the vehicle to plan its path in advance and automatically control the autonomous vehicle based on the planned path. The control conversion model can convert the generated expected trajectory into control instructions for controlling the autonomous vehicle, i.e., driving instructions such as left turn and acceleration. By optimizing the model parameters of the control conversion model, the control conversion model can better understand the expected trajectory and convert it into control instructions that the autonomous vehicle can understand.
[0233] In the connector, when the CCDP mechanism and the MEFA mechanism are introduced, in order to improve the connector's ability to extract core visual features, an auxiliary trainer is also introduced in the training stage in the embodiment of the present disclosure. The auxiliary trainer is only used for model training and can be pruned in the inference stage. During implementation, the fourth visual feature is reconstructed based on the auxiliary trainer to obtain the fifth visual feature. Thus, feature reconstruction is performed based on the auxiliary trainer, and the reconstruction loss (second loss) can be used to optimize the model parameters of the pruned part, thereby improving the model's ability to understand features. It can be implemented as follows: based on the first loss between the expected trajectory and the corresponding data true value, and the second loss between the fifth visual feature and the third visual feature (i.e., the reconstruction loss), the connector, the language model, and the auxiliary trainer are optimized.
[0234] In this disclosed embodiment, the fourth visual feature generated through Memory Enhanced Feature Aggregation (MEFA) is able to ideally encapsulate all visual features of the current frame. To further enhance the connector's ability to select key visual tokens and aggregate complete information, an auxiliary trainer is introduced. This allows the visual features ultimately output by the trained connector to further enhance the connector's ability to perform visual pruning and mine key visual information, thereby reducing visual misinterpretation.
[0235] When implementing, if Figure 12 As shown, feature reconstruction can be implemented as:
[0236] S1201, determining an input feature to be reconstructed based on the learnable embedding and the fourth visual feature; wherein the learnable embedding is used to restore feature information in the third visual feature that is not retained in the fourth visual feature.
[0237] When implementing, you can Figure 5 The auxiliary trainer part shown in the figure determines the input features to be reconstructed based on the following steps:
[0238] Step K1, performing an inversion operation on the mask information used to extract the core visual features to obtain a filtering mask;
[0239] Step K2, constructing potential unknown features based on learnable embedding and filter masks;
[0240] Step K3, constructing a known feature based on the difference feature between the fourth visual feature and the third visual feature, and the mask information;
[0241] Step K4: Fuse the unknown features and the difference features to obtain the input features to be reconstructed.
[0242] For example, the input features to be reconstructed can be described by formula (9):
[0243]
[0244] In formula (9), M i Represents mask information, which is used to represent core visual features; (1-M i ) represents the filter mask, which is obtained by inverting the mask information used to extract the core visual features; e is a learnable embedding, the initial value is the default value such as 0, and it can be updated by gradient backpropagation; e.(1-M i ) is to construct potential unknown features; Indicates that a feature transformation is performed on the fourth visual feature based on a multi-layer perceptron, and then a difference feature between the transformed fourth visual feature and the third visual feature is determined; Indicates the construction of known features; Represents the input features to be reconstructed.
[0245] In the disclosed embodiment, the mask information used when extracting the core visual features is inverted to obtain a filter mask, which is used to restore the visual features that were pruned during the core feature extraction process so that they can be reconstructed and enhanced in subsequent steps. Unknown features represent feature information that is not retained in the third visual feature. Learnable embedding helps the model extract useful feature information from these pruned feature information and converts it into a form that can be used by the model. Known features represent visual feature information that the model has been able to accurately capture and express, providing a basis for feature reconstruction.
[0246] S1202: Process the input feature to be reconstructed based on multiple Transformer blocks to obtain a fifth visual feature.
[0247] In this disclosed embodiment, the fourth visual feature is combined with a learnable embedding, which is used to recover this lost feature information, ensuring more complete visual data for subsequent processing. The Transformer block transforms and enhances the input features through a self-attention mechanism and a feedforward neural network, reconstructing a more complete, accurate, and semantically rich fifth visual feature, thereby improving the model's ability to express visual features.
[0248] In addition, in the embodiments of the present disclosure, not only can image sensors be used to obtain visual data, but visual data can also be obtained based on other visual sensors such as LiDAR. During implementation, radar data can use a radar encoder to obtain the sixth visual feature of the radar data, and convert the sixth visual feature into a language space through a connector of the radar data to obtain the seventh visual data. The seventh visual data is spliced with the second visual data (or the fourth visual data) to obtain visual fusion data, which participates in the subsequent processing flow as the new second visual data or the fourth visual data.
[0249] Of course, in another embodiment, the radar data can be mapped to the visual space of the image's visual encoder after being processed by its corresponding connector to obtain radar visual features. After the radar visual features and the first visual features (or third visual features) are fused (such as spliced), they participate in the subsequent processing flow as new first visual features (or third visual features).
[0250] Taking the image sensor to obtain visual data as an example, the end-to-end autonomous driving model provided by the embodiment of the present disclosure is as follows: Figure 13 As shown. The navigation instruction is processed by the Tokenizer (word segmenter) 1301 to obtain the first text word element of the instruction sample. The visual data sequence collected by the autonomous driving vehicle is input into the visual encoder 1302 to obtain the first visual feature of each frame of visual data. In order to reduce visual misunderstanding, the first visual feature in the embodiment of the present disclosure is processed by the connector 1303 to obtain the second visual feature of each frame of visual data. For the first visual feature of each frame of visual data, the connector 1303 introduces the CCDP mechanism to implement visual token pruning, thereby screening out key visual features. The first visual feature is processed by the MEFA mechanism, and time coding is introduced to enhance the first visual feature to obtain a memory enhancement feature, and then the second visual feature is projected into the language space based on the memory enhancement feature. The query transformer in the connector (not shown in the figure) takes the key visual feature as a query and introduces a memory enhancement feature, which helps to enhance the connector's ability to focus on objects that move significantly in time.
[0251] By introducing an auxiliary trainer during the model training phase, the first visual features extracted from the visual data can be reconstructed, thereby strengthening connector 1303's ability to mine key visual features and aggregate complete visual information. In summary, the explicit token reconstruction operation of the auxiliary trainer further strengthens the connector's ability to retain key visual features and increases the completeness of visual information.
[0252] In order to further reduce the number of model parameters, in the embodiment of the present disclosure, the language model 1304 adopts a lightweight language model.
[0253] In order to reduce instruction misunderstandings, the DDIA mechanism is introduced in the embodiment of the present disclosure, so that the language model 1304 can reduce misunderstandings in long-distance driving scenarios caused by the introduction of position encoding and improve the ability to understand navigation instructions.
[0254] Figure 13 In the figure, the modules with fire-shaped logos are modules whose model parameters need to be optimized during the training phase, and the modules with ice flower logos are modules that need to be frozen during the model training phase.
[0255] Compared with the language-guided autonomous driving model in the related art, such as LMDrive using a large language model, after simulation experiments on the same simulation platform, the end-to-end autonomous driving model provided by the embodiment of the present disclosure achieves better driving performance than LMDrive. Figure 14 shown.
[0256] Figure 14 In the LMDrive model, the number of parameters is about 7B (e.g. Figure 14 (a) in the figure), the number of parameters of the end-to-end autonomous driving model provided by the embodiment of the present disclosure is about 1B (e.g. Figure 14 The inventors found that the driving performance and the parameters of the model are not linearly related. They also analyzed the reasons for driving failure, such as Figure 14 As shown in part (c), driving failures are mostly caused by visual misunderstanding and instruction misunderstanding.
[0257] Three scenarios were simulated for closed-loop runtime: LangAutio-tiny (minimal closed-loop runtime), LangAutio-short (medium closed-loop runtime), and LangAutio (maximum closed-loop runtime). Driving performance was analyzed in each of these three scenarios, and visual and command misunderstandings were found to be key factors in driving failure.
[0258] On this basis, the present disclosure provides the following embodiments: Figure 13 The end-to-end autonomous driving model shown in Figure 1 is shown in Figure 2. Figure 14 Part (d) of Figure 3 compares the driving performance of this model with that of LMDrive and LMDrive-liteLM. The model performs better in all three scenarios, with improved driving accuracy in each scenario.
[0259] In summary, the disclosed embodiments demonstrate that insufficient visual understanding is a primary bottleneck in voice-guided driving, and that performance is not linearly correlated with language model size. Building on this research, the disclosed embodiments propose a new paradigm: a lightweight language model enhanced through visual augmentation strategies. This approach simultaneously reduces computational overhead and improves driving performance, facilitating practical deployment.
[0260] The end-to-end autonomous driving model proposed in the embodiments of the present disclosure is called VLDrive (Visual Feature Enhanced Lightweight Language Guided End-to-End Autonomous Driving Model). This model is a lightweight language-guided driving architecture enhanced by vision-centric strategies. The framework includes CCDP for adaptive visual signal extraction and MEFA for temporal information integration. In addition, DDIA promotes vision-language alignment and enables robust navigation command following. These complementary strategies together achieve more reliable autonomous driving.
[0261] The VLDrive proposed in the present disclosure was tested in a closed-loop simulation on the CARLA platform using a standard language-guided driving benchmark. VLDrive achieved state-of-the-art performance with a significantly reduced number of parameters.
[0262] Based on the same technical concept, the embodiment of the present disclosure also provides an end-to-end autonomous driving device 1500, such as Figure 15 Shown, including:
[0263] A first extraction unit 1501 is configured to extract a first visual feature from visual data collected by the autonomous driving vehicle;
[0264] A pruning unit 1502 is used to prune the first visual feature to obtain a key visual feature;
[0265] An alignment unit 1503 is configured to align the key visual feature to the language space to obtain a second visual feature corresponding to the key visual feature;
[0266] A prediction unit 1504 is configured to process the second visual feature and the navigation instruction based on the language model to obtain a future trajectory of the autonomous driving vehicle;
[0267] The control unit 1505 is configured to generate driving instructions for the autonomous vehicle based on the future trajectory.
[0268] In some embodiments, the pruning unit includes:
[0269] a prediction subunit, configured to predict probabilities of the plurality of first visual words in the first visual feature being key visual words;
[0270] The extraction subunit is configured to select the key visual word-grams from the plurality of first visual word-grams based on the probabilities that the plurality of first visual word-grams are the key visual word-grams, and obtain the key visual features.
[0271] In some embodiments, the prediction subunit is specifically configured to:
[0272] Extracting a first global feature and a first local feature from the first visual feature;
[0273] The probabilities of the plurality of first visual word-grams being key visual word-grams are predicted based on the first global feature and the first local feature.
[0274] In some embodiments, the prediction subunit is specifically configured to:
[0275] Performing feature transformation on the first visual feature to obtain a first local feature of the first visual feature;
[0276] A pooling operation is performed on the first local feature to obtain a first global feature of the first visual feature.
[0277] In some embodiments, the prediction subunit is specifically configured to:
[0278] Splicing the first global feature and the first local feature to obtain a first spliced feature;
[0279] Perform feature transformation on the first concatenated features to obtain probabilities of the plurality of first visual word-grams being key visual word-grams.
[0280] In some embodiments, the extraction subunit includes:
[0281] a mask acquisition subunit, configured to process the probabilities of the plurality of first visual word-grams being key visual word-grams based on a continuous classification distribution relaxation method to obtain a binary mask;
[0282] The acquisition subunit is configured to acquire key visual word-grams from the plurality of first visual word-grams based on the binary mask to obtain key visual features.
[0283] In some embodiments, the alignment unit comprises:
[0284] an enhancement subunit, configured to perform feature enhancement on the first visual feature of the visual data based on the first visual feature of multiple frames of reference data within a specified neighborhood of the visual data, to obtain a memory enhancement feature of the visual data;
[0285] The projection subunit is used to align the key visual features to the language space based on the memory enhancement features to obtain the second visual features corresponding to the key visual features.
[0286] In some embodiments, the projection subunit is specifically configured to:
[0287] Fusing the memory enhancement feature and the first visual feature to obtain a fused feature;
[0288] The fused feature is input into the query transformer, and the key visual feature is input into the query transformer as a query value to obtain a second visual feature corresponding to the key visual feature output by the query transformer.
[0289] In some embodiments, the prediction unit is specifically configured to process the second visual feature and the navigation instruction based on a lightweight language model to obtain a future trajectory for the autonomous driving vehicle.
[0290] In some embodiments, the prediction unit includes:
[0291] The first processing subunit is configured to perform the following operations based on the attention module of the language model: determining a first attention feature between the plurality of first text words based on position embeddings of the plurality of first text words in the navigation instruction; determining a second attention feature between the plurality of second visual words based on position embeddings of the plurality of second visual words in the second visual feature; and determining a third attention feature between the navigation instruction and the second visual feature based on the plurality of first text words and the plurality of second visual words;
[0292] The second processing subunit is used to process the first attention feature, the second attention feature and the third attention feature based on the feature processing module of the language model to obtain a future trajectory.
[0293] Based on the same technical concept, the embodiment of the present disclosure also provides an end-to-end autonomous driving model training device 1600, such as Figure 16 Shown, including:
[0294] A second extraction unit 1601 is configured to extract a third visual feature from the sample vision based on a visual encoder of an end-to-end autonomous driving model;
[0295] The sparse unit 1602 is configured to prune the third visual feature based on the connector of the end-to-end autonomous driving model to obtain a core visual feature; align the core visual feature to the language space to obtain a fourth visual feature corresponding to the core visual feature;
[0296] an estimation unit 1603 for processing the fourth visual feature and the instruction sample based on a language model of the end-to-end autonomous driving model to obtain an expected trajectory of the autonomous driving vehicle;
[0297] The optimization unit 1604 is configured to optimize the connector and the language model based on a first loss between the expected trajectory and the corresponding data ground truth.
[0298] In some embodiments, the sparse unit comprises:
[0299] a probability determination subunit, configured to predict probabilities of the plurality of third visual word-grams in the third visual feature being core visual word-grams;
[0300] The screening subunit is configured to screen out the core visual word from the plurality of third visual word-grams based on the probability that the plurality of third visual word-grams are the core visual word-grams, and obtain the core visual feature.
[0301] In some embodiments, the probability determination subunit is specifically configured to:
[0302] extracting a second global feature and a second local feature from the third visual feature;
[0303] The probabilities of the plurality of third visual word-grams being the core visual word-grams are predicted based on the second global feature and the second local feature.
[0304] In some embodiments, the probability determination subunit is specifically configured to:
[0305] Performing feature transformation on the third visual feature to obtain a second local feature;
[0306] A pooling operation is performed on the second local feature to obtain the second global feature.
[0307] In some embodiments, the probability determination subunit is specifically configured to:
[0308] Splicing the second global feature and the second local feature to obtain a second splicing feature;
[0309] Feature transformation is performed on the second concatenated features to obtain probabilities of the plurality of third visual word-grams being core visual word-grams.
[0310] In some embodiments, the screening subunit is specifically configured to:
[0311] The probability of multiple third visual words being core visual words is processed based on the continuous classification distribution relaxation method to obtain mask information;
[0312] A core visual word-gram is obtained from the plurality of third visual word-grams based on the mask information to obtain a core visual feature.
[0313] In some embodiments, the sparse unit comprises:
[0314] an enhancement subunit, configured to enhance the third visual feature of the sample vision based on the third visual feature of multiple frames of sample data within a preset neighborhood of the sample vision, so as to obtain a temporal coding feature of the sample vision;
[0315] The conversion subunit is used to align the core visual feature to the language space based on the time coding feature to obtain the fourth visual feature corresponding to the core visual feature.
[0316] In some embodiments, the conversion subunit is specifically configured to:
[0317] Fuse the time coding feature and the third visual feature to obtain the joint feature;
[0318] The joint feature is input into the query transformer of the connector, and the core visual feature is input into the query transformer as a query value to obtain a fourth visual feature corresponding to the core visual feature output by the query transformer.
[0319] In some embodiments, the language model is a lightweight language model.
[0320] In some embodiments, the estimation unit includes:
[0321] a third processing unit configured to perform the following operations based on the attention module of the language model: determining a fourth attention feature between the plurality of second text tokens based on position embeddings of the plurality of second text tokens in the instruction sample; determining a fifth attention feature between the plurality of fourth visual tokens based on position embeddings of the plurality of fourth visual tokens in the fourth visual feature; and determining a sixth attention feature between the instruction sample and the fourth visual feature based on the plurality of second text tokens and the plurality of fourth visual tokens;
[0322] The fourth processing unit is configured to process the fourth attention feature, the fifth attention feature, and the sixth attention feature based on the feature processing module of the language model to obtain an expected trajectory.
[0323] In some embodiments, the end-to-end autonomous driving model further includes a control conversion module, the control conversion module being configured to process the expected trajectory and obtain control instructions for the autonomous driving vehicle to travel along the expected trajectory;
[0324] Among them, the model parameters of the control conversion module are involved in the optimization.
[0325] In some embodiments, further comprising:
[0326] a reconstruction unit, configured to reconstruct the fourth visual feature based on the auxiliary training device to obtain a fifth visual feature;
[0327] The optimization unit is specifically used to optimize the connector, the language model, and the auxiliary trainer based on a first loss between the expected trajectory and the corresponding data true value, and a second loss between the fifth visual feature and the third visual feature.
[0328] In some embodiments, the reconstruction unit comprises:
[0329] a first reconstruction subunit, configured to determine an input feature to be reconstructed based on the learnable embedding and the fourth visual feature, wherein the learnable embedding is configured to restore feature information in the third visual feature that is not retained in the fourth visual feature;
[0330] The second reconstruction subunit is used to process the input features to be reconstructed based on multiple Transformer blocks to obtain the fifth visual feature.
[0331] In some embodiments, the first reconstruction subunit is specifically configured to:
[0332] Perform an inversion operation on the mask information used to extract the core visual features to obtain a filtering mask;
[0333] Constructing latent unknown features based on learnable embeddings and filter masks;
[0334] constructing a known feature based on the difference feature between the fourth visual feature and the third visual feature, and the mask information;
[0335] The unknown features and difference features are fused to obtain the input features to be reconstructed.
[0336] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0337] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0338] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0339] Figure 17 A schematic block diagram of an example electronic device 1700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0340] like Figure 17As shown, device 1700 includes a computing unit 1701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1702 or a computer program loaded from a storage unit 1708 into a random access memory (RAM) 1703. Various programs and data required for the operation of device 1700 can also be stored in RAM 1703. Computing unit 1701, ROM 1702, and RAM 1703 are connected to each other via a bus 1704. An input / output (I / O) interface 1705 is also connected to bus 1704.
[0341] Various components in device 1700 are connected to I / O interface 1705, including an input unit 1706, such as a keyboard, mouse, etc.; an output unit 1707, such as various types of displays, speakers, etc.; a storage unit 1708, such as a magnetic disk, optical disk, etc.; and a communication unit 1709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1709 allows device 1700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0342] Computing unit 1701 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of computing unit 1701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. Computing unit 1701 performs the various methods and processes described above, such as the end-to-end autonomous driving control method and / or the end-to-end driving model training method. For example, in some embodiments, the end-to-end autonomous driving control method and / or the end-to-end driving model training method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as storage unit 1708. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1700 via ROM 1702 and / or communication unit 1709. When the computer program is loaded into RAM 1703 and executed by computing unit 1701, one or more steps of the end-to-end autonomous driving control method and / or the end-to-end driving model training method described above may be performed. Alternatively, in other embodiments, computing unit 1701 may be configured to perform the end-to-end autonomous driving control method and / or the end-to-end driving model training method in any other appropriate manner (e.g., via firmware).
[0343] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0344] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0345] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0346] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0347] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0348] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0349] Based on the aforementioned electronic devices, the present disclosure also provides a vehicle, which may include electronic devices, and may also include communication components, a display screen for realizing a human-machine interface, and an information collection device for collecting surrounding environment information, etc. The communication components, the display screen, the information collection device and the electronic devices are communicatively connected.
[0350] According to an embodiment of the present disclosure, the electronic device may be integrated with the communication component, the display screen, and the information collection device, or may be separately provided with the communication component, the display screen, and the information collection device.
[0351] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0352] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. An end-to-end autonomous driving method, comprising: Extracting first visual features from visual data collected by the autonomous vehicle; Pruning the first visual features to obtain key visual features; Aligning the key visual feature to the language space to obtain a second visual feature corresponding to the key visual feature; Processing the second visual feature and the navigation instruction based on a language model to obtain a future trajectory of the autonomous driving vehicle; Driving instructions for the autonomous vehicle are generated based on the future trajectory.
2. The method according to claim 1, wherein The pruning of the first visual feature to obtain the key visual feature includes: predicting probabilities of the plurality of first visual words in the first visual feature being key visual words; Based on the probabilities that the plurality of first visual word-grams are key visual word-grams, key visual word-grams are screened out from the plurality of first visual word-grams to obtain the key visual features.
3. The method according to claim 2, wherein: The predicting the probability of the plurality of first visual words in the first visual feature being key visual words includes: extracting a first global feature and a first local feature from the first visual feature; The probabilities of the plurality of first visual word-grams being key visual word-grams are predicted based on the first global feature and the first local feature.
4. The method according to claim 3, wherein: The extracting a first global feature and a first local feature from the first visual feature includes: performing feature transformation on the first visual feature to obtain the first local feature of the first visual feature; A pooling operation is performed on the first local feature to obtain the first global feature of the first visual feature.
5. The method according to claim 3, wherein The predicting, based on the first global feature and the first local feature, probabilities of the plurality of first visual word-grams being key visual word-grams comprises: Splicing the first global feature and the first local feature to obtain a first spliced feature; Perform feature transformation on the first concatenated features to obtain probabilities that the plurality of first visual word-grams are key visual word-grams.
6. The method according to claim 2, wherein: The step of selecting the key visual word from the plurality of first visual word-grams based on the probability that the plurality of first visual word-grams are key visual word-grams to obtain the key visual feature includes: Processing the probabilities of the plurality of first visual word-grams as key visual word-grams based on a continuous classification distribution relaxation method to obtain a binary mask; A key visual word-gram is obtained from the plurality of first visual word-grams based on the binary mask to obtain the key visual feature.
7. The method according to any one of claims 1 to 6, wherein The step of aligning the key visual feature to the language space to obtain a second visual feature corresponding to the key visual feature includes: performing feature enhancement on the first visual feature of the visual data based on first visual features of multiple frames of reference data within a specified neighborhood of the visual data to obtain a memory enhancement feature of the visual data; Based on the memory enhancement feature, the key visual feature is aligned to the language space to obtain a second visual feature corresponding to the key visual feature.
8. The method according to claim 7, wherein: The step of aligning the key visual feature to a language space based on the memory enhancement feature to obtain a second visual feature corresponding to the key visual feature includes: fusing the memory enhancement feature and the first visual feature to obtain a fused feature; The fused feature is input into a query transformer, and the key visual feature is input into the query transformer as a query value, so as to obtain a second visual feature corresponding to the key visual feature output by the query transformer.
9. The method according to any one of claims 1 to 8, wherein The processing of the second visual feature and the navigation instruction based on the language model to obtain the future trajectory of the autonomous driving vehicle includes: The second visual feature and the navigation instruction are processed based on a lightweight language model to obtain a future trajectory of the autonomous driving vehicle.
10. The method according to any one of claims 1 to 9, wherein The processing of the second visual feature and the navigation instruction based on the language model to obtain the future trajectory of the autonomous driving vehicle includes: An attention module based on the language model performs the following operations: determining a first attention feature between a plurality of first text words based on position embeddings of the plurality of first text words in the navigation instruction; determining a second attention feature between the plurality of second visual words based on position embeddings of the plurality of second visual words in the second visual feature; and determining a third attention feature between the navigation instruction and the second visual feature based on the plurality of first text words and the plurality of second visual words; A feature processing module based on the language model processes the first attention feature, the second attention feature, and the third attention feature to obtain the future trajectory.
11. An end-to-end autonomous driving model training method, comprising: The visual encoder based on the end-to-end autonomous driving model extracts third-party visual features from the sample vision; pruning the third visual feature based on the connector of the end-to-end autonomous driving model to obtain a core visual feature; Aligning the core visual feature to the language space to obtain a fourth visual feature corresponding to the core visual feature; processing the fourth visual feature and the instruction sample based on a language model of the end-to-end autonomous driving model to obtain an expected trajectory of the autonomous driving vehicle; The connector and the language model are optimized based on a first loss between the expected trajectory and the corresponding ground truth data.
12. The method according to claim 11, wherein The pruning of the third visual feature to obtain the core visual feature includes: predicting probabilities of the plurality of third visual words in the third visual feature being core visual words; Based on the probabilities that the plurality of third visual word-grams are core visual word-grams, core visual word-grams are screened out from the plurality of third visual word-grams to obtain the core visual feature.
13. The method according to claim 12, wherein: The predicting the probability of the plurality of third visual words in the third visual feature being core visual words includes: extracting a second global feature and a second local feature from the third visual feature; The probabilities of the plurality of third visual word-grams being core visual word-grams are predicted based on the second global feature and the second local feature.
14. The method according to claim 13, wherein: Extracting the second global feature and the second local feature from the third visual feature includes: Performing feature transformation on the third visual feature to obtain the second local feature; A pooling operation is performed on the second local feature to obtain the second global feature.
15. The method according to claim 13, wherein The predicting, based on the second global feature and the second local feature, probabilities of the plurality of third visual word-grams being core visual word-grams comprises: Splicing the second global feature and the second local feature to obtain a second spliced feature; Perform feature transformation on the second concatenated features to obtain probabilities that the plurality of third visual word-grams are core visual word-grams.
16. The method according to claim 12, wherein The step of selecting a core visual word from the plurality of third visual words based on the probability that the plurality of third visual words are core visual words to obtain the core visual feature includes: processing the probabilities of the plurality of third visual word-grams as core visual word-grams based on a continuous classification distribution relaxation method to obtain mask information; A core visual word-gram is obtained from the plurality of third visual word-grams based on the mask information to obtain the core visual feature.
17. The method according to any one of claims 11 to 16, wherein The step of aligning the core visual feature to the language space to obtain a fourth visual feature corresponding to the core visual feature includes: Based on the third visual feature of multiple frames of sample data within a preset neighborhood of the sample vision, feature enhancement is performed on the third visual feature of the sample vision to obtain a temporal coding feature of the sample vision; Based on the time coding feature, the core visual feature is aligned to the language space to obtain a fourth visual feature corresponding to the core visual feature.
18. The method according to claim 17, wherein The step of aligning the core visual feature to the language space based on the time coding feature to obtain a fourth visual feature corresponding to the core visual feature includes: fusing the time coding feature and the third visual feature to obtain a joint feature; The joint feature is input into a query transformer of the connector, and the core visual feature is input into the query transformer as a query value, to obtain a fourth visual feature corresponding to the core visual feature output by the query transformer.
19. The method according to any one of claims 11 to 18, wherein The language model is a lightweight language model.
20. The method according to any one of claims 11 to 19, wherein The language model based on the end-to-end autonomous driving model processes the fourth visual feature and the instruction sample to obtain an expected trajectory of the autonomous driving vehicle, including: The attention module based on the language model performs the following operations: determining a fourth attention feature between the plurality of second text words based on position embeddings of the plurality of second text words in the instruction sample; determining a fifth attention feature between the plurality of fourth visual words based on position embeddings of the plurality of fourth visual words in the fourth visual feature; and determining a sixth attention feature between the instruction sample and the fourth visual feature based on the plurality of second text words and the plurality of fourth visual words; A feature processing module based on the language model processes the fourth attention feature, the fifth attention feature, and the sixth attention feature to obtain the expected trajectory.
21. The method according to claim 11, wherein The end-to-end autonomous driving model further includes a control conversion module, which is used to process the expected trajectory and obtain control instructions for the autonomous driving vehicle to travel along the expected trajectory; Among them, the model parameters of the control conversion module participate in the optimization.
22. The method according to any one of claims 11 to 21, further comprising: Reconstructing the fourth visual feature based on the auxiliary training device to obtain a fifth visual feature; Optimizing the connector and the language model based on a first loss between the expected trajectory and the corresponding data ground truth includes: The connector, the language model, and the auxiliary trainer are optimized based on a first loss between the expected trajectory and the corresponding data ground truth, and a second loss between the fifth visual feature and the third visual feature.
23. The method according to claim 22, wherein The reconstructing the fourth visual feature to obtain the fifth visual feature includes: Determining an input feature to be reconstructed based on the learnable embedding and the fourth visual feature, wherein the learnable embedding is used to restore feature information in the third visual feature that is not retained in the fourth visual feature; The input feature to be reconstructed is processed based on multiple Transformer blocks to obtain the fifth visual feature.
24. The method according to claim 23, wherein The determining of the input feature to be reconstructed based on the learnable embedding and the fourth visual feature includes: performing an inversion operation on the mask information used to extract the core visual features to obtain a filtering mask; constructing a potential unknown feature based on the learnable embedding and the filter mask; constructing a known feature based on a difference feature between the fourth visual feature and the third visual feature, and the mask information; The unknown features and the difference features are fused to obtain the input features to be reconstructed.
25. An end-to-end autonomous driving device, comprising: a first extraction unit, configured to extract a first visual feature from visual data collected by the autonomous driving vehicle; a pruning unit, configured to prune the first visual features to obtain key visual features; an alignment unit, configured to align the key visual feature to a language space to obtain a second visual feature corresponding to the key visual feature; a prediction unit, configured to process the second visual feature and the navigation instruction based on a language model to obtain a future trajectory of the autonomous driving vehicle; A control unit is configured to generate driving instructions for the autonomous vehicle based on the future trajectory.
26. An end-to-end autonomous driving model training device, comprising: A second extraction unit is configured to extract a third visual feature from the sample vision based on a visual encoder of an end-to-end autonomous driving model; a sparse unit, configured to prune the third visual feature based on a connector of the end-to-end autonomous driving model to obtain a core visual feature; Aligning the core visual feature to the language space to obtain a fourth visual feature corresponding to the core visual feature; an estimation unit, configured to process the fourth visual feature and the instruction sample based on a language model of the end-to-end autonomous driving model to obtain an expected trajectory of the autonomous driving vehicle; An optimization unit is configured to optimize the connector and the language model based on a first loss between the expected trajectory and the corresponding data ground truth.
27. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 24.
28. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-24.
29. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 24.
30. An autonomous driving vehicle comprising the electronic device of claim 27.
Citation Information
Cited By
Visual language model-based driving track planning method and intelligent driving system
CN121212372A
Motion trajectory planning method and device, intelligent system with body and automatic driving vehicle
CN121404319A