A multi-modal large language model-based automatic driving perception enhancement method and system
Patent Information
- Application Number
- CN202610804557.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明的目的在于克服上述缺陷,提供一种基于多模态大语言模型的自动驾驶感知增强方法及系统,解决BEV感知模型在复杂或边缘场景下易产生语义矛盾或高度不确定性预测,以及大语言模型难以直接应用于实时系统的问题
[0017]与现有技术相比,本发明的优势之处在于:
Smart Images

Figure CN122830719A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving technology, and in particular to an autonomous driving perception enhancement method and system based on a multimodal large language model. Background Technology
[0002] With the rapid development of autonomous driving technology, higher demands are being placed on the accuracy, robustness, and semantic understanding capabilities of vehicle environment perception systems. Bird's-eye view perception frameworks, through unified multi-sensor feature projection and fusion, have become one of the mainstream solutions for autonomous driving environment understanding. However, traditional BEV perception methods still face challenges in complex, variable, or rare edge case scenarios.
[0003] In existing technologies, BEV perception models are typically data-driven, and their performance is limited by the coverage of the training data. When encountering complex scenes not fully represented in the training data, severe occlusion, or sudden changes in lighting, the model is prone to producing semantically contradictory perception results or highly uncertain predictions. For example, the perception system might misdetect a vehicle image on a large billboard as a real vehicle, or misjudge the shadow of a pedestrian crossing the road as an obstacle. This unreliability at the perception level directly impacts downstream planning and control modules, posing a potential threat to driving safety. Multimodal large language models possess powerful cross-modal semantic understanding and reasoning capabilities, enabling comprehensive scene analysis by combining information such as images, point clouds, and text. However, directly applying them to real-time perception would face problems such as excessive computational overhead and latency that cannot meet the real-time requirements of autonomous driving.
[0004] Therefore, how to efficiently introduce the deep semantic understanding capabilities of multimodal large language models into the BEV perception framework, and specifically enhance the perception robustness and decision confidence in complex scenarios, has become an urgent technical problem to be solved. Summary of the Invention
[0005] The purpose of this invention is to overcome the above-mentioned defects and provide an autonomous driving perception enhancement method and system based on a multimodal large language model, which solves the problems that BEV perception models are prone to semantic contradictions or highly uncertain predictions in complex or edge scenarios, and that large language models are difficult to directly apply to real-time systems.
[0006] To achieve the above objectives, this invention proposes an autonomous driving perception enhancement method based on a multimodal large language model. The core of this method lies in its deep collaboration with a BEV-based perception framework, using the multimodal large language model as an enhancement module of the perception system and configuring a built-in monitoring module. When the built-in monitoring module detects semantic contradictions or high uncertainty in the output of the traditional perception process, it triggers the multimodal large language model to perform scene parsing and target re-identification, generating a structured semantic guidance signal. This signal is used to modulate the attention mechanism in the subsequent BEV feature generation process, thereby enhancing attention to key or questionable targets and generating semantically enhanced current frame BEV features, providing more reliable environmental perception input to the downstream control module.
[0007] The method includes the following steps: S1: Collect multi-sensor data and extract features, project the features onto the bird's-eye view BEV space for alignment and fusion, and obtain multimodal fusion features; S2: The historical frame BEV features are fused with the BEV query matrix and subjected to temporal enhancement processing to obtain temporally enhanced BEV features. The multimodal fused features and the temporally enhanced BEV features are then processed sequentially using the Transformer architecture, including spatial cross-attention, addition and regularization, feedforward network, and double addition and regularization, to generate the current frame BEV features. S3: Detect whether the perception result obtained based on the BEV features of the current frame has semantic contradictions or high uncertainty; If it does not exist, the current frame BEV feature is output directly; If present, the multimodal large language model is triggered to perform target detection and scene understanding on the original sensor data or intermediate features, generating a structured semantic guidance signal. The semantic guidance signal is converted into an attention mask corresponding to the BEV feature map space. When the Transformer architecture performs spatial cross-attention calculation, the attention mask is used to weight and modulate the attention score to improve the attention to key regions. The semantically enhanced current frame BEV features are then regenerated and output. S4: Generate or adjust vehicle control commands based on the semantically enhanced BEV features of the current frame in the final output.
[0008] Furthermore, in step S1, the multi-sensor data includes multi-view image data, lidar point cloud data, millimeter-wave radar point cloud data, and navigation and positioning information data.
[0009] Furthermore, step S2 specifically includes: combining historical frame BEV features with the BEV query matrix, obtaining temporally enhanced BEV features through temporal self-attention, addition, and regularization processing, and then using the Transformer architecture to sequentially process the multimodal fusion features and temporally enhanced BEV features through spatial cross-attention, addition and regularization, feedforward network, and secondary addition and regularization processing to generate environmentally enhanced current frame BEV features.
[0010] Furthermore, when performing spatial cross-attention calculation using the Transformer architecture, the attention score is weighted and modulated using the attention mask, specifically including: 1) The key matrix K, based on the BEV query matrix Q and the multimodal fusion features, is calculated using the following formula: Calculate the raw attention score ; 2) To adapt to parallel computation with multiple attention heads, the attention mask is... ( The dimension is adjusted to be compatible with the original attention score broadcast, denoted as... ; 3) Adjust the attention mask Compared with the original attention score Element-wise multiplication yields the modulated attention score, and the modulation formula is: = ⊙ To improve the attention response of the corresponding area within the grid coordinate range; 4) Perform Softmax normalization based on the modulated attention scores to obtain the modulated attention weights. ; 5) Modulate the attention weights The value matrix of the multimodal fusion features is weighted and aggregated to generate the semantically enhanced current frame BEV features.
[0011] Furthermore, in step S3, the process of detecting whether the perception result has semantic contradictions or high uncertainty includes: Preprocessing stage: Obtain multiple perception results output by the perception network, including 2D bounding boxes of target detection, semantic segmentation masks, drivable area masks, and corresponding prediction confidence scores; project the 2D bounding boxes output by the detection head onto the BEV coordinate system, map the semantic segmentation masks and drivable area masks onto a grid of the same size as the BEV feature map, and extract the multi-perception head prediction confidence vector and the average classification entropy of the last 5 frames; Semantic contradiction detection stage: Semantic contradiction detection is performed based on the multiple perception results. The semantic contradiction detection includes: spatial overlap conflict judgment between the target and static obstacles, conflict judgment between dynamic targets and non-drivable areas, and category mutual exclusion judgment of the same spatial coordinates. High uncertainty quantification stage: High uncertainty quantification is performed based on the prediction confidence of the multiple perception results. The high uncertainty quantification includes measuring classification fuzziness by target classification probability entropy and measuring detection reliability by target detection confidence. If either the semantic contradiction detection or the high uncertainty quantification condition is met, then the perception result is determined to have a semantic contradiction or high uncertainty.
[0012] Furthermore, in step S3, the structured semantic guidance signal includes at least: object category, bounding box in BEV, uncertainty score, and attention weight generated based on the uncertainty score; The generation process of the structured semantic guidance signal includes: 1) Extract image features, point cloud geometric features, BEV global semantic features, and local target features from the original sensor data or intermediate features through multimodal input parsing; 2) Cross-modal feature fusion strengthens the association of information within a modality through self-attention, and constructs a cross-modal cross-attention mapping based on text features to achieve multimodal information alignment; 3) Predict object categories using a Softmax classifier, predict BEV bounding boxes using a regression head, calculate uncertainty scores based on category prediction probabilities and bounding box regression confidence, and apply the mapping formula: It generates attention weights and encapsulates them into a structured dictionary output to achieve structured field prediction.
[0013] Furthermore, in step S3, the semantic guidance signal is converted into an attention mask corresponding to the BEV feature map space, specifically including: 1) Based on the grid size of the BEV feature map, create an initial mask matrix with an initial value of the first value; 2) Based on the grid resolution of the BEV feature map, map the physical coordinates of the BEV bounding box in the semantic guidance signal to the grid coordinates of the initial mask matrix; wherein, The formula for mapping the horizontal coordinate of the grid is: ; The formula for mapping the grid's ordinate is: ; 3) Replace the mask element values corresponding to the grid coordinate range with the attention weight values in the semantic guidance signal, and keep the element values outside the grid coordinate range as the first value. Perform linear interpolation on the edge grid of the BEV bounding box so that the weight gradually changes from the attention weight to the first value to form an attention mask. The attention weight value is greater than the first value.
[0014] This invention also proposes an autonomous driving perception enhancement system based on a multimodal large language model, used to implement the above-mentioned autonomous driving perception enhancement method based on a multimodal large language model. The system includes: Feature processing module: used to collect multi-sensor data and extract features, project the features onto the BEV space for alignment and fusion to obtain multimodal fusion features, and fuse historical frame BEV features with the BEV query matrix and perform temporal enhancement processing to generate current frame BEV features with enhanced environmental perception. The semantic enhancement module includes a monitoring submodule and a multimodal large language model submodule; among which, The monitoring submodule is used to detect whether there is semantic contradiction or high uncertainty in the perception result obtained based on the current frame BEV features. If not, the current frame BEV features are directly output. The multimodal large language model submodule is triggered when the monitoring submodule detects semantic contradictions or high uncertainty. It performs target detection and scene understanding on the original sensor data or intermediate features, generates a structured semantic guidance signal, converts the semantic guidance signal into an attention mask corresponding to the BEV feature map space, and uses the attention mask to weight and modulate the attention score when performing spatial cross-attention calculation, regenerates the semantically enhanced current frame BEV features and outputs them. Control interaction module: used to receive the current frame BEV features output by the feature processing module or the semantic enhancement module, and provide them to the downstream control module to support the generation of control commands or parameter adjustment; The downstream control module is used to generate or adjust vehicle control commands based on the received BEV features of the current frame.
[0015] Furthermore, the downstream control module is a controller based on optimization control, imitation learning, or reinforcement learning.
[0016] The present invention also proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described perception enhancement method for autonomous driving based on a multimodal large language model.
[0017] Compared with the prior art, the advantages of the present invention are: 1. This invention utilizes a multimodal large language model as a perception enhancement and semantic supervision module, deeply collaborating with the traditional BEV perception framework. When conventional perception encounters semantic contradictions or high uncertainty, the system leverages the deep semantic understanding and reasoning capabilities of the multimodal large language model to analyze the scene and generate semantic guidance signals. This modulates the attention mechanism in BEV feature generation, significantly improving the system's robustness and semantic consistency in perceiving key targets and rare edge cases. This semantically enhanced BEV feature provides downstream control modules with a more confident and interpretable environmental representation. The core of this invention lies in enhancing the reliability of the perception system in edge scenarios and, by providing higher-quality perception input, laying the foundation for improving the overall decision-making safety and adaptability of the vehicle in complex environments.
[0018] 2. This invention evaluates perception quality online through a built-in monitoring module, triggering the multimodal large language model only when semantic contradictions or high uncertainty are detected. This on-demand triggering mechanism solves the problems of excessive computational overhead and latency that cannot meet automotive-grade requirements caused by introducing large models for real-time full-scale inference. This enables the system to maintain a lightweight and fast response in most common scenarios, effectively balancing perception robustness and real-time performance.
[0019] 3. Compared to existing multi-model ensemble techniques that typically only perform result fusion or voting at the output level, this invention achieves deep and precise enhancement from semantic understanding to feature representation through an attention mask modulation mechanism. In this invention's method, when enhancement is triggered, the structured semantic guidance signal generated by the multimodal large language model is converted into an attention mask aligned with the BEV feature map space and directly applied to the score calculation process of cross-attention in the Transformer space. That is, by modulating the distribution of the original attention score through element-wise multiplication, the attention response of key regions is selectively amplified, and then the features are re-aggregated after Softmax normalization. This allows the semantic understanding capability of the multimodal large language model to directly shape the underlying representation of BEV features, which is smoother and more effective than simple result replacement or weighted fusion.
[0020] 4. The method of this invention cross-validates the perception results from two dimensions: semantic logic violation and insufficient confidence. Semantic contradiction detection covers three typical logical errors: spatial overlap between the target and static obstacles, conflict between dynamic targets and non-drivable areas, and category mutual exclusion. High uncertainty quantification measures the reliability of the perception output itself using both classification entropy and detection reliability. This multi-dimensional criterion parallel detection mechanism can cover a wider range of perception anomaly types, reduce the risk of missed detections, and ensure reliable triggering of the enhancement process when truly needed.
[0021] 5. In the method of this invention, the semantic guidance signal output by the multimodal large language model includes at least the object category, BEV bounding box, and attention weights. The attention weights are generated based on an uncertainty score mapping, meaning that targets with lower uncertainty receive higher attention weights. This mechanism gives the enhancement process a clear physical meaning and interpretability: the system does not blindly enhance, but rather allocates attention resources selectively based on the degree of certainty in the target judgment, thereby achieving precise and differentiated semantic guidance for key regions and providing an interpretable control method for the enhancement process.
[0022] 6. The semantically enhanced BEV features output by the method of this invention provide a more reliable and semantically consistent environmental representation while maintaining a consistent interface with the standard BEV space. The enhanced BEV features are compatible with and enable various downstream control modules, exhibiting good universality. They can serve as an independent perception enhancement front-end, seamlessly integrating with existing mainstream control architectures to provide higher-quality state inputs and indirectly improve decision-making security and adaptability across multiple technical approaches. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating the process of triggering the generation of semantically enhanced BEV features by a multimodal large language model in the method of Embodiment 1 of the present invention.
[0024] Figure 2 This is a schematic diagram of the overall operation framework of the autonomous driving perception enhancement method based on a multimodal large language model proposed in Embodiment 1 of the present invention.
[0025] Figure 3 A schematic diagram illustrating the process of generating structured semantic guidance signals for a multimodal large language model as exemplified in this application. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be further described below.
[0027] First, let's explain the relevant terms: BEV (Balanced Electric Vehicle) is a method for representing a scene in two-dimensional or simplified three-dimensional space from a vertical top-down perspective, and it is widely used in fields such as autonomous driving and robot navigation. Its core is to transform complex three-dimensional environmental information into a vehicle-centric top-down plane through the projection transformation of sensor data, thereby simplifying the understanding and calculation of spatial relationships.
[0028] Historical frame BEV features refer to the set of BEV features generated from multiple frames of sensor data in a continuous time series in an autonomous driving system.
[0029] The BEV query matrix is a core component in the autonomous driving perception model used to generate BEV features. It is usually composed of a set of grid-like learnable parameters and dynamically aggregates multi-view images and temporal information through an attention mechanism.
[0030] Example 1 This embodiment provides a method for enhancing perception in autonomous driving based on a multimodal large language model. For example... Figures 1-3 As shown, the method specifically includes the following steps: S1: Collect data from multiple sensors and extract features. Project the extracted features onto the bird's-eye view BEV space for alignment and fusion to obtain multimodal fusion features. In this embodiment, multi-sensor data is collected using devices such as cameras, LiDAR, millimeter-wave radar, and GPS / BeiDou satellites. This multi-sensor data includes multi-view image data, LiDAR point cloud data, millimeter-wave radar point cloud data, and navigation and positioning information data. For example... Figure 2 As shown, multi-view images can include camera data from the front, rear, left front, left rear, right front, and right rear of the vehicle; LiDAR point cloud data can be used to provide accurate three-dimensional geometric information, such as obstacle contour recognition and lane line detection; millimeter-wave radar point cloud data provides target relative speed information, such as the relative speed information of the vehicle and surrounding targets; and navigation positioning information data provides the absolute or relative position information of the vehicle and surrounding vehicles.
[0031] Furthermore, feature extraction can employ specialized networks for different modalities. For example, in this embodiment, a CNN network is used to extract local features from image data, and a 3D convolutional or point cloud feature encoding network is used to extract geometric and spatial features from LiDAR and millimeter-wave radar point clouds, obtaining heterogeneous feature vectors from multiple perspectives. Subsequently, cross-attention mechanisms such as Transformer are used to project the extracted 2D / 3D features onto a unified BEV coordinate system, achieving projection and alignment of 2D / 3D features into the BEV space. This unifies the features from different sensors into the BEV coordinate system at the vehicle center, forming multimodal fusion features through alignment and fusion, providing a unified representation for subsequent temporal enhancement and perception.
[0032] S2: The historical frame BEV features are fused with the BEV query matrix and temporally enhanced to obtain temporally enhanced BEV features. The multimodal fused features and temporally enhanced BEV features are then processed sequentially using the Transformer architecture, including spatial cross-attention, addition and regularization, feedforward network, and double addition and regularization, to generate the current frame BEV features. In this embodiment, the first step is to perform fusion and temporal feature enhancement: combining historical frame BEV features with a learnable BEV query matrix, and then sequentially performing temporal self-attention calculation, addition, and regularization to obtain temporally enhanced BEV features. Temporal self-attention can effectively capture the motion trend of dynamic targets and the stability of static scenes by sampling features in historical frames through deformable attention.
[0033] Then, the multimodal fusion features and temporally enhanced BEV features are deeply fused using the Transformer architecture. Specifically, the temporally enhanced BEV features are used as the query, and the multimodal fusion features obtained in step S1 are used as the key and value. Spatial cross-attention, addition and regularization, feedforward network, and double addition and regularization are performed sequentially to finally generate the BEV features of the current frame. This feature has fused multi-sensor spatial information and temporal dynamic information, and can characterize the environmental state at the current moment.
[0034] S3: Detect whether there is semantic contradiction or high uncertainty in the perception result obtained based on the current frame BEV features; if not, directly output the current frame BEV features; if so, trigger the multimodal large language model to perform target detection and scene understanding on the original sensor data or intermediate features, generate a structured semantic guidance signal, convert the generated semantic guidance signal into an attention mask corresponding to the BEV feature map space, and use the attention mask to weight and modulate the attention score when performing spatial cross-attention calculation in the Transformer architecture to improve the attention to key regions, regenerate the semantically enhanced current frame BEV features and output them.
[0035] In this embodiment, the detection process includes a preprocessing stage, a semantic contradiction detection stage, a high uncertainty quantification stage, and conflict aggregation and final determination: The preprocessing stage includes: projecting the 2D bounding box output by the target detection head onto the BEV coordinate system using camera intrinsic and extrinsic parameters, resulting in a format of [ , , , The bounding box of the BEV is obtained; the semantic segmentation mask and the drivable area mask are mapped to a grid of the same size as the BEV feature map through the projection matrix to ensure that the spatial coordinates are aligned one by one; the prediction confidence vector of the multi-sensor head is extracted and the average classification entropy of the last 5 frames is cached for time-series auxiliary judgment of uncertainty assessment.
[0036] Semantic contradiction detection phase: Semantic contradiction detection is performed in the aligned BEV space in the order of "spatial logical conflict → category mutual exclusion conflict" to discover perceptual results that violate common sense or physical laws. This includes, but is not limited to, the following detection judgments: spatial logical conflict detection judgment and category mutual exclusion conflict detection judgment.
[0037] The spatial logic conflict detection and judgment includes the spatial overlap conflict judgment between the target and static obstacles and the conflict judgment between the dynamic target and the non-drivable area. The parameters involved are explained as follows: = [ , , , The dynamic targets output by the detection head include the bounding boxes of vehicles, pedestrians, cyclists, etc. in the BEV space; The static scene categories output by the semantic segmentation head, including walls, fences, flower beds, etc., are denoted as a set. = {wall, barrier, flowerbed,...} corresponds to the BEV mask region; This is the BEV mask for the drivable area output by the drivable area prediction head, where 1 indicates drivable and 0 indicates non-drivable. Category mutual exclusion conflict detection includes category mutual exclusion judgment for the same spatial coordinates, involving the following parameters: [Settings would be inserted here]. To detect the head in BEV coordinates Semantic tags at the location, For semantic segmentation head in Semantic tags at the location, The predefined set of mutually exclusive pairs for each category includes {(vehicle, wall), (pedestrian, building), (lane line, obstacle), ...}.
[0038] The specific judgments for each conflict detection are as follows: (1) Spatial overlap and conflict judgment between target and static obstacle: Calculate the dynamic target bounding box Masking of static obstacles (such as walls, fences, flower beds) intersection ratio ( , If the intersection-union ratio exceeds the threshold, i.e. ( , )= > If the expression is incorrect, a semantic contradiction is determined to exist. The empirical threshold for the intersection-union ratio is set to... = 0.3, and .
[0039] (2) Conflict judgment between dynamic targets and non-drivable areas: Calculate the bounding box of dynamic targets. The percentage of the drivable area; if the percentage is less than a threshold, i.e. ( )= < If so, a conflict is determined. Among these, the percentage threshold... = 0.5, and the target category is the type that needs to move within the drivable area, including vehicles, pedestrians, cyclists, etc.
[0040] (3) Class mutual exclusion judgment at the same spatial coordinate: If at a certain coordinate of BEV, the semantic label output by the detection head and the semantic label output by the segmentation head belong to a predefined mutually exclusive pair (such as vehicle and wall, pedestrian and building), and the confidence of both exceeds the threshold, then it is determined that there is a class mutual exclusion contradiction.
[0041] High uncertainty quantification stage: used to assess the unreliability of the perception network's output itself. It includes: (a) Classification fuzziness quantification: For fuzzy classification results, the fuzziness is measured by the target classification probability entropy. Specifically, this involves extracting the category probability vector output by the perceptual network. Through calculation formula Calculate the information entropy value. Set the trigger threshold. (For example =1.2), if the entropy value exceeds the preset threshold, it indicates that the classification result is highly ambiguous, and it is judged that the classification uncertainty exceeds the standard.
[0042] (b) Quantification of Detection Reliability: The reliability of the perceived results is measured by the target detection reliability score. Specifically, this involves extracting the confidence score output by the detection head. This score can be calculated from the classification branch probability and the localization branch probability. The predicted values are obtained by weighted fusion, and the formula is: = Among them, the weighting coefficient =0.7, the predicted probability of the category is ,position Predicted value ; Set the detection confidence threshold (For example =0.5 (verified by actual vehicles; detection results below this value have a false positive rate exceeding 40%), if the confidence score is... Below the detection confidence threshold, i.e., when If the test result is too high, it indicates that the test result is unreliable and is judged as exceeding the test uncertainty standard.
[0043] Conflict aggregation and final determination: If at least one of the above semantic contradiction detection criteria is met, it is determined that "a semantic contradiction exists"; if at least one of the high uncertainty quantification criteria is met, it is determined that "high uncertainty exists". Meeting either criterion triggers the enhancement process, such as... Figure 1 As shown: If the determination result is negative, it indicates that the current perception result is reliable. Then, the current frame BEV feature generated in step S2 is directly output without introducing additional computational overhead, thus ensuring real-time performance.
[0044] If the determination result is yes, the multimodal large language model is triggered for enhancement processing. The multimodal large language model performs target detection and scene understanding on the raw sensor data or intermediate features, generating a structured semantic guidance signal. This model can employ a four-step process: "multimodal input parsing → cross-modal feature fusion → structured field prediction → signal formatting," to achieve target detection and scene understanding on the raw sensor data or intermediate features, ultimately outputting a structured semantic guidance signal containing object category, BEV bounding box, uncertainty score, and attention weights. Figure 3 As shown, the specific operation process is as follows: a. Multimodal input parsing: Based on different input types, targeted parsing is performed on the two input paths, raw sensor data and intermediate features, to ensure effective extraction of input information.
[0045] For raw sensor data, multi-scale features of the image can be extracted using the visual encoder built into the MLLM, outputting an image feature sequence that focuses on capturing target contours, textures, and spatial relationships. For point cloud analysis, the LiDAR point cloud can be voxelized to filter out invalid voxels, and then voxel-level geometric features can be extracted through 3D convolution. At the same time, the statistical information of the point cloud can be converted into a structured text description.
[0046] For navigation data parsing, the navigation positioning data is converted into standardized text.
[0047] For intermediate feature input parsing, this includes BEV feature block parsing, multi-sensor head intermediate feature parsing, and conflict region feature maps from the segmentation head. Specifically: BEV feature block parsing involves flattening the BEV feature blocks into a 1D sequence along the spatial dimension, normalizing it through a LayerNorm layer, and then inputting it into the feature encoding layer of a multimodal large language model to extract global semantic features in the BEV space. Multi-sensor head intermediate feature parsing involves concatenating candidate box features with the flattened segmentation features into a unified sequence, which is then linearly projected to a dimension consistent with the image features, serving as a supplement to local target features.
[0048] b. Cross-modal feature fusion: The multimodal large language model achieves deep fusion of multimodal features through the "double cross-attention mechanism" to ensure the complementarity of information from images, point clouds, text, and intermediate features.
[0049] Specifically, self-attention calculation is performed separately for each modal feature to strengthen the information association within the modality; then, based on the text features, cross-attention mappings of text-image, text-point cloud, and text-intermediate features are constructed to achieve cross-modal information alignment and output a unified fusion feature vector.
[0050] c. Generate structured semantic guidance signals: A multimodal large language model is used to predict four core fields: object category, BEV bounding box, uncertainty score, and attention weight, ensuring the accuracy and computability of the output. Specifically: Object category prediction uses a Softmax classifier, which outputs a predefined set of traffic object categories, including the probability distribution of vehicles, pedestrians, cyclists, guardrails, lane lines, etc. The category with the highest probability is selected as the prediction result, and the category probability is denoted as . ; BEV bounding box prediction uses the regression head to output four-dimensional coordinates in the BEV coordinate system, denoted as ( (Unit: m), the regression loss uses smooth L1 loss to ensure the smoothness of coordinate prediction; Uncertainty score prediction uses a regression head activated by Sigmoid, outputting a value between 0 and 1. It comprehensively reflects the probability dispersion of the category prediction and the confidence level of the bounding box regression. The calculation formula is as follows: ,in The confidence level of the bounding box regression; Attention weight prediction is generated based on an uncertainty score mapping, and the mapping formula is: To ensure that targets with lower uncertainty receive higher attention weights, the default weight value for key regions in this embodiment is 2.4.
[0051] Finally, the prediction results of the above four fields are encapsulated into a structured dictionary format to ensure that downstream modules can directly call them.
[0052] like Figure 3 As shown, in this embodiment, the clearest original front / around view image is directly input, or a key region image back-projected from BEV features. The structured semantic guidance signal includes object category, BEV bounding box, uncertainty score, and attention weight. The BEV bounding box is directly the coordinates in the BEV space, and the attention weight is determined based on the target importance and uncertainty score. The uncertainty score predicts probability through category. BEV bounding box regression confidence The weighted calculation is derived based on the weighted formula, which is: Attention weights are generated based on an uncertainty score mapping, and the mapping formula is as follows: Within the bounding box area, a higher weight value of 2.4 is assigned based on the importance and uncertainty of the target, while the weight of non-critical areas is 1.0.
[0053] d. Convert to attention mask M mask The generated semantic guidance signal is converted into a two-dimensional prior weight map corresponding to the BEV feature map space, i.e., the attention mask M. maskThe specific conversion steps are as follows: (1) Initialize the mask matrix: using the current BEV feature map grid size, i.e. Create an initial mask matrix based on a 200×200 matrix. All elements have a value of 1.0, representing the basic weight of non-critical areas.
[0054] (2) Mapping BEV bounding box coordinates to the grid range: Based on the grid resolution of the BEV feature map (e.g., each grid represents 0.5 meters), the physical coordinates of the BEV bounding box are mapped to BEV grid coordinates. The mapping formula is: The formula for mapping the horizontal coordinate of the grid is: ; The formula for mapping the grid's ordinate is: ; It should be noted that the grid size, grid resolution, and specific coordinate values of the BEV feature map can be adjusted according to the actual driving scenario. For example, the BEV feature map uses a default grid size of 200×200 and a default grid resolution of 0.5m. The BEV bounding box follows the format ([x_{min}, y_{min}, x_{max}, y_{max}]), with the x-axis along the vehicle's forward direction, the y-axis perpendicular to the x-axis to the left, and the origin at the center of the vehicle's rear axle. The unit is meters, and the specific coordinate values are generated by the multimodal large language model based on the original sensor data or intermediate feature inference. Those skilled in the art can adjust the grid resolution to any value between 0.25m and 1.0m according to different scenarios such as highways and urban roads, and the grid size can be adjusted accordingly without departing from the protection scope of this invention.
[0055] (3) Weight assignment and mask generation: Replace the mask elements within the grid range obtained by the above mapping with the attention weights in the semantic guidance signal, and keep the initial value of 1.0 for the elements outside the grid range to form the final attention mask, dimension and BEV features. Figure 1 To, that is .
[0056] (4) Mask smoothing: In order to avoid feature distortion caused by abrupt changes in weight between the target boundary mesh and the non-boundary mesh, linear interpolation is performed on the edge mesh of the BEV bounding box to gradually change the weight from a high value of 2.4 to 1.0, so as to ensure the smoothness of attention modulation and avoid abrupt changes.
[0057] After the mask is constructed, it is introduced into the spatial cross-attention calculation of the Transformer in the BEV space. Weighted modulation is achieved by multiplying the result by the attention score. The specific process is as follows: First, calculate the dot product between the BEV query matrix Q and the multimodal fusion feature key matrix K. The formula is as follows: The original attention score S is obtained. raw ; Then, to accommodate parallel computation with multiple attention heads, the attention mask is... ( Adjusted to broadcastable dimension , recorded as Ensure that it matches the original attention score S. raw Element-wise multiplication. Using the element-wise multiplication operator ⊙, and the adapted mask. raw attention score Weighted modulation is applied to enhance the attentional response in key areas, using the following modulation formula: S mod = S raw ⊙M adapt The modulated fraction S is obtained. mod At this point, the attention score for key areas is amplified, while that for non-key areas remains unchanged.
[0058] Next, the modulated attention score S mod According to the last dimension Perform Softmax normalization to obtain the modulated attention weights. Pay attention weights Value matrix of multimodal fusion features Multiply, where The spatial cross-attention feature aggregation is completed. Then, following the subsequent steps of S2, the processes of "addition and regularization, feedforward network, and second addition and regularization" are executed sequentially to generate and output the semantically enhanced BEV features of the current frame. It should be noted that when the multimodal large language model is not triggered, the attention mask... All elements remain at 1.0; at this point, S mod = S raw This is completely consistent with the original spatial cross-attention calculation logic in this step, with no additional performance loss.
[0059] S4: Generate or adjust vehicle control commands based on the semantically enhanced BEV features of the current frame in the final output.
[0060] This embodiment's method performs online semantic logic verification and uncertainty assessment on the perception results. It only triggers a multimodal large language model for deep analysis when anomalies are detected, and uses the prior knowledge output by the model to correct the BEV feature generation process through attention modulation. This mechanism significantly improves the system's perceptual robustness and semantic consistency in edge cases and complex scenarios while maintaining low latency in normal scenarios.
[0061] Example 2 This embodiment 2 proposes an autonomous driving perception enhancement system based on a multimodal large language model, used to implement the autonomous driving perception enhancement method based on a multimodal large language model in embodiment 1 above. The system includes: Feature processing module: used to perform steps S1 and S2, namely, to collect multi-sensor data and extract features, project the features onto the BEV space for alignment and fusion to obtain multimodal fusion features, and fuse historical frame BEV features with the BEV query matrix and perform temporal enhancement processing to generate current frame BEV features with enhanced environmental perception.
[0062] The semantic enhancement module includes a monitoring submodule and a multimodal large language model submodule, which are used to execute step S3; among them, The monitoring submodule is used to execute the detection logic as in Example 1, that is, to detect whether there is semantic contradiction or high uncertainty in the perception result obtained based on the current frame BEV features. If not, the current frame BEV features are directly output. The multimodal large language model submodule is triggered when the monitoring submodule detects semantic contradiction or high uncertainty. It performs target detection and scene understanding on the original sensor data or intermediate features, generates a structured semantic guidance signal, converts the semantic guidance signal into an attention mask corresponding to the BEV feature map space, and uses the attention mask to weighted modulate the attention score when performing spatial cross-attention calculation, regenerates the semantically enhanced current frame BEV features and outputs them.
[0063] Control interaction module: Used to receive the current frame BEV features output by the feature processing module or semantic enhancement module and provide them to the downstream control module, supporting the generation of control commands or parameter adjustment; Downstream control module: This module executes step S4, which involves generating or adjusting vehicle control commands based on the received BEV features of the current frame. This downstream control module is a controller based on optimized control, imitation learning, or reinforcement learning. The semantically enhanced BEV features serve as its input, providing a more reliable and interpretable environmental representation, thereby supporting the control module in generating safer vehicle control commands, such as steering, acceleration, and braking.
[0064] Example 3 This embodiment 3 proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the autonomous driving perception enhancement method based on a multimodal large language model described in embodiment 1 above.
[0065] The above are merely preferred embodiments of the present invention and do not constitute any limitation on the present invention. Any equivalent substitutions or modifications made by those skilled in the art to the technical solutions and content disclosed in the present invention without departing from the scope of the present invention shall be deemed to have remained within the scope of protection of the present invention.
Claims
1. A method for enhancing perception in autonomous driving based on a multimodal large language model, characterized in that, Includes the following steps: S1: Collect multi-sensor data and extract features, project the features onto the bird's-eye view BEV space for alignment and fusion, and obtain multimodal fusion features; S2: The historical frame BEV features are fused with the BEV query matrix and subjected to temporal enhancement processing to obtain temporally enhanced BEV features. The multimodal fused features and the temporally enhanced BEV features are then processed sequentially using the Transformer architecture, including spatial cross-attention, addition and regularization, feedforward network, and double addition and regularization, to generate the current frame BEV features. S3: Detect whether the perception result obtained based on the BEV features of the current frame has semantic contradictions or high uncertainty; If it does not exist, the current frame BEV feature is output directly; If present, the multimodal large language model is triggered to perform target detection and scene understanding on the original sensor data or intermediate features, generating a structured semantic guidance signal. The semantic guidance signal is converted into an attention mask corresponding to the BEV feature map space. When the Transformer architecture performs spatial cross-attention calculation, the attention mask is used to weight and modulate the attention score to improve the attention to key regions. The semantically enhanced current frame BEV features are then regenerated and output. S4: Generate or adjust vehicle control commands based on the semantically enhanced BEV features of the current frame in the final output.
2. The autonomous driving perception enhancement method based on a multimodal large language model according to claim 1, characterized in that, In step S1, the multi-sensor data includes multi-view image data, lidar point cloud data, millimeter-wave radar point cloud data, and navigation and positioning information data.
3. The autonomous driving perception enhancement method based on a multimodal large language model according to claim 1, characterized in that, Step S2 specifically includes: combining the historical frame BEV features with the BEV query matrix, and obtaining temporally enhanced BEV features through temporal self-attention, addition, and regularization processing.
4. The autonomous driving perception enhancement method based on a multimodal large language model according to claim 1, characterized in that, Step S3, the process of detecting whether the perception result has semantic contradictions or high uncertainty includes: Preprocessing stage: Obtain multiple perception results output by the perception network, including 2D bounding boxes of target detection, semantic segmentation masks, drivable area masks, and corresponding prediction confidence scores; project the 2D bounding boxes output by the detection head onto the BEV coordinate system, map the semantic segmentation masks and drivable area masks onto a grid of the same size as the BEV feature map, and extract the multi-perception head prediction confidence vector and the average classification entropy of the last 5 frames; Semantic contradiction detection stage: Semantic contradiction detection is performed based on the multiple perception results. The semantic contradiction detection includes: spatial overlap conflict judgment between the target and static obstacles, conflict judgment between dynamic targets and non-drivable areas, and category mutual exclusion judgment of the same spatial coordinates. High uncertainty quantification stage: High uncertainty quantification is performed based on the prediction confidence of the multiple perception results. The high uncertainty quantification includes measuring classification fuzziness by target classification probability entropy and measuring detection reliability by target detection confidence. If either the semantic contradiction detection or the high uncertainty quantification condition is met, then the perception result is determined to have a semantic contradiction or high uncertainty.
5. The autonomous driving perception enhancement method based on a multimodal large language model according to claim 1, characterized in that, In step S3, the structured semantic guidance signal includes at least: object category, bounding box in BEV, uncertainty score, and attention weight generated based on the uncertainty score; The generation process of the structured semantic guidance signal includes: 1) Extract image features, point cloud geometric features, BEV global semantic features, and local target features from the original sensor data or intermediate features through multimodal input parsing; 2) Cross-modal feature fusion strengthens the association of information within a modality through self-attention, and constructs a cross-modal cross-attention mapping based on text features to achieve multimodal information alignment; 3) Predict object categories using a Softmax classifier, predict BEV bounding boxes using a regression head, calculate uncertainty scores based on category prediction probabilities and bounding box regression confidence, and apply the mapping formula: It generates attention weights and encapsulates them into a structured dictionary output to achieve structured field prediction.
6. The autonomous driving perception enhancement method based on a multimodal large language model according to claim 1, characterized in that, In step S3, the semantic guidance signal is converted into an attention mask corresponding to the BEV feature map space, specifically including: 1) Based on the grid size of the BEV feature map, create an initial mask matrix with an initial value of the first value; 2) Based on the grid resolution of the BEV feature map, map the physical coordinates of the BEV bounding box in the semantic guidance signal to the grid coordinates of the initial mask matrix; wherein, The formula for mapping the horizontal coordinate of the grid is: ; The formula for mapping grid ordinates is: ; 3) Replace the mask element values corresponding to the grid coordinate range with the attention weight values in the semantic guidance signal, and keep the element values outside the grid coordinate range as the first value. Perform linear interpolation on the edge grid of the BEV bounding box so that the weight gradually changes from the attention weight to the first value to form an attention mask. The attention weight value is greater than the first value.
7. The autonomous driving perception enhancement method based on a multimodal large language model according to claim 1, characterized in that, When performing spatial cross-attention calculation using the Transformer architecture, the attention score is weighted and modulated using the attention mask, specifically including: 1) The key matrix K, based on the BEV query matrix Q and the multimodal fusion features, is calculated using the following formula: Calculate the raw attention score ; 2) To adapt to parallel computation with multiple attention heads, the attention mask is... ( The dimension is adjusted to be compatible with the original attention score broadcast, denoted as... ; 3) Adjust the attention mask Compared with the original attention score Element-wise multiplication yields the modulated attention score, and the modulation formula is: = ⊙ To improve the attention response of the corresponding area within the grid coordinate range; 4) Perform Softmax normalization based on the modulated attention scores to obtain the modulated attention weights. ; 5) Modulate the attention weights The value matrix of the multimodal fusion features is weighted and aggregated to generate the semantically enhanced current frame BEV features.
8. An autonomous driving perception enhancement system based on a multimodal large language model, used to implement the autonomous driving perception enhancement method based on a multimodal large language model as described in any one of claims 1-7, characterized in that, include: Feature processing module: used to collect multi-sensor data and extract features, project the features onto the BEV space for alignment and fusion to obtain multimodal fusion features, and fuse historical frame BEV features with the BEV query matrix and perform temporal enhancement processing to generate current frame BEV features with enhanced environmental perception. The semantic enhancement module includes a monitoring submodule and a multimodal large language model submodule; among which, The monitoring submodule is used to detect whether there is semantic contradiction or high uncertainty in the perception result obtained based on the current frame BEV features. If not, the current frame BEV features are directly output. The multimodal large language model submodule is triggered when the monitoring submodule detects semantic contradictions or high uncertainty. It performs target detection and scene understanding on the original sensor data or intermediate features, generates a structured semantic guidance signal, converts the semantic guidance signal into an attention mask corresponding to the BEV feature map space, and uses the attention mask to weight and modulate the attention score when performing spatial cross-attention calculation, regenerates the semantically enhanced current frame BEV features and outputs them. Control interaction module: used to receive the current frame BEV features output by the feature processing module or the semantic enhancement module, and provide them to the downstream control module; The downstream control module is used to generate or adjust vehicle control commands based on the received BEV features of the current frame.
9. The autonomous driving perception enhancement system based on a multimodal large language model according to claim 8, characterized in that, The downstream control module is a controller based on optimization control, imitation learning, or reinforcement learning.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the autonomous driving perception enhancement method based on a multimodal large language model as described in any one of claims 1 to 7.