Anti-interference target detection method and system based on automatic driving scene multi-modal fusion

By adopting a multimodal fusion anti-interference target detection method in unmanned driving scenarios, combining image and lidar data, and utilizing temporal self-attention and caching mechanisms, the accuracy and reliability issues of occluded object detection are solved, achieving more efficient object detection.

CN120689585APending Publication Date: 2025-09-23WUXI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510535414.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In existing unmanned driving technologies, when an object is obscured, the detection method does not pay enough attention to the features of the obscured object, making it difficult to accurately detect the obscured part of the information, and is unable to fully explore the correlation between continuous time series data, affecting the accuracy and timeliness of detection.

Method used

An anti-interference target detection method based on multimodal fusion of autonomous driving scenarios is adopted. Image features are processed by combining semantic segmentation with instance segmentation. Temporal self-attention mechanism and caching mechanism are introduced. Combined with lidar point cloud data, two-dimensional view and bird's-eye view features are generated for target detection and segmentation.

Benefits of technology

It improves the detection accuracy and reliability of obscured objects, provides a more efficient and accurate object detection solution, and enhances the ability to capture the features of obscured objects and identify occlusion situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689585A_ABST
    Figure CN120689585A_ABST
Patent Text Reader

Abstract

The invention provides an anti-interference target detection method and system based on automatic driving scene multi-modal fusion, and the method comprises the steps: employing a semantic segmentation and instance segmentation combination method based on a camera image in an image view coding layer, and employing a Mask-RCNN model, and precisely judging a possible shielded object and region; in a feature processing link, a time self-attention mechanism is introduced, a feature map is weighted from a time dimension, an occluded object is focused, irrelevant information is inhibited, and the capability of capturing features of the occluded object is enhanced; besides, a caching mechanism is arranged in the model, a target confidence coefficient change method is adopted, the detection confidence coefficient of continuous video frames is stored, and the shielding condition is recognized by analyzing the confidence coefficient change condition of a target object in the continuous video frames. According to the invention, the detection precision and reliability of the shielded object are improved, and a more efficient and more accurate solution is provided for the object detection of the unmanned driving technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of partially obscured object detection in unmanned driving, and more specifically to an anti-interference target detection method and system based on multimodal fusion of autonomous driving scenarios. Background Art

[0002] Amid the booming development of autonomous driving technology, achieving safe and efficient driving relies on precise environmental perception. Accurate detection of surrounding objects is crucial. In complex and ever-changing real-world traffic scenarios, cameras, lidar, and other sensors work together but face numerous challenges. More critically, in dynamic traffic scenarios, the state of objects is constantly changing, and the movement of vehicles, pedestrians, and other objects causes occlusions to occur at any time and in a time-series manner. Existing methods, when utilizing continuous data, fail to effectively exploit the correlation between data at different times, and are unable to accurately track and detect objects at different stages of occlusion. This significantly impacts the accuracy and timeliness of autonomous driving system decisions, becoming a significant bottleneck hindering the widespread application of autonomous driving technology.

[0003] In the existing technology, there are problems in the unmanned driving scenario such as insufficient attention to the features of the occluded objects, difficulty in accurately detecting the information of the occluded parts, and inability to fully explore the correlation between continuous time series data. Summary of the Invention

[0004] In order to solve the problems of existing occluded object detection methods, such as insufficient attention to the characteristics of occluded objects, difficulty in accurately detecting the occluded part information, and inability to fully explore the correlation between time-series continuous data, the present invention proposes an anti-interference target detection method and system based on multimodal fusion of autonomous driving scenarios.

[0005] In order to achieve the above technical effects, the technical solutions of the present invention are as follows:

[0006] An anti-interference target detection method and system based on multimodal fusion in autonomous driving scenarios, comprising the following steps:

[0007] S1: Obtain a dataset, which includes multi-view images and lidar point cloud data;

[0008] S2: Multi-view image features are processed using a combination of semantic segmentation and instance segmentation to obtain image features. The image features are processed by adding a temporal self-attention mechanism module. The processed features are input into a cache mechanism module and then compressed into a two-dimensional bird's-eye view feature through three-dimensional coordinate projection.

[0009] S3: Generate 2D lidar bird's-eye view features from lidar point cloud data by voxelizing the point cloud and aggregating along the height dimension;

[0010] S4: Obtain fusion features based on the two-dimensional view bird's-eye view features and the two-dimensional lidar bird's-eye view features; input the fusion features into the target detection head module to obtain the target detection results and the segmentation results of the fusion features.

[0011] The present invention also provides an anti-interference target detection system based on multimodal fusion of autonomous driving scenarios, comprising:

[0012] A data set acquisition module is used to acquire a data set, which includes multi-view images and lidar point cloud data;

[0013] The 2D bird's-eye view feature acquisition module is used to process multi-view image features using a combination of semantic segmentation and instance segmentation to obtain image features. The image features are processed by adding a temporal self-attention mechanism module, and the processed features are input into the cache mechanism module. The features are then compressed into 2D bird's-eye view features after 3D coordinate projection.

[0014] The 2D LiDAR bird's-eye view feature acquisition module is used to generate 2D LiDAR bird's-eye view features by voxelizing LiDAR point cloud data and aggregating them along the height dimension.

[0015] The feature fusion and detection module is used to obtain fusion features based on the two-dimensional view bird's-eye view features and the two-dimensional lidar bird's-eye view features; the fusion features are input into the target detection head module to obtain the target detection results and the segmentation results of the fusion features.

[0016] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0017] At the image view coding level, this invention uses a combination of semantic segmentation and instance segmentation based on camera images, with the help of the Mask-RCNN model, to accurately determine possible occluded objects and areas. In the feature processing stage, a temporal self-attention mechanism is introduced to weight the feature map from the temporal dimension, focusing on occluded objects, suppressing irrelevant information, and enhancing the ability to capture their features. In addition, a caching mechanism is set up in the model, and a target confidence change method is used to store the detection confidence of consecutive video frames. Occlusion is identified by analyzing the confidence change of the target object in the consecutive video frames. This invention improves the detection accuracy and reliability of occluded objects and provides a more efficient and accurate solution for object detection in unmanned driving technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 The flowchart of the anti-interference target detection method based on multimodal fusion of autonomous driving scenarios is shown in an embodiment of the present invention.

[0019] Figure 2This is an architectural flow chart of an anti-interference target detection method based on multimodal fusion in autonomous driving scenarios, shown in an embodiment of the present invention.

[0020] Figure 3 This is a flow chart of the BevFusion framework data processing architecture according to an embodiment of the present invention.

[0021] Figure 4 This is a structural diagram of a view encoder module according to an embodiment of the present invention.

[0022] Figure 5 This is a feature pyramid network structure diagram shown in an embodiment of the present invention.

[0023] Figure 6 This is a structural diagram of a feature adaptation module according to an embodiment of the present invention.

[0024] Figure 7 This is a structural diagram of a lidar encoder module according to an embodiment of the present invention.

[0025] Figure 8 This is a structural diagram of the dynamic fusion module shown in an embodiment of the present invention.

[0026] Figure 9 This is a structural diagram of a detection head module according to an embodiment of the present invention.

[0027] Figure 10 This is an architecture diagram of an anti-interference target detection system based on multimodal fusion in autonomous driving scenarios, shown in an embodiment of the present invention. DETAILED DESCRIPTION

[0028] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.

[0029] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a," "the," and "the" used in this invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0030] It should be understood that although the terms "first," "second," "third," etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information, without departing from the scope of the present invention. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."

[0031] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0032] Example 1

[0033] This embodiment proposes an anti-interference target detection method based on multimodal fusion of autonomous driving scenarios, such as Figure 1 FIG. 1 is a flow chart of this embodiment.

[0034] S1: Obtain a dataset, which includes multi-view images and lidar point cloud data;

[0035] S2: Multi-view image features are processed using a combination of semantic segmentation and instance segmentation to obtain image features. The image features are processed by adding a temporal self-attention mechanism module. The processed features are input into a cache mechanism module and then compressed into a two-dimensional bird's-eye view feature through three-dimensional coordinate projection.

[0036] S3: Generate 2D lidar bird's-eye view features from lidar point cloud data by voxelizing the point cloud and aggregating along the height dimension;

[0037] S4: Obtain fusion features based on the two-dimensional view bird's-eye view features and the two-dimensional lidar bird's-eye view features; input the fusion features into the target detection head module to obtain the target detection results and the segmentation results of the fusion features.

[0038] In this embodiment, at the image view coding level, the present invention uses a combination of semantic segmentation and instance segmentation based on camera images, and with the help of the Mask-RCNN model, accurately determines the possible occluded objects and areas; in the feature processing link, a temporal self-attention mechanism is introduced to weight the feature map from the time dimension, focusing on the occluded objects, suppressing irrelevant information, and enhancing the ability to capture its features; in addition, a caching mechanism is set in the model, and a target confidence change method is adopted to store the detection confidence of continuous video frames, and the occlusion situation is identified by analyzing the confidence changes of the target objects in the continuous video frames.

[0039] Example 2

[0040] On the basis of Example 1, the technical effects of this solution are further demonstrated. Specifically:

[0041] The architecture flow chart of this embodiment is as follows Figure 2 As shown, this embodiment is built on the BevFusion framework, and its architecture flow chart is as follows Figure 3 shown.

[0042] In an optional embodiment, in step S1, the nuscenes dataset is used as the input data source; the multi-view images in the nuscenes dataset are collected by cameras surrounding the vehicle, and the point cloud data in the nuscenes dataset is obtained by a rotating lidar.

[0043] In an optional embodiment, the data set also includes detailed annotation information and map data of the city.

[0044] Furthermore, in complex urban traffic intersection scenarios, the present invention uses the NuScenes dataset as the input data source. The NuScenes dataset is a highly influential large-scale, comprehensive dataset published by nuTonomy in the field of autonomous driving. Its original purpose was to provide solid data support for various key tasks in autonomous driving, such as perception and decision-making.

[0045] The NuScenes dataset is massive, collected from Boston and Singapore, two cities with complex traffic conditions and challenging driving scenarios. It fully accounts for the differences in geography, traffic regulations, and environmental conditions. The dataset contains 1,000 20-second driving scenes, encompassing approximately 1.4 million camera images, 390,000 LiDAR scans, and 1.4 million radar scans.

[0046] This dataset covers two core data types: multi-view images and point cloud data. Multi-view images are collected synchronously at a 12Hz frequency by six cameras surrounding the vehicle. These images, with their different perspectives, capture the vehicle's driving environment from all angles. They not only present rich environmental textures but also clearly show details such as road signs, surrounding buildings, and other traffic participants, providing intuitive visual information for algorithms.

[0047] Point cloud data is acquired via a rotating lidar (Velodyne HDL 32E) with a 20Hz acquisition frequency. It has 32 channels, a 360° horizontal field of view, and a +10° to -30° vertical field of view. Each ring has approximately 1,080 (±10) points, a measurement range of 80 to 100 meters, an effective echo distance of up to 70 meters, and an accuracy of ±2 cm. It can generate up to approximately 1.39 million points per second, and its unique feature is its extremely precise spatial information. The laser beam emitted by the lidar interacts with surrounding objects and then reflects back. Based on information such as the reflection time, a three-dimensional point cloud model of the target object can be constructed. This point cloud data can accurately outline the target object and clarify the positional relationship between objects, clearly presenting pedestrians, vehicles nearby, and obstacles in the distance.

[0048] In addition to multi-view images and point cloud data, the dataset also includes detailed annotation information. For 23 target object categories, accurate 3D bounding boxes are annotated at a frequency of 2Hz throughout the dataset. Object attributes such as visibility, activity state, and posture are also annotated. Furthermore, the dataset provides city map data with rich details, including road layout, traffic sign locations, lane information, etc., to assist autonomous driving systems in localization and path planning tasks.

[0049] In this example, the nuscenes dataset was selected, which contains multi-view images and point cloud data. The multi-view images provide rich texture details, while the point cloud data provides precise spatial information. The combination of the two provides a comprehensive data foundation for detecting occluded objects. The multi-view images reveal the appearance of objects, providing clues even when partially occluded. The point cloud data outlines the object's contours and positional relationships, helping to determine the actual location of occluded objects and improving detection accuracy and reliability.

[0050] In an optional embodiment, the step S2 of processing the multi-view image features using a combination of semantic segmentation and instance segmentation to obtain image features includes the following steps:

[0051] S21: input the multi-view image into the view encoder module to obtain multi-view image features;

[0052] S22: Use the Mask-RCNN algorithm to perform image segmentation and enhancement on multi-view image features and generate masks of occluded areas;

[0053] S23: Input the multi-view image features corresponding to the occluded area mask into the temporal self-attention mechanism module based on the Transformer architecture to obtain features processed by the temporal self-attention mechanism;

[0054] S24: storing the features processed by the temporal self-attention mechanism into the cache mechanism module, and obtaining the features after the target confidence level is changed by calculating the confidence level change and comparing it with the preset confidence level change threshold;

[0055] S25: Convert the 2D image coordinates in the feature into 3D ego-vehicle coordinates through the view projector module;

[0056] S26: Input the converted three-dimensional coordinate data into the bird's-eye view encoder module to obtain the two-dimensional view bird's-eye view features.

[0057] In an optional embodiment, the view encoder module further includes a feature pyramid network and a feature adaptation module; the feature pyramid network fuses features of different scales through top-down and lateral connections; the feature adaptation module upsamples multi-scale features through upsampling operations, adaptive average pooling operations, and convolution layers, connects the sampled features together, and then passes them through convolution layers to obtain image features.

[0058] In an optional embodiment, the stored continuous multi-frame detection confidences are C1, C2, ..., C T , T represents the cache queue length, and the confidence change is calculated as follows:

[0059] ΔC n =max(|C n -C n-1 |,|C n -C n+1 |)

[0060] Where, ΔC n Represents the confidence change, and n represents the index of the current frame.

[0061] Furthermore, in step S21, after obtaining multi-view images suitable for complex urban traffic intersection scenes from the nuscenes dataset, the multi-view images first enter the view encoder module, namely the image view encoder module, whose structure is as follows: Figure 4 This module uses the Dual-Swin-Tiny backbone network, a lightweight network based on the Swin Transformer architecture. It boasts efficient feature extraction capabilities and low computational complexity, making it suitable for image feature extraction tasks in resource-constrained scenarios. In complex urban traffic intersections, large amounts of image data require rapid processing. The Dual-Swin-Tiny backbone network effectively extracts key features from images, such as those of vehicles, pedestrians, and traffic signs, while maintaining computational efficiency.

[0062] This module is also paired with a feature pyramid network (FPN), whose structure is as follows Figure 5The Feature Pyramid Network (FPN) is a network structure used for multi-scale feature fusion. It fuses feature maps of different scales through top-down and lateral connections, enabling the model to simultaneously utilize high-resolution detail information and low-resolution semantic information, thereby enhancing the detection capability of objects of different sizes.

[0063] In addition, the module contains a simple feature adaptation module (ADP), whose structure is as follows Figure 6 For example, for each view image with an input shape of H×W×3 (the “3” here usually corresponds to the three color channels of an RGB image), after passing through the backbone network and the feature pyramid network, multi-scale features F2, F3, F4, and F5 will be output, and their shapes are respectively (“C” here refers to the number of channels of these feature maps. The number of channels will change according to the network structure and calculation process. Different channels can capture different types of feature information.) The feature adaptation module then upsamples the multi-scale features to The shape is then connected together after sampling, and then passed through a 1×1 convolution layer to obtain a shape of image features, so as to adaptively fuse features of different scales and enhance the model's ability to capture features of objects of different sizes.

[0064] Furthermore, in step S22, the Mask-RCNN algorithm is developed based on Faster R-CNN. By adding a branch for predicting instance masks, it enables instance segmentation during object detection. It excels in both object detection and instance segmentation, and is particularly suitable for object recognition in complex scenes. The Mask-RCNN workflow primarily includes the following key steps: backbone network feature extraction, region proposal network (RPN) candidate box generation, region of interest (RoI) pooling, classification, regression, and mask generation.

[0065] At complex urban traffic intersections, multi-view cameras continuously capture image information from all angles of the intersection. These images contain a wealth of information, such as vehicles waiting for traffic lights, pedestrians crossing the road, and roadside traffic facilities, and are often obstructed. When these images are fed into the Mask-RCNN algorithm, the backbone network processes them. For example, at a densely populated intersection, the backbone network can extract low-level features such as vehicle edges and textures, as well as high-level features such as the vehicle's overall shape and category. Based on these feature maps, the Region Proposal Network (RPN) generates a series of candidate boxes that may contain objects. In intersection scenarios, the RPN quickly identifies areas where targets such as vehicles and pedestrians may be present. Because candidate boxes vary in size and position, the region of interest pooling operation aligns and samples the corresponding feature map areas, converting them to a fixed size for easier processing. During the classification, regression, and mask generation stages, the classification branch accurately determines whether the candidate box contains a vehicle, pedestrian, or other object. The regression branch fine-tunes the candidate box's position and size to more accurately enclose the object. The mask branch generates precise instance masks. For example, when multiple vehicles are occluding each other at an intersection, the mask branch can generate a separate mask for each vehicle, clearly defining its boundaries.

[0066] However, actual autonomous driving scenarios are full of challenges. Factors such as lighting changes, object occlusion, and varying weather conditions can affect algorithm performance. To improve the Mask-RCNN model's ability to segment occluded objects in complex scenarios, this example uses the following data augmentation techniques during training:

[0067] Lighting:

[0068] Brightness Adjustment: Randomly adjust the image brightness within a certain range. Specifically, the brightness coefficient is set between [0.5, 1.5] to simulate scenes with varying light intensities. For example, in the early morning or evening, when light intensity is low, lowering the brightness coefficient can make the training images appear similar to low-light conditions, allowing the model to learn the characteristics of objects under these conditions. In bright midday environments, raising the brightness coefficient can simulate overly bright scenes, enhancing the model's adaptability to varying light intensities.

[0069] Contrast Transformation: Randomly adjusts the image contrast, with the contrast coefficient ranging from 0.8 to 1.2. This helps the model learn the differences in object features under different contrast conditions. For example, on cloudy days, image contrast is low. By reducing the contrast coefficient to simulate such a scene, the model can accurately extract object features under such blurred visual conditions.

[0070] Saturation Modification: Randomly modulates the image saturation, with the saturation coefficient set between [0.7, 1.3]. In real-world scenarios, different weather and environmental conditions can affect image saturation; for example, rainy days can reduce image saturation. This data augmentation allows the model to adapt to images with varying saturation levels and better identify occluded objects.

[0071] Geometric aspects:

[0072] Random Flip: Set the horizontal and vertical flip probabilities of the random flip operation to 0.5. This operation simulates observation scenes from different angles, allowing the model to learn the characteristics of objects in different flipped states, enhancing its ability to recognize objects from multiple angles. For example, if a vehicle originally traveling forward is horizontally flipped, the model can learn the characteristics of its reverse state, allowing it to accurately recognize similar situations in real-world scenarios.

[0073] Rotation: The rotation angle range is controlled within [-15°, 15°], with a step size of 1°. This allows the model to learn about the changes in object features after rotation within a certain angle range, adapting to the differences in features of objects at different angles in real-world scenes. For example, pedestrians may have different postures and angles while walking. By learning these rotated image features, the model can better identify pedestrians at various angles.

[0074] Adding Gaussian noise: Adding Gaussian noise with a mean of 0 and a standard deviation of 0.01 simulates image noise interference in real-world scenarios. This helps the model learn to extract object features in noisy environments and improves its robustness to noise. For example, in rainy or low-light environments, images captured by the camera may contain noise. Models trained with Gaussian noise can more reliably recognize objects.

[0075] Through these data augmentation techniques, the training dataset has been significantly enriched. Training on diverse data allows the model to learn a wider range of object features, effectively enhancing its generalization capabilities and enabling it to better identify occluded objects in various complex scenarios in real-world applications.

[0076] Furthermore, in step S23, at complex urban traffic intersections, the movement of vehicles and pedestrians causes the occlusion of objects to change at any time. The attention mechanism module plays a key role in this scenario. Suppose that at an intersection, a car is gradually blocked by a bus in front while driving. The self-attention mechanism based on the Transformer architecture processes feature maps at different time steps. For this blocked vehicle, the feature maps at different time steps are linearly transformed to obtain query, key and value vectors, and then the attention score and weight matrix are calculated. Through such operations, the model can capture the correlation between feature maps at different times before the vehicle is blocked, during the occlusion process, and when it is briefly exposed, enhance the coherence and importance of features related to vehicle occlusion in the time series, suppress irrelevant feature information, and thus more accurately judge the state and position of the vehicle.

[0077] Therefore, specifically in the temporal dimension, a self-attention mechanism based on the Transformer architecture is introduced. Assume that the input feature map sequence is {F1, F2, ... F T}, where T represents the time step. First, for each time step feature map F t ∈R C ×H×W (C is the number of channels, H is the height, and W is the width), and the query, key, and value vectors are obtained through linear transformation. Let the query matrix be WQ, the key matrix be WK, and the value matrix be WV, then: the query vector Q t =F t W Q , key vector K t =F t W K , value vector V t =F t W V Then calculate the attention scores between different time steps and obtain the attention score matrix A through dot product operation: Where dk is the dimension of the key vector, i,j∈{1,2,…,T}. Then the attention score matrix A is normalized by the Softmax function to obtain the normalized attention weight matrix Finally, according to the normalized attention weight matrix Perform weighted summation on the value vector to obtain the feature map processed by the temporal attention mechanism

[0078] Furthermore, in step S24, a caching mechanism is crucial for detecting occluded objects at complex urban intersections. In this model, a dedicated caching mechanism is implemented to store detection results from consecutive video frames. This caching mechanism utilizes a queue data structure. After extensive experiments and in-depth analysis of complex traffic scenarios, the queue length was set to a range of 10-15 frames. This range was chosen because in real-world autonomous driving scenarios, particularly in complex areas like large transportation hubs and dense urban intersections, object motion trajectories are extremely complex, and occlusions are frequent and long-lasting. For example, in large transportation hubs, with vehicles and pedestrians constantly traversing, an object may experience multiple, long, and complex occlusions, with brief re-occlusions and re-occlusions. If the caching mechanism is too short, such as only 3-5 frames, it is highly likely that the confidence level of an object will not be fully captured throughout the occlusion process, leading to inaccurate judgment of its occlusion status. A length of 10-15 frames, on the other hand, provides a richer and more comprehensive confidence level trajectory, providing sufficient evidence for accurately determining an object's occlusion stage and status. This interval length can fully adapt to the needs of complex scenarios, but will not be too long to cause excessive consumption of computing resources and excessive storage burden.

[0079] Assume that the stored continuous multi-frame detection confidences are C1, C2, ..., C T (T is the cache queue length,

[0080] 10≤T≤15), define the confidence change ΔC n (n represents the index of the current frame, 1<n<T):

[0081] ΔC n =max(|C n -C n-1 |,|C n -C n+1 |)

[0082] This formula compares the absolute values ​​of the detection confidence differences between the current frame, the previous frame, and the next frame, and takes the maximum value as the confidence change. This more comprehensively reflects the changes in the confidence of the current frame, avoiding ignoring the actual trend of change by only considering the difference with a single frame.

[0083] By introducing the target confidence change method, we conduct an in-depth analysis of the detection confidence recorded by the cache mechanism. When a specific condition is met, that is, the confidence change exceeds the confidence change thresholds threshold1 and threshold2 determined by the experiment (threshold1<ΔC n<threshold2), and part of the object is not detected in the current frame (taking the nth frame as an example, the value range of n is within the frames covered by the cache queue), which can be detected by comparing the complete mask M of the object. full It is judged by the mask area Mn detected in the current frame. Assume that the complete mask area of ​​the object is S full , the mask area detected in the current frame is Sn, when it satisfies:

[0084]

[0085] It can be inferred that the object is partially occluded in the current frame.

[0086] For example, in a continuous 15-frame monitoring of a complex traffic scene, a pedestrian is continuously detected. The confidence level is 0.9 in the 5th frame, drops sharply to 0.2 in the 6th frame, and rises back to 0.6 in the 7th frame. The confidence level also fluctuates in subsequent frames. First calculate the confidence level change in the 6th frame, ΔC6 = max(|0.2-0.9|,|0.2-0.6|) = 0.7. After comparative analysis of the mask area, it is found that the mask area of ​​the pedestrian in the 6th frame is reduced by 25% compared to the complete mask area, that is, it meets the requirements. Combined with the significant drop in confidence (0.7 exceeds the corresponding threshold), it can be determined that the pedestrian is highly likely occluded in frame 6. This approach, leveraging information from multiple frames in the cache mechanism, allows for more accurate identification of object occlusion in dynamic scenes, providing more reliable data support for autonomous driving system decision-making.

[0087] Furthermore, in step S25, the view projector module plays an indispensable role at complex urban traffic intersections. The cameras at the intersection capture two-dimensional images, while the lidar obtains three-dimensional point cloud data, and the two have different coordinate systems. The view projector module uses the camera's internal and external parameters, as well as the vehicle's position information, to achieve coordinate conversion. For example, when an unmanned vehicle is driving at an intersection, the view projector module uses internal parameter data such as the camera's focal length and optical center position, combined with external parameter data such as the camera's installation position and posture on the vehicle, to map each pixel in the image to a three-dimensional spatial coordinate system centered on the vehicle, completing the conversion from two-dimensional image coordinates to three-dimensional ego-vehicle coordinates, laying the foundation for the subsequent generation of bird's-eye view (BEV) features.

[0088] Furthermore, in step S26, the data processed by the view projector module then enters the bird's-eye view encoder module. The bird's-eye view encoder module is designed to further extract and encode these converted coordinate data to generate a feature map with a bird's-eye view perspective. In complex urban traffic intersection scenarios, the bird's-eye view encoder module contains multiple convolutional layers and pooling layers. The spatial features and semantic information in the data are captured through convolution operations. Taking a series of 3×3 convolutional layers as an example, the convolution kernel slides on the data to extract features from local areas, which can effectively capture detailed information such as the edges and shapes of objects; the pooling layer is used to reduce the data dimension and reduce the amount of calculation while retaining key features. During the encoding process, features of objects of different positions and categories will be enhanced according to the characteristics of the unmanned driving scene. For object areas that may be obscured, the intensity of feature extraction will be increased to retain more details and improve the accuracy of subsequent detection.

[0089] In this embodiment, Dual-Swin-Tiny is used as the backbone network, in conjunction with FPN and ADP modules. Dual-Swin-Tiny efficiently extracts key features, FPN fuses multi-scale features, and ADP adaptively fuses features of different scales, enhancing the ability to capture features of objects of different sizes. This allows for more accurate feature extraction in complex intersection scenarios, even if an object is partially occluded, providing richer and more accurate information for subsequent occlusion detection. In complex urban traffic intersections, the Mask-RCNN algorithm extracts image features through the backbone network, RPN generates candidate frames, and the region of interest pooling operation unifies the size of the candidate frames. Classification, regression, and mask generation determine the object category, position, and mask. In the case of object occlusion, the mask branch generates an independent mask for each object, clearly defines the boundary, and accurately determines whether the object is occluded and the occluded area. Data enhancement technology improves the model's ability to segment occluded objects in complex scenes. The self-attention mechanism based on the Transformer architecture processes feature maps at different time steps. Vehicles and pedestrians move frequently, and object occlusion situations vary. This mechanism captures the correlation between feature maps at different times before an object is occluded, during occlusion, and when it is briefly exposed. This enhances the coherence and importance of occlusion-related features in the time series, suppresses irrelevant information, and more accurately determines the state and position of an object, improving the accuracy of occluded object detection. The cache mechanism stores detection results from consecutive video frames, with a queue length set to 10-15 frames. Objects at intersections have complex motion trajectories, and occlusions are frequent and long-lasting. This length fully captures the confidence changes of objects during occlusion. By defining the confidence change amount, combining the confidence change threshold with the object mask area comparison, it accurately determines whether an object is occluded, providing reliable data support for autonomous driving system decision-making. The view projector module uses camera intrinsic and extrinsic parameters and vehicle pose information to convert 2D image coordinates into 3D ego-vehicle coordinates. Intersection cameras capture 2D images, while lidar acquires 3D point cloud data. These two systems have different coordinate systems. Coordinate conversion lays the foundation for the subsequent generation of bird's-eye view (BEV) features, enabling the fusion of data from different sensors in a unified coordinate system to more accurately detect the position and state of occluded objects. The convolutional and pooling layers of the BEV encoder module process the converted data. The convolutional layers capture spatial and semantic features, while the pooling layers reduce the data dimensionality, enhancing feature extraction in areas potentially subject to occlusion. The generated BEV features share the same perspective and dimensionality as the lidar BEV features, providing a unified representation for multimodal data fusion and improving the accuracy of occluded object detection.

[0090] In an optional embodiment, step S3 specifically includes the following steps:

[0091] S31: Obtain the lidar point cloud data features through the lidar encoder module;

[0092] S32: Input the lidar point cloud data features into the bird's-eye view encoder module, and based on the voxelized features, aggregate them along the height dimension through the convolution layer to generate two-dimensional lidar bird's-eye view features.

[0093] Furthermore, in step S31, the laser radar point cloud data is processed by the laser radar encoder module, the structure of which is as follows: Figure 7 As shown in Figure 2, the LiDAR encoder module consists of two main parts: voxelization and backbone network.

[0094] In the voxelization stage, the point cloud data is processed by VoxelNet (a neural network model that converts point cloud data into voxel representation for processing, which can effectively use the spatial structure information of the point cloud for feature extraction) to convert the point cloud into voxel representation and then extract the lidar features. In VoxelNet, after dividing the point cloud into voxels, the features of the points in each voxel are aggregated. Assuming that there are N points in a voxel, the feature representation of each point is (i=1,…,N, Dp is the point feature dimension), and the features of these points are aggregated through a multi-layer perceptron (MLP) to obtain voxel features. For a simple MLP structure, including a layer of linear transformation and activation function (such as ReLU), the calculation formula is:

[0095] h i =ReLU(W1p i +b1)

[0096] in is the weight matrix, is the bias vector, and Dh is the hidden layer dimension. Then the intermediate features hi of all points in the voxel are average pooled to obtain the voxel feature v, as follows:

[0097]

[0098] This formula aggregates the features of multiple points into one voxel feature by averaging all intermediate features within the voxel, so that the voxel can represent the comprehensive features of the points inside it.

[0099] After completing the above operations, VoxelNet uses convolution operations to further extract features. Here, a 3×3×3 three-dimensional convolution kernel is used. Assuming that the convolution kernel is K, the convolution operation is performed on the voxel feature v. The calculation formula for the output feature vout is: Here, (m, n, l) represents the position of the output feature in voxel space. This formula indicates that during the convolution operation, the convolution kernel K slides over the voxel feature v in a 3×3×3 window, performing a weighted summation of the voxel features within each window to obtain the new output feature vout. This approach effectively captures the spatial characteristics of the point cloud. In complex urban traffic intersection scenarios, point cloud data contains a large amount of information such as vehicles, pedestrians, and surrounding buildings. These VoxelNet operations can effectively extract the features of the target object. Even if the object is partially obscured, its existence can be detected through the spatial structure of the point cloud.

[0100] The backbone network uses a sparse encoder. In complex urban traffic intersection scenarios, point cloud data is sparse, meaning that most spaces are devoid of point cloud data. The sparse encoder effectively processes this sparse data, avoiding the computational waste associated with traditional encoders. The sparse encoder performs feature extraction and transformation on the voxelized data through a series of sparse convolution operations. Sparse convolutions only operate on non-zero elements, significantly reducing computational effort. For example, in open areas at intersections, where point cloud data is very sparse, the sparse encoder can quickly skip these areas and perform detailed feature extraction only on areas containing target objects, improving computational efficiency. Furthermore, the sparse encoder continuously extracts and abstracts features from the point cloud data through multiple layers of convolution and pooling operations, ultimately generating representative lidar features. These features are then fused with features extracted from multi-view images for subsequent occluded object detection.

[0101] Furthermore, in step S32, at complex urban intersections, the LiDAR continuously scans the surrounding environment, acquiring a large amount of point cloud data. After step S31, this data is further processed by the bird's-eye view encoder module to generate LiDAR bird's-eye view features (LiDAR BEV Features). This process is crucial for improving data utilization efficiency and enhancing detection accuracy.

[0102] The bird's-eye view encoder module contains multiple convolutional layers, with a series of 3×3 convolutional layers as the core component, which deeply processes the LiDAR features to enhance the feature expression capability.

[0103] Convolution kernel settings: The size of each 3×3 convolution kernel is fixed to 3×3×Cin, where Cin represents the number of channels in the input feature map. The convolution kernel acts as an information filter, performing a convolution operation on the input feature map in the form of a sliding window. For example, when the input feature map is of size H×W×Cin (H represents height and W represents width), each sliding convolution kernel performs element-by-element multiplication and accumulation operations on a local area of ​​3×3×Cin, thereby generating an element of the output feature map. This operation method can carefully capture the local spatial information in the lidar features and accurately extract key features.

[0104] Stride and padding strategy: To ensure the integrity and continuity of feature information, the stride is set to 1. This allows the convolution kernel to move only one pixel horizontally and vertically at a time, preserving the details of the input feature map to the greatest extent possible and avoiding information loss due to excessive stride. At the same time, zero padding (Padding = 1) is used to add a circle of zero-valued pixels to the edge of the input feature map. This ensures that the size of the output feature map is consistent with the input feature map, ensuring seamless connection between layers while preventing boundary information from being ignored during the convolution process.

[0105] Channel Adjustment Mechanism: The 3×3 convolutional layer has the ability to flexibly adjust the number of channels. In practical applications, the first 3×3 convolutional layer converts the input Cin channel feature map into a Cout1 channel feature map. Subsequent convolutional layers continue to adjust the number of channels accordingly, such as converting Cout1 to Cout2. The adjustment of the number of channels depends on the model complexity and task requirements. For example, in complex scenarios, to obtain richer feature information, the number of channels may be gradually increased, from an initial 32 channels to 64 or even 128 channels through multiple layers of convolution.

[0106] Activation function enhances nonlinearity: After each 3×3 convolutional layer, a ReLU activation function is applied. The ReLU function expression is f(x) = max(0,x). It introduces nonlinear transformations into the model, effectively solving the vanishing gradient problem, accelerating model training, and enabling the model to learn more complex and discriminative feature representations, thereby significantly improving the expressive power of LiDAR features. Through the progressive processing of these 3×3 convolutional layers, LiDAR features are deeply mined and enhanced from a bird's-eye view perspective, providing a solid data foundation for subsequent multimodal fusion and occluded object detection.

[0107] LiDAR features are processed through a series of 3×3 convolutional layers. Convolution operations at different levels extract feature information at different levels. Shallow convolutional layers primarily capture local, low-level features such as edges and corners. As the number of convolutional layers increases, deeper layers learn more abstract and advanced semantic features, such as the overall shape and category of an object. The combination of these different levels of feature information significantly enhances the expressive power of LiDAR features from a bird's-eye view perspective, providing richer and more accurate feature representations for subsequent multimodal fusion and object detection tasks.

[0108] After processing by the bird's-eye view encoder module, the lidar bird's-eye view features are ultimately generated. These features have a unified bird's-eye view perspective and match the image bird's-eye view features in terms of dimension and perspective, facilitating subsequent effective multimodal fusion, thereby improving the autonomous driving system's ability to perceive and understand the surrounding environment.

[0109] In this embodiment, point cloud data is processed through VoxelNet voxelization and a sparse encoder. VoxelNet uses the spatial structural information of point clouds to extract features. Even if an object is partially occluded, its existence can be detected based on the spatial distribution of the point cloud. The sparse encoder effectively processes sparse data, reducing the waste of computing resources. At the same time, it extracts abstract features through sparse convolution and multi-layer operations, improving the efficiency and accuracy of detecting occluded objects. The 3×3 convolution layer of the bird's-eye view encoder module deeply processes LiDAR features. The convolution kernel setting, stride and padding strategy, channel number adjustment mechanism, and the use of ReLU activation function enable LiDAR features to be deeply mined and enhanced from a bird's-eye view perspective. Different layers of convolution operations extract features at different levels. The shallow layer captures local low-level features, and the deep layer learns abstract high-level semantic features, providing a solid data foundation for multimodal fusion and occluded object detection.

[0110] In an optional embodiment, step S4 specifically includes the following steps:

[0111] S41: A dynamic fusion method based on channel attention is used to fuse the 2D lidar bird's-eye view features and the 2D view bird's-eye view features to output features for detection;

[0112] S42: Input the features used for detection into the target detection head module to obtain the target detection result and the segmentation result of the fused features.

[0113] Furthermore, the pre-processed lidar bird's-eye view features are fused with the camera image bird's-eye view features, a crucial step in fully leveraging the advantages of multi-sensor data. At complex urban intersections, the lidar bird's-eye view features contain precise three-dimensional spatial structure information about the environment, accurately reflecting the position and outline of objects; while the camera image bird's-eye view features provide rich texture and semantic details, helping to identify object categories.

[0114] In order to effectively fuse these two features, a dynamic fusion method based on channel attention is adopted. Inspired by the squeeze incentive mechanism, the fusion process can be expressed as:

[0115] F fused =f adaptive (f static ([F Camera ,F LiDAR ]))

[0116] Among them, [·,·] represents the concatenation operation of the camera image bird's-eye view feature (FCamera) and the lidar bird's-eye view feature (FLiDAR) along the channel dimension. fstatic is a static channel and spatial fusion function implemented by a 3×3 convolutional layer, which is used to reduce the channel dimension of the concatenated features. Assuming that the number of channels of the lidar bird's-eye view feature map is CLiDAR and the number of channels of the camera image bird's-eye view feature map is CCamera, the number of feature channels will change after concatenation, and the dimension will be adjusted after the fstatic operation. For the input features The calculation expression of fadaptive is:

[0117] f adaptive (F)=σ(Wf avg (F))·F

[0118] Where W represents a linear transformation matrix (e.g., a 1×1 convolution), favg represents global average pooling, and σ represents the sigmoid function. This fusion approach uses a simple channel attention module to select important fusion features, combining the strengths of both sensor data to provide more comprehensive information for subsequent detection. This fusion feature combines the strengths of both sensor data, providing more comprehensive information for subsequent detection.

[0119] Dynamic fusion module (structural perspective): The fused features enter the dynamic fusion module, and its structure is as follows Figure 8 As shown in the figure, the module sequentially undergoes 3×3 convolution, global average pooling, 1×1 convolution and other operations to deeply process the fusion features and output the final detection results.

[0120] 3×3 convolutional layer: Assume that the input feature map is The output feature map is The convolution kernel is Bias The convolution calculation process is:

[0121]

[0122] Where i = 0,…,H′-1; j = 0,…,W′-1; k = 0,…,C out -1, where it is assumed that the convolution stride is 1 and the padding is 1 to keep the feature map size unchanged (H′=H, W′=W). In the specific implementation, the number of convolution kernels in the 3×3 convolution layer can be set according to the dimension of the feature and the computing resources, for example, to 64. These 64 convolution kernels can be regarded as 64 different feature detectors, each of which focuses on capturing specific local features in the fused features, such as the edges, corners, and texture details of objects. Different convolution kernels learn different features, and they work together to extract rich local information from the fused features, providing more valuable data for subsequent global information aggregation and dimensionality adjustment.

[0123] Global average pooling: After extracting local features through the 3×3 convolutional layer, a global average pooling operation is performed. This step converts the feature map into a fixed-length feature vector and averages the feature map in the spatial dimensions (height and width). Assuming the input feature map is of size H×W×C, after global average pooling, the feature map of each channel will be compressed into a single value, that is, a feature vector of length C is obtained. This operation can aggregate global information, reduce feature dimensions, and highlight the overall feature distribution. It prevents the model from over-focusing on local details and ignoring overall features, allowing the model to have a clearer grasp of the overall characteristics of the object and enhancing the model's robustness to changes in different scales and positions.

[0124] 1×1 convolution layer (adjusting feature dimensions): The final 1×1 convolution operation is performed, which mainly adjusts the feature dimensions. The size of the 1×1 convolution kernel is 1×1×C in , which can change the number of channels through convolution operation without changing the spatial size (height and width) of the feature map. For example, if the number of channels of the input feature vector is C in , through the 1×1 convolution layer, the number of channels can be adjusted to C according to the needs of subsequent processing out In this detection method, by reasonably adjusting the number of channels, the model can process fusion features more effectively, further optimize feature representation, and enhance the model's ability to detect occluded objects.

[0125] In this embodiment, a dynamic fusion method based on channel attention is used to fuse lidar bird's-eye view features and camera image bird's-eye view features. The lidar features provide precise three-dimensional spatial structure information, while the camera image features provide rich texture and semantic details. This combination of features combines the advantages of both sensor data. By applying the channel attention module, important fusion features are selected, providing more comprehensive information for subsequent detection and improving the accuracy of detecting occluded objects. The 3×3 convolutional layer of the dynamic fusion module extracts local features, global average pooling aggregates global information, and the 1×1 convolutional layer adjusts the feature dimension. The multiple convolution kernels of the 3×3 convolutional layer capture the specific local features of the fused features, global average pooling highlights the overall feature distribution, and the 1×1 convolutional layer optimizes the feature representation, enhancing the model's ability to detect occluded objects and making the detection results more accurate.

[0126] In an optional embodiment, the target detection head module, i.e., the head module, includes a target detection branch and a bird's-eye view segmentation branch; the target detection branch includes a SeparateHead class module, a DCNSeparateHead class module, a CenterHead class module and a first loss calculation and target generation module; the bird's-eye view segmentation branch includes a BEVGridTransform class module, a classifier module and a second loss calculation and target generation module.

[0127] Furthermore, the features output by the dynamic fusion module enter the target detection head module. The structure of the target detection head module is shown in the figure below. Figure 9 This module focuses on the target detection part, and its output is center (center coordinate), dim (target size), height (target height), rot (rotation angle), and vel (speed).

[0128] This module is divided into two branches, one for 3D object detection and the other for bird's-eye view segmentation. The following will give a detailed introduction to these two branches:

[0129] 3D object detection branch:

[0130] It mainly consists of the SeparateHead class module, the DCNSeparateHead class module, the CenterHead class module, and the first loss calculation and target generation module. It adopts an architecture based on convolutional neural networks (CNN) and deep learning frameworks (such as PyTorch). The core structure includes the following parts:

[0131] SeparateHead class: This is a separate head class used to handle different object detection tasks. It receives the input feature map, extracts features through multiple convolutional layers, and generates different outputs for each task (such as heatmaps, regression values, scales, rotation angles, etc.). Each task (such as object localization and regression) has its own independent processing flow.

[0132] DCNSeparateHead class: This is an extension of SeparateHead that adds a deformable convolutional network (DCN) to better adapt to features and improve the performance of classification and regression tasks. Through the DCN layer, the feature map is adaptively adjusted to facilitate prediction of classification and regression tasks.

[0133] CenterHead class: This is the core of the model, responsible for fusing features from different heads to generate a predicted 3D bounding box. It includes multiple subtasks, such as object classification (heatmap prediction) and object location (regression), and calculates a loss function to optimize the model. During inference, it also uses post-processing (such as non-maximum suppression, NMS) to refine the final detection results.

[0134] Loss calculation and target generation: During training, the model generates training targets and uses Gaussian focal loss and L1 regression loss to calculate the losses for classification and regression tasks. During the inference phase, NMS is used to generate the final bounding box to ensure the validity of each detection result.

[0135] In summary, this branch handles different tasks through multiple heads (SeparateHead, DCNSeparateHead and CenterHead), and gradually completes the 3D object detection task through steps such as convolutional layers, feature adaptation, regression and classification loss calculation, and NMS.

[0136] Bird's-eye view split branch:

[0137] This branch defines a head BEVSegmentationHead for bird's-eye view (BEV) segmentation, which mainly consists of three parts: first, the BEVGridTransform class, which implements the spatial transformation of the input feature map, converting the input image from one coordinate system to another; then the classifier part, which processes the transformed feature map through multiple convolutional layers and activation functions and outputs the segmentation results for each category; finally, the loss calculation part, which calculates the loss of each category through cross entropy loss (xent) and focal loss (focal) to optimize the model during training. During the training phase, the model calculates and returns the loss of each category; while in the inference phase, the model directly returns the category prediction after sigmoid activation.

[0138] Overall, this branch implements the transformation, classification, and loss calculation of BEV feature maps, which is suitable for multi-class object segmentation tasks.

[0139] Regarding the choice of loss function, L1 loss is used for regression tasks (including regression of target size, target height, rotation angle, and speed). L1 loss measures the absolute error between the predicted value and the true value, helping the model to more accurately approximate the true value in regression tasks.

[0140] Focal Loss is used for category classification tasks. It is designed to address the category imbalance problem in object detection. It dynamically adjusts the weights of samples from different categories during training, allowing the model to focus more on samples from the minority category, thereby improving detection performance for objects of different categories.

[0141] For the heatmap part, that is, the prediction of the center coordinates, Gaussian Focal Loss is used. Gaussian Focal Loss enhances the positioning of the center coordinate prediction by introducing Gaussian distribution, which can better process the location information of the target in the image and improve the accuracy of the center coordinate prediction.

[0142] In this embodiment, the object detection head module uses L1 loss, focal loss, and Gaussian focal loss. When detecting partially occluded vehicles, L1 loss enables the model to more accurately approximate the true value in regression tasks, focal loss addresses class imbalance, and Gaussian focal loss improves the accuracy of center coordinate prediction. Ultimately, it outputs detailed vehicle information, providing a key basis for autonomous vehicle decision-making and control at intersections, ensuring safe driving.

[0143] Example 3

[0144] This embodiment proposes an anti-interference target detection system based on multimodal fusion of autonomous driving scenarios, and applies the anti-interference target detection method based on multimodal fusion of autonomous driving scenarios proposed in Example 1. Figure 10 The figure shows the architecture diagram of the system in this embodiment.

[0145] The system of this embodiment includes:

[0146] A data set acquisition module is used to acquire a data set, which includes multi-view images and lidar point cloud data;

[0147] The 2D bird's-eye view feature acquisition module is used to process multi-view image features using a combination of semantic segmentation and instance segmentation to obtain image features. The image features are processed by adding a temporal self-attention mechanism module, and the processed features are input into the cache mechanism module. The features are then compressed into 2D bird's-eye view features after 3D coordinate projection.

[0148] The 2D LiDAR bird's-eye view feature acquisition module is used to generate 2D LiDAR bird's-eye view features by voxelizing LiDAR point cloud data and aggregating them along the height dimension.

[0149] The feature fusion and detection module is used to obtain fusion features based on the two-dimensional view bird's-eye view features and the two-dimensional lidar bird's-eye view features; the fusion features are input into the target detection head module to obtain the target detection results and the segmentation results of the fusion features.

[0150] Each embodiment of the present invention is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiment. The device embodiment described above is merely exemplary. The modules described as separate components may or may not be physically separated. When implementing the scheme of the present invention, the functions of each module can be implemented in the same one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the scheme of this embodiment.

[0151] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. An anti-interference target detection method based on multimodal fusion in autonomous driving scenarios, characterized by: The following steps are involved: S1: Obtain a dataset, which includes multi-view images and lidar point cloud data; S2: Multi-view image features are processed using a combination of semantic segmentation and instance segmentation to obtain image features. The image features are processed by adding a temporal self-attention mechanism module. The processed features are input into a cache mechanism module and then compressed into a two-dimensional bird's-eye view feature through three-dimensional coordinate projection. S3: Generate 2D lidar bird's-eye view features from lidar point cloud data by voxelizing the point cloud and aggregating along the height dimension; S4: Obtain fusion features based on the 2D bird's-eye view features and the 2D lidar bird's-eye view features; The fused features are input into the target detection head module to obtain the target detection results and the segmentation results of the fused features.

2. The anti-interference target detection method based on multimodal fusion of autonomous driving scenarios according to claim 1 is characterized in that: In step S1, the nuscenes dataset is used as the input data source; the multi-view images in the nuscenes dataset are collected by cameras surrounding the vehicle, and the point cloud data in the nuscenes dataset is obtained by a rotating lidar.

3. The anti-interference target detection method based on multimodal fusion of autonomous driving scenarios according to claim 2 is characterized in that: The dataset also includes detailed annotation information and map data of the city.

4. The anti-interference target detection method based on multimodal fusion of autonomous driving scenarios according to claim 1 is characterized in that: Step S2 includes the following steps: S21: input the multi-view image into the view encoder module to obtain multi-view image features; S22: Use the Mask-RCNN algorithm to perform image segmentation and enhancement on multi-view image features and generate masks of occluded areas; S23: Input the multi-view image features corresponding to the occluded area mask into the temporal self-attention mechanism module based on the Transformer architecture to obtain features processed by the temporal self-attention mechanism; S24: storing the features processed by the temporal self-attention mechanism into the cache mechanism module, and obtaining the features after the target confidence level is changed by calculating the confidence level change and comparing it with the preset confidence level change threshold; S25: Convert the 2D image coordinates in the feature into 3D ego-vehicle coordinates through the view projector module; S26: Input the converted three-dimensional coordinate data into the bird's-eye view encoder module to obtain the two-dimensional view bird's-eye view features.

5. The anti-interference target detection method based on multimodal fusion of autonomous driving scenarios according to claim 4 is characterized in that: The view encoder module also includes a feature pyramid network and a feature adaptation module; The feature pyramid network fuses features of different scales through top-down and lateral connections; The feature adaptation module upsamples the multi-scale features through upsampling operations, adaptive average pooling operations and convolution layers, connects the sampled features together, and then passes them through the convolution layer to obtain image features.

6. The anti-interference target detection method based on multimodal fusion of autonomous driving scenarios according to claim 4 is characterized in that: Assume that the stored continuous multi-frame detection confidences are C1, C2, ..., C T , T represents the cache queue length, and the confidence change is calculated as follows: ΔC n =max(|C n -C n-1 |,|C n -C n+1 |) Where, ΔC n Represents the confidence change, and n represents the index of the current frame.

7. The anti-interference target detection method based on multimodal fusion of autonomous driving scenarios according to claim 1 is characterized in that: Step S3 specifically includes the following steps: S31: Obtain the lidar point cloud data features through the lidar encoder module; S32: Input the lidar point cloud data features into the bird's-eye view encoder module, and based on the voxelized features, aggregate them along the height dimension through the convolution layer to generate two-dimensional lidar bird's-eye view features.

8. The anti-interference target detection method based on multimodal fusion of autonomous driving scenarios according to claim 1 is characterized in that: Step S4 specifically includes the following steps: S41: A dynamic fusion method based on channel attention is used to fuse the 2D lidar bird's-eye view features and the 2D view bird's-eye view features to output features for detection; S42: Input the features used for detection into the target detection head module to obtain the target detection result and the segmentation result of the fused features.

9. The anti-interference target detection method based on multimodal fusion of autonomous driving scenarios according to claim 1 is characterized in that: The target detection head module includes a target detection branch and a bird's-eye view segmentation branch; the target detection branch includes a SeparateHead class module, a DCNSeparateHead class module, a CenterHead class module and a first loss calculation and target generation module; the bird's-eye view segmentation branch includes a BEVGridTransform class module, a classifier module and a second loss calculation and target generation module.

10. A detection system for use in the anti-interference target detection method based on multimodal fusion of autonomous driving scenarios according to any one of claims 1 to 9, characterized in that: Includes: A data set acquisition module is used to acquire a data set, which includes multi-view images and lidar point cloud data; The 2D bird's-eye view feature acquisition module is used to process multi-view image features using a combination of semantic segmentation and instance segmentation to obtain image features. The image features are processed by adding a temporal self-attention mechanism module, and the processed features are input into the cache mechanism module. The features are then compressed into 2D bird's-eye view features after 3D coordinate projection. The 2D LiDAR bird's-eye view feature acquisition module is used to generate 2D LiDAR bird's-eye view features by voxelizing LiDAR point cloud data and aggregating them along the height dimension. A feature fusion and detection module is used to obtain fusion features based on the two-dimensional bird's-eye view features and the two-dimensional lidar bird's-eye view features; The fused features are input into the target detection head module to obtain the target detection results and the segmentation results of the fused features.

Citation Information

Cited By

  • Three-dimensional target detection method and system based on time sequence multi-view information fusion

    CN121438255A

  • A three-dimensional target detection method and system based on time-series multi-view information fusion

    CN121438255B