Intelligent driving car target detection system and method based on multi-sensor fusion

The intelligent driving vehicle target detection system, which integrates multiple sensors, utilizes phased feature extraction and information fusion of image and point cloud data to solve the problem of insufficient detection accuracy of a single sensor, achieving more accurate target detection and improving the safety of intelligent driving.

CN120876838BActive Publication Date: 2025-11-25XI'AN PETROLEUM UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511385102.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-11-25
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing single-sensor detection technologies cannot meet the perception requirements in intelligent driving vehicles, while multi-sensor fusion methods suffer from insufficient detection accuracy or the inability to obtain reliable target detection results.

Method used

An intelligent driving vehicle target detection system based on multi-sensor fusion is adopted. Through phased feature extraction and information fusion of image data and point cloud data, combined with a target detection box fusion optimization strategy, more accurate target detection boxes are generated.

Benefits of technology

It improves the accuracy of target detection and the safety of intelligent driving, stably completes target detection, enhances the utilization of multimodal information, and avoids information loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876838B_ABST
    Figure CN120876838B_ABST
Patent Text Reader

Abstract

The application discloses a multi-sensor fusion-based intelligent driving automobile target detection system and method, and belongs to the technical field of image processing, which comprises the following steps: collecting and processing image data and point cloud data; extracting image features from the processed image data in four stages; extracting point cloud features from the processed point cloud data in four stages; fusing the image feature maps extracted in each stage with the point cloud features extracted in the corresponding stage; the point cloud feature extraction is based on the fusion of the point cloud features and the image feature maps in the previous stage; the image feature maps extracted in the four stages are spliced, and multi-scale image feature maps are obtained by image feature extraction; the multi-scale image feature maps are fused with the point cloud features extracted in the fourth stage to generate a target detection frame; and the target detection frames meeting the conditions are fused to obtain an accurate target detection frame, thereby improving the safety of intelligent driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a target detection system and method for intelligent driving vehicles based on multi-sensor fusion. Background Technology

[0002] Target detection systems for intelligent driving vehicles have a wide range of applications, using various technologies to detect, analyze, and determine targets. Three-dimensional target detection technology is one of the key modules in these systems, providing crucial spatial information to ensure safe autonomous driving. With the development of intelligent driving technology, single-sensor detection techniques are no longer sufficient for vehicle perception requirements. Therefore, multi-sensor fusion has become a research hotspot. Utilizing multi-sensor fusion to compensate for the shortcomings of single sensors can effectively improve perception accuracy. Among these methods, LiDAR and camera fusion is currently the most widely used and important approach. It uses multiple sensors to collect raw data and combines data processing and information fusion techniques to detect the appearance and distance of targets. However, in practical applications, problems still exist such as insufficient detection accuracy or even the inability to obtain reliable target detection results. Summary of the Invention

[0003] To address the problems in the existing technology, this invention provides an intelligent driving vehicle target detection system and method based on multi-sensor fusion.

[0004] On the one hand, a target detection method for intelligent driving vehicles based on multi-sensor fusion is provided, the method comprising:

[0005] S1: Image data and point cloud data acquisition and processing;

[0006] S2: Extract image features from the processed image data in four stages, extract point cloud features from the processed point cloud data in four stages, and fuse the image feature map extracted in each stage with the point cloud feature extracted in the corresponding stage to enhance the point cloud features.

[0007] Starting from the second stage, each point cloud feature extraction is performed based on the fusion of the point cloud features and image feature maps from the previous stage.

[0008] S3: The image feature maps extracted in the four stages are stitched together to obtain a new image feature map. The new image feature map is then subjected to image feature extraction again to obtain a multi-scale image feature map. The multi-scale image feature map is then fused with the point cloud features extracted in the fourth stage.

[0009] S4: Generate target detection boxes based on the fusion results in S3, and fuse the target detection boxes that meet the conditions to obtain more accurate target detection boxes.

[0010] Furthermore, in S1, the image data and processing specifically include: acquiring image data of the surrounding environment, and then using three methods—random rotation, color transformation, and random noise—to perform data enhancement;

[0011] Point cloud data acquisition and processing specifically includes: acquiring point cloud data of the surrounding environment, dividing the point cloud data into effective regions, and filtering the point cloud data of the effective regions.

[0012] Furthermore, in S2 and S3, image feature extraction includes:

[0013] Image feature extraction is performed by using the initially processed image data, the image feature map extracted in the previous stage, or the new image feature map through two-dimensional convolution. The image feature map extracted after two-dimensional convolution is used as input to generate an aggregated feature map through global average pooling.

[0014] Then, the kernel size is adaptively selected based on the mechanism of grouped convolution;

[0015] After one-dimensional convolution, the Sigmoid activation function is used to obtain the channel attention feature map. Then, the image feature map extracted after two-dimensional convolution is multiplied element-wise with the channel attention feature map to generate a weighted feature map.

[0016] Using the weighted feature map as input, the aggregated feature map is first generated by max pooling and average pooling operations, then concatenated and processed by convolution.

[0017] Next, the spatial attention feature map is obtained using the Sigmoid activation function; finally, the weighted feature map is multiplied element-wise with the spatial attention feature map to generate the final weighted feature map, which is the image feature map extracted at each stage or the multi-scale image feature map.

[0018] Furthermore, the adaptive selection of the convolution kernel size based on the grouped convolution mechanism includes:

[0019] One-dimensional convolution kernel Size and channel dimension They are set proportionally, and there is a mapping relationship between them:

[0020] (1)

[0021] In equation (1), Indicates the kernel size. This represents the channel dimension in the aggregated feature map. and It is a constant of a linear function in one variable. Indicates a linear mapping;

[0022] Channel dimension Setting it to a multiple of 2, equation (1) is expanded into a nonlinear function formula:

[0023] (2)

[0024] Solve :

[0025] (3)

[0026] in, Indicates separation The most recent odd number, express .

[0027] Furthermore, in S2 and S3, information fusion includes:

[0028] SA1: Decorates point cloud features onto the image feature map to generate an enhanced image feature map;

[0029] SA2: Combines point cloud features with enhanced image feature maps to obtain enhanced point cloud features.

[0030] Furthermore, the SA1 specifically includes:

[0031] A blank feature map with the same size as the image feature map is constructed, and the blank feature map is filled with projected point cloud features. Then, it is decorated onto the image feature map to obtain the enhanced image feature map. During the filling process, multiple point cloud features projected to the same location are averaged and pooled. : (4)

[0032] In equation (4), The feature map representing the feature filling of the point cloud. Indicates projection to The point cloud features within the location (u, v) are represented by n, which is the number of projection points.

[0033] Furthermore, the SA2 specifically includes:

[0034] The feature map filled with point cloud features and the enhanced image feature map are mapped to the same channel dimension through convolution and element-wise addition. Nonlinear compression is performed using the hyperbolic tangent function and convolution to generate a single-channel point cloud feature map.

[0035] To further obtain the importance weight of each point, the sigmoid activation function is used to normalize the weight mapping to the range [0,1], thus obtaining the weight map. ,

[0036] (5)

[0037] In equation (5), , , F represents the learnable weight matrix in the data fusion layer, σ represents the sigmoid activation function, and F... P F represents point cloud features. EI表示 Enhanced image feature maps;

[0038] The weighted graph obtained after the above processing is obtained by element-wise multiplication. The feature map is combined with the point cloud feature map, and then concatenated with the enhanced image feature map to obtain the enhanced point cloud features:

[0039] (6)

[0040] In equation (6), F fuse This indicates enhanced point cloud features. Feature map representing point cloud feature filling;

[0041] The above process enables the point cloud features and image features to be fully integrated.

[0042] Furthermore, S4 specifically includes:

[0043] S41: Generate target detection boxes based on the information fusion results of multi-scale image feature maps and point cloud features extracted in the fourth stage. Sort the target detection boxes from high to low according to their score confidence and store the sorting results in an array. Each element in the array records the score and coordinates of the corresponding target detection box.

[0044] S42: Create two empty containers, one is a clustering container (Clusters) to store candidate target detection boxes belonging to the same target, and the other is a fusion container (Fusions) to store fused boxes after weighted fusion.

[0045] S43: Select the target detection box with the highest confidence score from the sorted target detection boxes as the first cluster box and the initial fusion result, namely cluster1 and fusion1;

[0046] S44: Calculate the Euclidean distance (AED) between other target detection boxes and the fusion boxes in the current Fusions container. If the AED value of a target detection box is determined to be less than a preset threshold, add the target detection box to the corresponding clustering container and perform weighted fusion using equation (7) to update the confidence level.

[0047] (7)

[0048] In equation (7), This represents the updated confidence level of the target detection box. This represents the confidence level of candidate bounding boxes in the clustering list. This indicates the number of candidate bounding boxes in the clustering list;

[0049] Using equation (8) for weighted fusion of x-coordinates: (8)

[0050] In equation (8), This represents the x-coordinate of the updated target detection box. This represents the confidence level of candidate bounding boxes in the clustering list. Represents the x-coordinate of the candidate object detection boxes in the clustering list;

[0051] Using equation (9) for weighted fusion of y-coordinates:

[0052] (9)

[0053] In equation (9), This represents the y-coordinate of the updated target detection box. This represents the confidence level of candidate bounding boxes in the clustering list. Represents the y-coordinate of the candidate target detection boxes in the cluster list;

[0054] The above fusion process yields a new fusion box. If the calculated AED is greater than the threshold, the target detection box is considered a new target and placed into a new clustering container for further fusion operations. Finally, the fusion result is output.

[0055] On the other hand, a target detection system for intelligent driving vehicles based on multi-sensor fusion is provided to implement the aforementioned target detection method for intelligent driving vehicles based on multi-sensor fusion. The system includes:

[0056] LiDAR sensors, installed in autonomous vehicles, are used to collect point cloud data;

[0057] Camera sensors, installed in intelligent driving vehicles, are used to collect image data;

[0058] The data processing unit is used for processing point cloud data and image data;

[0059] The feature extraction module is used to extract image features from the processed image data in four stages, and also to extract point cloud features from the processed point cloud data in four stages. Starting from the second stage, each point cloud feature extraction is based on the fusion of the point cloud features and image feature maps from the previous stage.

[0060] It is also used to stitch together the image feature maps extracted step by step from the four stages to obtain a new image feature map, and then to extract image features from the new image feature map again to obtain a multi-scale image feature map;

[0061] The feature fusion module is used to fuse the image feature maps extracted at each stage with the point cloud features extracted at the corresponding stage, and also to fuse the multi-scale image feature maps with the point cloud features extracted in the fourth stage.

[0062] The target detection box module is used to generate target detection boxes based on the fusion results of multi-scale image feature maps and point cloud features extracted in the fourth stage.

[0063] The target detection box fusion module is used to fuse target detection boxes that meet the conditions to obtain more accurate target detection boxes.

[0064] The beneficial effects of the technical solution provided by this invention are as follows: First, the collected image data and point cloud data are processed. Then, image feature extraction and point cloud feature extraction are performed in four stages. The point cloud data is enhanced by fusion of image information in stages to ensure full utilization of multimodal information and prevent information loss when the point cloud data is projected onto a two-dimensional image. Furthermore, combined with the target detection box fusion optimization strategy, target detection can be completed stably and accurately, improving the safety and reliability of intelligent driving.

[0065] Secondly, to address the problem of insufficient image feature representation when relying solely on two-dimensional convolution for feature extraction, this invention proposes an image feature extraction network (HA-Net). This image feature extraction network can enhance the expressive power of image feature maps, thereby improving the accuracy of three-dimensional target detection.

[0066] In addition, during the object detection stage, object detection boxes are generated based on the fusion results. Candidate object detection boxes are filtered by aggregating Euclidean distance, and the object detection boxes that meet the conditions are further fused to make full use of the information of each candidate object detection box, so that the generated object detection boxes are closer to the real object detection boxes. Attached Figure Description

[0067] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0068] Figure 1 This is a flowchart of a target detection method for intelligent driving vehicles based on multi-sensor fusion provided by the present invention;

[0069] Figure 2 This is a flowchart of image feature extraction, point cloud feature extraction, and information fusion provided by the present invention;

[0070] Figure 3 This invention provides a flowchart for image feature extraction.

[0071] Figure 4 This invention provides an information fusion flowchart;

[0072] Figure 5 This is a flowchart of a target detection box fusion method provided by the present invention;

[0073] Figure 6 This is a diagram illustrating the weighted fusion process of candidate target detection boxes provided by the present invention;

[0074] Figure 7 This is an installation diagram of an intelligent driving vehicle target detection system based on multi-sensor fusion provided by the present invention.

[0075] Reference numerals: 1-LiDAR sensor; 2-Camera sensor; 3-User display; 4-Intelligent driving vehicle; 5-Data processing unit. Detailed Implementation

[0076] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0077] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0078] Example 1

[0079] See Figure 1 A target detection method for intelligent driving vehicles based on multi-sensor fusion includes the following steps:

[0080] Step (1): Collect raw image data of the environment surrounding the intelligent driving vehicle to provide necessary data information for subsequent data processing and fusion in the system. Specifically, this includes:

[0081] The first step is to collect a large amount of image data under different types of environments (environments in different places). For each environment, sample photos should be taken from different angles and under different lighting conditions to increase the diversity of the raw data. For example, 10,000 raw images should be taken around the predetermined road.

[0082] Step 2: Convert the collected JPG images to PNG format and save them.

[0083] Step 3: Considering the common transformations of objects in the test environment, three methods are used for data augmentation: random rotation, color transformation, and random noise. For example, one variant is generated for each sample photo and each augmentation method. Therefore, the total number of samples after data augmentation is 40,000 (10,000 PNG images, 10,000 randomly transformed images, 10,000 color-transformed images, and 10,000 random noise images). This method increases the number and diversity of samples, resulting in better training performance and improved model robustness.

[0084] Step (2): Point cloud data acquisition and processing, specifically including:

[0085] Step 1 involves collecting a large amount of point cloud data under different environments. For each environment, sample point cloud data under different environmental conditions should be collected as much as possible to improve the diversity of the raw data. For example, 10,000 frames of discontinuous raw point cloud data should be collected around a predetermined road.

[0086] Step 2 involves dividing the point cloud data into effective regions. If the point cloud density is too low to accurately represent obstacle information within 100 meters directly in front of the LiDAR, the point cloud data beyond 100 meters should be removed, and the point cloud data within 20 meters behind the LiDAR should be retained. Therefore, the range of the autonomous vehicle in the X direction is (-20, 100), and the point cloud data within 10 meters on both sides of the autonomous vehicle's Y axis should be retained, i.e., the range in the Y direction is (-10, 10), and the range in the Z direction is (-2, 5).

[0087] Step 3: Filter the point cloud data. For example, voxel filtering can be performed on the point cloud data to remove some dense and useless points from the original point cloud data, reducing the number of points without affecting the geometric features of objects in the point cloud.

[0088] Step (3): Extract image features from the processed image data in four stages, extract point cloud features from the processed point cloud data in four stages, and fuse the image feature map extracted in each stage with the point cloud features extracted in the corresponding stage to enhance the point cloud features; wherein, starting from the second stage, each point cloud feature extraction is based on the fusion of the point cloud features and image feature map of the previous stage.

[0089] The image feature maps extracted in the four stages are stitched together to obtain a new image feature map. The new image feature map is then subjected to image feature extraction again to obtain a multi-scale image feature map. The multi-scale image feature map is then fused with the point cloud features extracted in the fourth stage.

[0090] It should be noted that, see Figure 2 , Figure 2 The process includes an Image stream and a Point Cloud stream. In the Image stream, the processed image data is used as input. The first feature extraction is performed through a 2D convolution and a feature extraction network (HA-Net) to obtain image feature map F1. Then, F1 is subjected to a second feature extraction through a 2D convolution and a feature extraction network to obtain image feature map F2. Then, F2 is subjected to a third feature extraction through a 2D convolution and a feature extraction network to obtain image feature map F3. Then, F3 is subjected to a fourth feature extraction through a 2D convolution and a feature extraction network to obtain image feature map F4. Then, F1, F2, F3, and F4 are subjected to a 2D transposed convolution layer and concatenated to obtain a new image feature map. Finally, feature extraction is performed through a 2D convolution and a feature extraction network to obtain multi-scale image feature maps.

[0091] The Point Cloud flow includes four paired Set Abstractions and a Feature Propagation Layer. The Propagation layer is used for point cloud feature extraction. The outputs of the abstraction and feature propagation layers are denoted as SAi and FPi (i=1,2,3,4), respectively. First, the processed point cloud data is used as input for the first point cloud feature extraction to obtain SA1. Information is fused with F1 and SA1 to enhance the point cloud features. Based on the information fusion of F1 and SA1, the second point cloud feature extraction is performed to obtain SA2. Information is fused with F2 and SA2. Then, the third point cloud feature extraction is performed to obtain SA3. Information is fused with F3 and SA3. Then, the fourth point cloud feature extraction is performed to obtain SA4. The point cloud features extracted in the fourth stage (the output of SA4) are fused with the multi-scale image feature map through the fusion module (LI-Fusion) to obtain a compact and discriminative feature representation. Then, it is input into the object detection box fusion module (AED-WBF) to generate and fuse object detection boxes to obtain more accurate object detection boxes. Finally, it is sent to the detection head for foreground point segmentation and 3D proposal generation to obtain segmentation and detection results.

[0092] It should be noted that the point cloud feature extraction and detection at each stage uses the PointNet++ object detection method for feature extraction.

[0093] Image feature extraction at each stage is performed using a feature extraction network (HA-Net), see [link to documentation]. Figure 3 Specifically, this includes: extracting image features from the initially processed image data, the image feature map extracted in the previous stage, or the new image feature map through two-dimensional convolution; using the image feature map extracted after two-dimensional convolution as input; firstly, generating aggregated feature maps through global average pooling (GAP); and to avoid wasting computational resources due to manual adjustment, adaptively selecting the kernel size based on the mechanism of grouped convolution. Specifically, the one-dimensional convolution kernel... Size and channel dimension They are set proportionally, and there is a mapping relationship between them:

[0094] (1)

[0095] In equation (1), Indicates the kernel size. This represents the channel dimension in the aggregated feature map. and It is a constant of a linear function in one variable. Indicates a linear mapping;

[0096] Then, in order to enhance the nonlinear characteristics of the mapping, the channel dimension is increased. Setting it to a multiple of 2 fully utilizes computational resources and enhances the model's expressiveness, equation (1) is extended to a nonlinear function formula:

[0097] (2)

[0098] Solve :

[0099] (3)

[0100] in, Indicates separation The most recent odd number, express .

[0101] Next, after using one-dimensional convolution, the sigmoid activation function (σ) is used to obtain the channel attention feature map. Then, the image feature map extracted after two-dimensional convolution is multiplied element-wise with the channel attention feature map to generate a weighted feature map.

[0102] Using the weighted feature map as input, the system first generates an aggregated feature map by performing max pooling and average pooling operations. The two feature maps from max pooling and average pooling are then concatenated (element-wise addition) and processed through a series of convolutions (3×3conv, 7×7conv). Next, the sigmoid activation function is used to obtain the spatial attention feature map. Finally, the weighted feature map is multiplied element-wise with the spatial attention feature map to generate the final weighted feature map. This final weighted feature map is either the image feature map extracted at each stage or a multi-scale image feature map.

[0103] It should also be noted that the information fusion between the image feature maps extracted in each stage and the point cloud features extracted in the corresponding stage, as well as the information fusion between the multi-scale image feature maps and the point cloud features extracted in the fourth stage, are all achieved through methods such as... Figure 4 The fusion process shown specifically includes:

[0104] First, construct a feature map F corresponding to the image feature map. IConsistently sized blank feature maps, and projected (dense) point cloud features F P Filling in blank feature maps involves addressing the issue that, as the network deepens layer by layer, the resolution of the image feature maps decreases while their receptive field increases. This can lead to multiple point clouds being projected onto the same location (u, v) in the image. To avoid information redundancy and improve feature quality, average pooling is performed on the features of multiple point clouds projected onto the same location during the filling process. : (4)

[0105] In equation (4), Representing point cloud features F P The filled feature map Indicates projection to The point cloud features within the location (u, v) are represented by n, which is the number of projection points.

[0106] Then, the point cloud features F P Filled feature map Decorate to image features F I Enhanced image feature map F is formed in the middle. I '.

[0107] Then the point cloud features F P Filled feature map With enhanced image feature map F I To construct a more discriminative and compact feature representation, the two modalities are combined. First, features from both modalities are convolved and element-wise added to map to the same channel dimension for feature alignment and information complementarity. Then, non-linear compression using the hyperbolic tangent function (tanh) and convolution generates a single-channel point cloud feature map to capture key semantic responses. To further obtain the importance weights for each point, the sigmoid activation function is used to normalize the weight mapping to the [0,1] range, resulting in a weight map. ,

[0108] (5)

[0109] In equation (5), , , F represents the learnable weight matrix in the data fusion layer, σ represents the sigmoid activation function, and F... P F represents point cloud features. EI This represents the enhanced image feature map (i.e., F). I').

[0110] Then, the weighted graph processed above is joined by element-wise multiplication. Feature map filled with point cloud features Combined, and then with enhanced image features F EI By stitching together, enhanced point cloud features F are obtained. fuse :

[0111] (6)

[0112] In equation (6), F fuse This indicates enhanced point cloud features. A feature map representing the feature filling of a point cloud.

[0113] The above process enables the point cloud features and image features to be fully integrated.

[0114] Step (4): Generate target detection boxes based on the fusion results of the multi-scale image feature maps and the point cloud features extracted in the fourth stage, and fuse the target detection boxes that meet the conditions to obtain more accurate target detection boxes. See [link to relevant documentation]. Figure 5 The specific process includes:

[0115] Step 1: Sort the target detection boxes from highest to lowest score confidence (Score ranking) and store the ranking results in an array. Each element in the array records the score and coordinates of the corresponding target detection box.

[0116] Step 2: Create two empty containers. One container, called the Clusters container, will store the clustering results, and the other container, called the Fusions container, will store the fusion results. The Clusters container stores candidate bounding boxes belonging to the same target, while the Fusions container stores the weighted fused bounding boxes.

[0117] Step 3: First, perform the fusion of the first three object detection boxes. During the fusion process, select the object detection box with the highest confidence score from the sorted object detection boxes as the first cluster box and the initial fusion result, namely cluster1 and fusion1.

[0118] Step 4: Calculate the Euclidean distance (AED) between other target detection boxes and the fusion boxes in the current Fusions container. If the AED value of a target detection box is found to be less than a preset threshold, add the target detection box to the corresponding cluster container for weighted fusion (as shown in cluster1), and update the confidence using equation (7):

[0119] (7)

[0120] In equation (7), This represents the updated confidence level of the target detection box. This represents the confidence level of the target detection boxes in the target candidate cluster list. This indicates the number of target detection boxes among the target candidates in the clustering list;

[0121] Using equation (8) for weighted fusion of x-coordinates: (8)

[0122] In equation (8), This represents the x-coordinate of the updated target detection box. This represents the confidence level of the target detection boxes in the target candidate cluster list. This represents the x-coordinate of the target detection box in the cluster list;

[0123] Using equation (9) for weighted fusion of y-coordinates:

[0124] (9)

[0125] In equation (9), This represents the y-coordinate of the updated target detection box. This represents the confidence level of the target detection boxes in the target candidate cluster list. This represents the y-coordinate of the target detection box in the cluster list;

[0126] The above fusion process yields a new fusion box. If the calculated AED is greater than the threshold, the target detection box is considered a new target and placed into a new cluster container for further fusion operations (as shown in cluster2). This process is repeated for the first four target detection boxes, the first five target detection boxes, and so on, until the final fusion result is output.

[0127] See Figure 6 Box A and box B are two object detection boxes to be fused. X1 A X2A Y1 A Y2 Let A be the coordinates of the target detection box, and B be the coordinates of the target detection box. X1 B X2 B Y1 B Y2 Let A be the coordinates of the bounding box for object B. C Score B is the confidence score of the detection bounding box A. C Box B represents the confidence score of the target detection box, and box C represents the fused target detection box. X1 C X2 C Y1 C Y2 Score C represents the coordinates of the merged object detection bounding box. C This represents the confidence score of the merged target detection bounding box.

[0128] It is worth noting that in this invention, the collected image data and point cloud data are first processed, and then image feature extraction and point cloud feature extraction are performed in four stages. The point cloud data is enhanced by fusion of image information in stages to ensure full utilization of multimodal information, so that no information loss occurs when the point cloud data is projected onto the two-dimensional image. Furthermore, combined with the target detection box fusion optimization strategy, target detection can be completed stably and accurately, improving the safety and reliability of intelligent driving.

[0129] Secondly, to address the problem of insufficient image feature representation when relying solely on two-dimensional convolution for feature extraction, this invention proposes an image feature extraction network (HA-Net). This image feature extraction network can enhance the expressive power of image feature maps, thereby improving the accuracy of three-dimensional target detection.

[0130] In addition, during the object detection stage, object detection boxes are generated based on the fusion results. Candidate object detection boxes are filtered by aggregating Euclidean distance, and the object detection boxes that meet the conditions are further fused to make full use of the information of each candidate object detection box, so that the generated object detection boxes are closer to the real object detection boxes.

[0131] Example 2

[0132] See Figure 7 A target detection system for intelligent driving vehicles based on multi-sensor fusion is used to implement the target detection method for intelligent driving vehicles based on multi-sensor fusion in Embodiment 1, including a lidar sensor 1, a camera sensor 2, a user display 3, an intelligent driving vehicle 4, and a data processing unit 5.

[0133] The lidar sensor 1 is installed on the top of the intelligent driving vehicle 4. The camera sensor 2 is installed around the lidar sensor 1 with the lidar sensor 1 as the center. The user display 3 is installed in the driver's cabin of the intelligent driving vehicle 4 and is mainly used to display the working status of the entire system. The data processor 5 is installed inside the intelligent driving vehicle 4 next to the user display 3 and is mainly used for processing and fusing point cloud data and image data to improve detection accuracy.

[0134] The data processing unit 5 includes a feature extraction module, which is used to extract image features from the processed image data in four stages, and to extract point cloud features from the processed point cloud data in four stages. Starting from the second stage, each point cloud feature extraction is based on the fusion of the point cloud features and image feature maps from the previous stage. The module is also used to stitch together the image feature maps extracted in the four stages to obtain a new image feature map, and to extract image features from the new image feature map again to obtain a multi-scale image feature map.

[0135] The data processing unit 5 also has a feature extraction module, which is used to fuse the image feature map extracted at each stage with the point cloud feature extracted at the corresponding stage, and also to fuse the multi-scale image feature map with the point cloud feature extracted in the fourth stage.

[0136] The data processing unit 5 also includes a target detection box module, which generates target detection boxes based on the fusion results of multi-scale image feature maps and point cloud features extracted in the fourth stage.

[0137] The data processing unit 5 also includes a target detection box fusion module, which is used to fuse target detection boxes that meet the conditions to obtain more accurate target detection boxes.

[0138] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0139] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A target detection method for intelligent driving vehicles based on multi-sensor fusion, characterized in that, The method includes: S1: Image data and point cloud data acquisition and processing; S2: Extract image features from the processed image data in four stages, extract point cloud features from the processed point cloud data in four stages, and fuse the image feature map extracted in each stage with the point cloud feature extracted in the corresponding stage to enhance the point cloud features. Starting from the second stage, each point cloud feature extraction is performed based on the fusion of the point cloud features and image feature maps from the previous stage. S3: The image feature maps extracted in the four stages are stitched together to obtain a new image feature map. The new image feature map is then subjected to image feature extraction again to obtain a multi-scale image feature map. The multi-scale image feature map is then fused with the point cloud features extracted in the fourth stage. S4: Generate target detection boxes based on the fusion results in S3, and fuse the target detection boxes that meet the conditions to obtain more accurate target detection boxes, specifically including: S41: Generate target detection boxes based on the information fusion results of multi-scale image feature maps and point cloud features extracted in the fourth stage. Sort the target detection boxes from high to low according to their score confidence and store the sorting results in an array. Each element in the array records the score and coordinates of the corresponding target detection box. S42: Create two empty containers, one is a clustering container (Clusters) to store candidate target detection boxes belonging to the same target, and the other is a fusion container (Fusions) to store fused boxes after weighted fusion. S43: Select the target detection box with the highest confidence score from the sorted target detection boxes as the first cluster box and the initial fusion result, namely cluster1 and fusion1; S44: Calculate the Euclidean distance (AED) between other target detection boxes and the fusion boxes in the current Fusions container. If the AED value of a target detection box is determined to be less than a preset threshold, add the target detection box to the corresponding clustering container for weighted fusion and update the confidence using equation (7): (7) In equation (7), This represents the updated confidence level of the target detection box. This represents the confidence level of candidate bounding boxes in the clustering list. This indicates the number of candidate bounding boxes in the clustering list; Using equation (8) for weighted fusion of x-coordinates: (8) In equation (8), This represents the x-coordinate of the updated target detection box. This represents the confidence level of candidate bounding boxes in the clustering list. Represents the x-coordinate of the candidate object detection boxes in the clustering list; Using equation (9) for weighted fusion of y-coordinates: (9) In equation (9), This represents the y-coordinate of the updated target detection box. This represents the confidence level of candidate bounding boxes in the clustering list. This represents the y-coordinate of the candidate target detection boxes in the clustering list; The above fusion process yields a new fusion box. If the calculated AED is greater than the threshold, the target detection box is considered a new target and placed into a new clustering container for further fusion operations. Finally, the fusion result is output.

2. The intelligent driving vehicle target detection method based on multi-sensor fusion according to claim 1, characterized in that, In S1, the image data and processing specifically include: acquiring image data of the surrounding environment, and then using three methods—random rotation, color transformation, and random noise—to enhance the data; Point cloud data acquisition and processing specifically includes: acquiring point cloud data of the surrounding environment, dividing the point cloud data into effective regions, and filtering the point cloud data of the effective regions.

3. The intelligent driving vehicle target detection method based on multi-sensor fusion according to claim 2, characterized in that, In S2 and S3, image feature extraction includes: Image feature extraction is performed by using the initially processed image data, the image feature map extracted in the previous stage, or the new image feature map through two-dimensional convolution. The image feature map extracted after two-dimensional convolution is used as input to generate an aggregated feature map through global average pooling. Then, the kernel size is adaptively selected based on the mechanism of grouped convolution; After one-dimensional convolution, the Sigmoid activation function is used to obtain the channel attention feature map. Then, the image feature map extracted after two-dimensional convolution is multiplied element-wise with the channel attention feature map to generate a weighted feature map. Using the weighted feature map as input, the aggregated feature map is first generated by max pooling and average pooling operations, then concatenated and processed by convolution. Next, the spatial attention feature map is obtained using the Sigmoid activation function; finally, the weighted feature map is multiplied element-wise with the spatial attention feature map to generate the final weighted feature map, which is the image feature map extracted at each stage or the multi-scale image feature map.

4. The intelligent driving vehicle target detection method based on multi-sensor fusion according to claim 3, characterized in that, The adaptive selection of kernel size based on the mechanism of grouped convolution includes: One-dimensional convolution kernel Size and channel dimension They are set proportionally, and there is a mapping relationship between them: (1) In equation (1), Indicates the kernel size. This represents the channel dimension in the aggregated feature map. and It is a constant of a linear function in one variable. Indicates a linear mapping; Channel dimension Setting it to a multiple of 2, equation (1) is expanded into a nonlinear function formula: (2) Solve : (3) in, Indicates separation The most recent odd number, express .

5. The intelligent driving vehicle target detection method based on multi-sensor fusion according to claim 4, characterized in that, In S2 and S3, information fusion includes: SA1: Decorates point cloud features onto the image feature map to generate an enhanced image feature map; SA2: Combines point cloud features with enhanced image feature maps to obtain enhanced point cloud features.

6. The intelligent driving vehicle target detection method based on multi-sensor fusion according to claim 5, characterized in that, SA1 specifically includes: A blank feature map with the same size as the image feature map is constructed, and the blank feature map is filled with projected point cloud features. Then, it is decorated onto the image feature map to obtain the enhanced image feature map. During the filling process, multiple point cloud features projected to the same location are averaged and pooled. : (4) In equation (4), The feature map representing the feature filling of the point cloud. Indicates projection to The point cloud features within the location (u, v) are represented by n, which is the number of projection points.

7. The intelligent driving vehicle target detection method based on multi-sensor fusion according to claim 6, characterized in that, The SA2 specifically includes: The feature map filled with point cloud features and the enhanced image feature map are mapped to the same channel dimension through convolution and element-wise addition. Nonlinear compression is performed using the hyperbolic tangent function and convolution to generate a single-channel point cloud feature map. To further obtain the importance weight of each point, the sigmoid activation function is used to normalize the weight mapping to the range [0,1], thus obtaining the weight map. , (5) In equation (5), , , F represents the learnable weight matrix in the data fusion layer, σ represents the sigmoid activation function, and F... P F represents point cloud features. EI Represents the enhanced image feature map; The weighted graph obtained after the above processing is obtained by element-wise multiplication. The feature map is combined with the point cloud feature map, and then concatenated with the enhanced image feature map to obtain the enhanced point cloud features: (6) In equation (6), F fuse This indicates enhanced point cloud features. Feature map representing point cloud feature filling; The above process enables the point cloud features and image features to be fully integrated.

8. A target detection system for intelligent driving vehicles based on multi-sensor fusion, characterized in that, The system is used to implement the multi-sensor fusion-based intelligent driving vehicle target detection method according to any one of claims 1-7, the system comprising: LiDAR sensors, installed in autonomous vehicles, are used to collect point cloud data; Camera sensors, installed in intelligent driving vehicles, are used to collect image data; The data processing unit is used for processing point cloud data and image data; The feature extraction module is used to extract image features from the processed image data in four stages, and also to extract point cloud features from the processed point cloud data in four stages. Starting from the second stage, each point cloud feature extraction is based on the fusion of the point cloud features and image feature maps from the previous stage. It is also used to stitch together the image feature maps extracted step by step from the four stages to obtain a new image feature map, and then to extract image features from the new image feature map again to obtain a multi-scale image feature map; The feature fusion module is used to fuse the image feature maps extracted at each stage with the point cloud features extracted at the corresponding stage, and also to fuse the multi-scale image feature maps with the point cloud features extracted in the fourth stage. The target detection box module is used to generate target detection boxes based on the fusion results of multi-scale image feature maps and point cloud features extracted in the fourth stage. The target detection box fusion module is used to fuse target detection boxes that meet the conditions to obtain more accurate target detection boxes.

Citation Information

Patent Citations

  • Three-dimensional target detection method based on multi-modal fusion and deformable attention

    CN117975436A

  • Three-dimensional object detection method based on multi-modal fusion and deep attention mechanism

    WO2024217115A1