Roadside Object Detection Method, System and Device for Multimodal Feature Alignment and Fusion

Through the multimodal feature alignment and fusion method, the problem of space-time and space-time aberration in roadside object detection is solved, efficient feature fusion and accurate three-dimensional object detection are achieved, and detection accuracy and speed are improved.

CN116844129BActive Publication Date: 2025-07-04BEIJING UNIV OF CHEM TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310909012.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-24
Publication Date
2025-07-04
Estimated Expiration
2043-07-24

AI Technical Summary

Technical Problem

In the prior art, there is an asynchronous deviation in time and space of multimodal data, which hinders the fusion of multimodal features. Especially in roadside target detection, the alignment difficulty of camera and lidar data is greatly increased, resulting in a decrease in detection accuracy.

Method used

The roadside object detection method is adopted for multimodal feature alignment and fusion, and by obtaining multi-scale feature maps of image and point cloud data, using attention weight matrix and deconvolution processing, the effective fusion of image and point cloud features is achieved, and a three-dimensional bounding box and classification score are generated.

Benefits of technology

It effectively compensates for the spatial and temporal async problem of data of different modalities, improves detection accuracy and processing speed, fully retains image semantic information, reduces false detection targets, and improves detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116844129B_ABST
    Figure CN116844129B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of computer vision, and specifically relates to a roadside target detection method, system and device for multi-modal feature alignment and fusion, aiming to solve the problem that different modal data have spatio-temporal asynchronization deviations, which hinder multi-modal feature fusion. The method of the present invention includes: obtaining a multi-scale feature map group of an image to be processed by using a convolutional neural network; obtaining a set of point features after feature extraction by using the PointNet++ method; searching for adjacent regions and fusing the multi-scale feature map group with point sets of different sizes; introducing global information by adding a channel-level information interaction layer in the feature fusion module; generating three-dimensional bounding boxes and classification scores through a detection head to perform target detection on the point cloud to be processed. The method of the present invention adopts a search alignment method, which can compensate for the adverse effects of spatio-temporal asynchronization problems existing in different modal data on multi-modal feature fusion and accurately obtain detection results.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Object detection is a core component of autonomous driving and intelligent transportation. Currently, in 3D object detection tasks in road traffic scenarios, data such as images and point clouds are usually used as inputs to predict the geometric and semantic information of key elements on the road. Image-based detection has rich image texture information but lacks the dimension of spatial information and cannot accurately restore the position of spatial information. LiDAR has the advantages of long detection distance, being unaffected by light, and being able to accurately obtain target distance information, etc., and can make up for the shortcomings of camera images. LiDAR-based detection provides rich 3D structure information, but there is a problem of sparse point clouds. Therefore, autonomous driving object detection mainly uses multi-sensor fusion, especially the fusion of LiDAR and cameras.

[0003] Features of different modalities have different representation methods, and how to align multiple modality features for fusion is the key. No matter which method, an important prerequisite for multi-modal data fusion is to calibrate the data of different sensors into the same coordinate system. Specifically, the corresponding relationship between the data of different sensors should be accurate. For traffic datasets in real-world scenarios, due to errors in the dataset calibration and synchronization process, it is a common problem that camera and LiDAR data are out of sync in time and space. At the same time, we also need to consider that in feature-level fusion methods, new errors will occur in the corresponding relationship determined by calibration parameters after the raw data of different modalities are feature-extracted. Since these features are often enhanced and aggregated, a key challenge in fusion is how to effectively align the features transformed from the two modalities. When there are deviations in obtaining a rough corresponding relationship through the calibration parameters of the raw data, the difficulty of combining the advantages of the two modalities is greatly increased.

[0004] In recent years, multi-modal 3D object detection has been a research focus. Currently, popular multi-modal 3D object detection methods can be divided into data-level, feature-level, and decision-level fusion according to the fusion timing. Among them, data-level fusion mainly fuses raw or preprocessed sensor data, makes full use of the original information of the data, and has relatively low computational requirements, but is not flexible enough. Decision-level fusion combines the decision outputs of different data modality network structures, has high flexibility and modularity, but has a high computational cost and will lose a lot of intermediate features. Feature-level fusion fuses features in the intermediate layer, enabling the network to learn different feature representations, and the difficulty lies in the selection of the fusion timing.

[0005] An important prerequisite for multimodal fusion is to align the data or features of different sensors. Most of the existing methods require a large amount of computing power to support the global interaction of multimodal information, and extracting point cloud features based on methods such as voxels will lose some information. On the other hand, these methods are studied on public datasets, and it is difficult to define the effect of feature alignment and its impact on detection performance. Our experimental verification of the roadside dataset with spatio-temporal asynchronization problems can supplement the research in this field.

[0006] Based on this, the present invention provides a roadside object detection method, system and device for multimodal feature alignment and fusion. Summary of the Invention

[0007] In order to solve the above problems in the prior art, that is, the research datasets in the prior art have spatio-temporal asynchronization biases, which hinder multimodal feature fusion, the present invention provides a roadside object detection method, system and device for multimodal feature alignment and fusion.

[0008] On one hand of the present invention, a roadside object detection method for multimodal feature alignment and fusion is proposed, and the method includes:

[0009] Step S10, obtaining an image for three-dimensional object detection to be processed and its corresponding point cloud data; extracting multi-scale feature maps of the input image; extracting a plurality of point feature sets of the point cloud data;

[0010] Step S20, respectively mapping the point features in the plurality of point feature sets onto the multi-scale feature maps to obtain corresponding coordinates; obtaining regional values according to the coordinates, and obtaining image features within a plurality of regions after adding the regional values on the multi-scale feature maps as the first fusion image features;

[0011] Step S30, respectively fusing the image features of the multi-scale feature maps and each first fusion image feature with a pre-constructed attention weight matrix, and splicing the image features of the fused multi-scale feature maps with the fused first fusion image features respectively to obtain second fusion image features, and splicing the second fusion image features with the corresponding point feature sets to obtain enhanced point cloud features;

[0012] Step S40, performing deconvolution processing on the image features of each multi-scale feature map, and splicing the image features after deconvolution processing to obtain spliced image features, fusing the spliced images with the enhanced point cloud features to obtain multimodal feature fusion point clouds; generating three-dimensional bounding boxes and classification scores through a detection head for the multimodal feature fusion point clouds, and outputting them as three-dimensional object detection results; the detection head is constructed based on convolutional layers.

[0013] In some preferred embodiments, the method for obtaining the first fused image feature is as follows:

[0014] Step S21: Obtain the point cloud feature F in the point feature set p For the coordinate point P on the lidar coordinate, map the P on the multi-scale feature as the key point coordinate P c , and use the P c as a pointer to search on the multi-scale feature map to obtain the image feature F of the corresponding multi-scale feature map i ;

[0015] Step S22: Add the regional value P c around the P offset , and use the sum of the regional value P offset and the P c as a pointer to search to obtain the first fused image feature F′ i :

[0016] P offset = R n × sigmoid(W1F p );

[0017] P′ c = P c + P offset ;

[0018] wherein, the R n is the neighborhood range parameter, the W1 is the learnable weight matrix, and P′ c represents the sum of the regional value P offset and P c .

[0019] In some preferred embodiments, the method for obtaining the enhanced point cloud feature is as follows:

[0020] Step S31: Weight the point cloud feature F p in the point feature set, the image feature F i and the F′ i through multiple learnable weight matrices to obtain the attention weight matrices W attention , W′ attention :

[0021] W attention = sigmoid(W4tanh(W2F p + W3F i ));

[0022] W′ attention = sigmoid(W4tanh(W2F p+W3F′ i ));

[0023] wherein, W2, W3, and W4 are learnable weight matrices;

[0024] Step S32, according to the W attention , the W′ attention , the F i , the F′ i obtain the enhanced point cloud feature F′ p :

[0025] F′ p = C(C(W attention F i ∪ W′ attention F′ i ), F p );

[0026] wherein, C represents the feature concatenation operation, and ∪ represents the union operation.

[0027] In some preferred embodiments, the method for obtaining the concatenated image feature F w is as follows:

[0028] F w = C(F1 + ∑Deconv(F n ));

[0029] wherein, F1 is the image feature of the largest scale, F n is the image feature other than F1, and Deconv represents the deconvolution operation.

[0030] In some preferred embodiments, for the model corresponding to the multi-modal feature alignment and fusion roadside object detection method, the loss function during training includes classification loss, regression loss, and forced consistency loss;

[0031] wherein, the classification loss uses the Focal loss function; the regression parameters are optimized by the Smooth L1 loss function, and the regression loss is increased for the three parameters in the x-axis, y-axis, and z-axis directions.

[0032] In some preferred embodiments, the method for extracting multiple point feature sets of the point cloud data includes the PointNet++ method.

[0033] On the other hand, the present invention proposes a multi-modal feature alignment and fusion roadside object detection system, based on a multi-modal feature alignment and fusion roadside object detection method, the system includes:

[0034] An extraction module configured to obtain an image to be subjected to 3D object detection and its corresponding point cloud data; extract multi-scale feature maps of the input image; extract multiple point feature sets of the point cloud data;

[0035] A first fusion module configured to map point features in multiple point feature sets onto the multi-scale feature maps respectively to obtain corresponding coordinates; obtain region values according to the coordinates, and obtain image features within multiple regions after adding the region values on the multi-scale feature maps as first fusion image features;

[0036] A second fusion module configured to fuse the image features of the multi-scale feature maps and each first fusion image feature with a pre-constructed attention weight matrix respectively, and splice the image features of the fused multi-scale feature maps and the fused first fusion image features respectively to obtain second fusion image features, and splice the second fusion image features with the corresponding point feature sets to obtain enhanced point cloud features;

[0037] A result output module configured to perform deconvolution processing on the image features of each multi-scale feature map, and splice the image features after deconvolution processing to obtain spliced image features, fuse the spliced image with the enhanced point cloud features to obtain multi-modal feature fusion point clouds; generate 3D bounding boxes and classification scores through a detection head for the multi-modal feature fusion point clouds and output them as 3D object detection results; the detection head is constructed based on convolutional layers.

[0038] In a third aspect of the present invention, a storage device is proposed, in which multiple programs are stored, and the programs are adapted to be loaded and executed by a processor to implement a roadside object detection method for multi-modal feature alignment and fusion.

[0039] In a fourth aspect of the present invention, a processing device is proposed, including a processor and a storage device; the processor is adapted to execute each program; the storage device is adapted to store multiple programs; the programs are adapted to be loaded and executed by the processor to implement a roadside object detection method for multi-modal feature alignment and fusion.

[0040] Advantages of the present invention:

[0041] (1) The method of the present invention adopts a search alignment method, which can compensate for the deviation brought by the spatio-temporal asynchrony problem of different modal data to multi-modal feature fusion, has a high processing speed, and accurately obtains detection results.

[0042] (2) The method of the present invention obtains feature representations of different scales for the input image through a convolutional neural network, searches for key positions in the image target area, and gives higher weights to interact with point cloud feature information, effectively enhancing the point cloud features.

[0043] (3) When the method of the present invention performs feature fusion, it introduces global feature information to update the weights, which concentrates the underlying computing resources on the more important parts of the image and more fully preserves the effectiveness of the image semantic information. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Other features, objects, and advantages of the present application will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:

[0045] Figure 1 is a schematic flowchart of the roadside object detection method for multi-modal feature alignment and fusion of the present invention;

[0046] Figure 2 is an example diagram of the image feature extraction module of an embodiment of the roadside object detection method for multi-modal feature alignment and fusion of the present invention;

[0047] Figure 3 is an example diagram of the multi-modal feature search and alignment module of an embodiment of the roadside object detection method for multi-modal feature alignment and fusion of the present invention;

[0048] Figure 4 is an example diagram of the full-channel fusion of the multi-modal feature fusion module of an embodiment of the roadside object detection method for multi-modal feature alignment and fusion of the present invention;

[0049] Figure 5 is an example diagram of the overall model structure of an embodiment of the roadside object detection method for multi-modal feature alignment and fusion of the present invention;

[0050] Figure 6 is an example diagram of the object detection result of an embodiment of the roadside object detection method for multi-modal feature alignment and fusion of the present invention;

[0051] Figure 7 is a schematic diagram of the structure of the computer system of the server for implementing the method, system, and device embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The present application will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention and are not intended to limit the invention. Additionally, it should be noted that only the parts related to the relevant invention are shown in the drawings for the convenience of description.

[0053] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.

[0054] As Figures 1-6 shown, refer toFigure 1 , in the first embodiment of the present invention, a roadside object detection method for multi-modal feature alignment and fusion is provided. The method includes:

[0055] Step S10: Obtain the three-dimensional object detection image to be processed and its corresponding point cloud data; extract the multi-scale feature maps of the input image; extract multiple point feature sets of the point cloud data;

[0056] Step S20: Map the point features in multiple point feature sets onto the multi-scale feature maps respectively to obtain corresponding coordinates; obtain region values according to the coordinates, and obtain the image features within multiple regions after adding the region values on the multi-scale feature maps as the first fusion image features;

[0057] Step S30: Fuse the image features of the multi-scale feature maps and each first fusion image feature with a pre-constructed attention weight matrix respectively, and splice the image features of the fused multi-scale feature maps with the fused first fusion image features respectively to obtain the second fusion image features, and splice the second fusion image features with the corresponding point feature sets to obtain enhanced point cloud features;

[0058] Step S40: Perform deconvolution processing on the image features of each multi-scale feature map, and splice the image features after deconvolution processing to obtain spliced image features. Fuse the spliced image with the enhanced point cloud features to obtain multi-modal feature fusion point clouds; generate three-dimensional bounding boxes and classification scores through a detection head for the multi-modal feature fusion point clouds, and output them as three-dimensional object detection results; the detection head is constructed based on convolutional layers.

[0059] To more clearly illustrate a roadside object detection method for multi-modal feature alignment and fusion of the present invention, the following combines Figure 1 to elaborate on each step in the embodiments of the present invention in detail.

[0060] A roadside object detection method for multi-modal feature alignment and fusion in the first embodiment of the present invention includes Step S10 - Step S40, and each step is described in detail as follows:

[0061] Step S10: Obtain the three-dimensional object detection image to be processed and its corresponding point cloud data; extract the multi-scale feature maps of the input image; extract multiple point feature sets of the point cloud data;

[0062] Among them, the method for extracting multiple point feature sets of the point cloud data includes the PointNet++ method.

[0063] In a preferred embodiment of the present invention, image feature extraction can be constructed by selecting convolutional layers, normalization layers, and activation layers in a convolutional neural network. The reduction multiples of the resolutions of different image feature maps relative to the input image are 2, 4, 8, and 16 respectively, and the number of channels are 64, 128, 256, and 512 respectively. As Figure 2 shown, the image features of the point cloud data are obtained by deconvolving different scale feature maps of the reduction multiples to splice the original size feature maps. In a preferred embodiment of the present invention, four layers can be selected to encode the features of the point cloud data and four layers to decode. During the encoding process, the point features are further enriched by combining the image features. The image feature extraction network uses basic convolutional layers, corresponding to the point cloud feature extraction network.

[0064] Step S20: Map the point features in multiple point feature sets onto the multi-scale feature maps respectively to obtain corresponding coordinates; obtain region values according to the coordinates, and obtain the image features within multiple regions after adding the region values on the multi-scale feature maps as the first fused image features.

[0065] Common publicly available datasets provide calibration parameters with high accuracy, and the time synchronization error of the sampled data of the camera and the lidar is small. According to the position transformation matrix between the lidar and the camera and the camera internal parameters, the corresponding image features of the point cloud can be found, and the fused features are more representative of the two types of features. However, for the roadside dataset in the real scene, due to errors in the dataset calibration and synchronization process, it is a common problem that the camera and lidar data are out of sync in time and space. At the same time, in the feature-level fusion method, new errors will occur in the corresponding relationship determined by the calibration parameters after the original data of different modalities are extracted. Since these features are often enhanced and aggregated, a key challenge in fusion is how to effectively align the features converted from the two modalities. When there are biases in obtaining the rough corresponding relationship through the calibration parameters of the original data, the difficulty of combining the advantages of the two modalities is greatly increased.

[0066] The present invention performs multi-scale search alignment on multi-modal features. First, find the corresponding relationship between the point cloud and the image according to the calibration parameters. Determine the target area based on the projection position, and through multiple iterations, find the optimal local image features that match the point cloud in the target area. Such a design compensates for the deviation caused by the spatio-temporal out-of-sync problem of the dataset on the one hand, and improves the alignment problem after the features are continuously enhanced and aggregated on the other hand. As Figure 3 shown, different search ranges can be adopted in this embodiment to implement the above operations, for example:

[0067] Increase the search for adjacent regions of image features to find the optimal local image features that match the point cloud data to compensate for the deviation between the point cloud data and the multi-scale image feature mapping. First, calculate the key point coordinates of the point feature set mapped on the image through the calibration matrix, increase the region value based on the two-dimensional pixel coordinates, realize the interaction between the point feature and the key image feature domain feature, and achieve the goal of searching for image features within a limited range. The specific steps are as follows:

[0068] Step S21: Obtain the point cloud feature F in the point feature set p At the coordinate point P on the lidar coordinate, map the P on the multi-scale feature as the key point coordinate P c , and use the P c as a pointer to search on the multi-scale feature map to obtain the image feature F of the corresponding multi-scale feature map i ;

[0069] Step S22: Increase the region value P c around the P offset , and use the sum of the region value P offset and the P c as a pointer to search to obtain the first fused image feature F' i :

[0070] P offset =R n ×sigmoid(W1F p );

[0071] P' c =P c +P offset ;

[0072] Among them, the R n is the neighborhood range parameter, the W1 is the learnable weight matrix, and P' c represents the sum of the region value P offset and P c .

[0073] Step S30: Fuse the image feature of the multi-scale feature map and each first fused image feature with the pre-constructed attention weight matrix respectively, splice the fused image feature of the multi-scale feature map with the fused first fused image features respectively to obtain the second fused image feature, and splice the second fused image feature with the corresponding point feature set to obtain the enhanced point cloud feature;

[0074] Multi-modal feature alignment helps to integrate the advantages of different modal features. To integrate multi-modal features, effective and discriminative features need to be learned. In the point cloud, both pedestrians and railings are composed of vertically distributed sparse point clouds, and more information is required to handle these similar situations. To enhance the learning of fused features, the method of the present invention introduces global information, such as Figure 4 As shown, it is an example diagram of full-channel fusion of a multi-modal feature fusion module according to an embodiment of the present invention, providing semantic information of different dimensions of image features for each point.

[0075] Step S31, perform weighted processing on the point cloud feature F p in the set of point features, the image feature F i and the F′ i through multiple learnable weight matrices to obtain attention weight matrices W attention and W′ attention :

[0076] W attention = sigmoid(W4tanh(W2F p + W3F i ));

[0077] W′ attention = sigmoid(W4tanh(W2F p + W3F i ));

[0078] wherein, W2, W3, and W4 are learnable weight matrices;

[0079] Step S32, obtain the enhanced point cloud feature F′ attention according to the W attention , the W′ i , the F i , and the F′ p ;

[0080] F′ p = C(C(W attention F i ∪ W′ attention F′ i ), F p );

[0081] wherein, C represents the feature concatenation operation, and ∪ represents the union operation.

[0082] The method adopted by the present invention more fully combines the effectiveness of image semantic information, and in the detection of small targets that are difficult to identify, it reduces misdetected targets, thereby improving the detection accuracy.

[0083] Step S40: Perform deconvolution processing on the image features of each multi-scale feature map, splice the image features after deconvolution processing to obtain spliced image features, fuse the spliced image with the enhanced point cloud features to obtain multi-modal feature fused point clouds; generate 3D bounding boxes and classification scores through a detection head for the multi-modal feature fused point clouds and output them as 3D object detection results; the detection head is constructed based on convolutional layers.

[0084] Among them, the spliced image feature F w , and its acquisition method is:

[0085] F w = C(F1 + ∑Deconv(F n ));

[0086] Among them, F1 is the image feature of the largest scale, F n is the image feature other than F1, and Deconv represents the deconvolution operation.

[0087] For the model corresponding to the roadside object detection method of multi-modal feature alignment and fusion, its loss function during the training process includes classification loss, regression loss, and forced consistency loss;

[0088] Among them, the classification loss uses the Focal loss function; optimize the regression parameters through the Smooth L1 loss function and add regression losses to the three parameters in the x-axis, y-axis, and z-axis directions.

[0089] As Figure 5 shown, the 3D detection framework mainly consists of three parts: a feature extraction network, feature alignment and fusion, and a refinement network. Finally, the features generate 3D bounding boxes and classification scores through a detection head.

[0090] As Figure 6 shown, it is an example diagram of the object detection result of an embodiment of the present invention. The first row in the figure is the input image, and the second and third rows in the figure are the object detection result diagrams of the lidar point cloud data. Among them, the second row is the detection result of the baseline model, and the 3D ground truth box, 3D detection box, and square border are marked in the figure. There are missed detection and misdetection situations in the square border. In the case of missed detection, only the 3D ground truth box exists for the object in the square border; in the case of misdetection, only the 3D detection box exists for the object in the square border. In the detection result diagram of the present invention in the third row, the missed detection object is detected with a 3D bounding box, and the 3D detection box of the misdetected object is reduced, improving the detection performance.

[0091] Although the steps are described in the above order in the above embodiments, those skilled in the art can understand that, in order to achieve the effects of this embodiment, different steps do not have to be executed in such an order, and they can be executed simultaneously (in parallel) or in reverse order, and these simple changes are all within the protection scope of the present invention.

[0092] A roadside target detection system for multimodal feature alignment and fusion according to the second embodiment of the present invention is based on a roadside target detection method for multimodal feature alignment and fusion. The system includes:

[0093] An extraction module configured to obtain an image to be subjected to three-dimensional target detection and its corresponding point cloud data; extract a multi-scale feature map of the input image; extract a plurality of point feature sets of the point cloud data;

[0094] A first fusion module configured to map the point features in each of the plurality of point feature sets onto the multi-scale feature map to obtain corresponding coordinates; obtain a region value according to the coordinates, and obtain image features within a plurality of regions after adding the region value on the multi-scale feature map as first fusion image features;

[0095] A second fusion module configured to fuse the image features of the multi-scale feature map and each first fusion image feature with a pre-constructed attention weight matrix respectively, and splice the image features of the fused multi-scale feature map and the fused first fusion image features respectively to obtain second fusion image features, and splice the second fusion image features with the corresponding point feature sets to obtain enhanced point cloud features;

[0096] A result output module configured to perform deconvolution processing on the image features of each multi-scale feature map, splice the deconvolved image features to obtain spliced image features, fuse the spliced image with the enhanced point cloud features to obtain multimodal feature fusion point clouds; generate three-dimensional bounding boxes and classification scores through a detection head for the multimodal feature fusion point clouds and output them as three-dimensional target detection results; the detection head is constructed based on convolutional layers.

[0097] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process and related descriptions of the above-described system can refer to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0098] It should be noted that the roadside target detection system for multimodal feature alignment and fusion provided in the above embodiments is only illustrated by dividing the above functional modules. In practical applications, the above functions can be assigned to different functional modules according to needs, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be combined into one module, or further split into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing each module or step, and are not regarded as an improper limitation of the present invention.

[0099] As Figure 7 shown, a storage device according to a third embodiment of the present invention stores multiple programs, and the programs are suitable for being processed, loaded, and executed by a processor to implement a roadside target detection method for multimodal feature alignment and fusion.

[0100] As Figure 7 shown, a processing device according to a fourth embodiment of the present invention includes a processor and a storage device; the processor is suitable for executing each program; the storage device is suitable for storing multiple programs; the programs are suitable for being loaded and executed by the processor to implement a roadside target detection method for multimodal feature alignment and fusion.

[0101] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes and related descriptions of the above-described storage device and processing device can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0102] Those skilled in the art should be able to realize that the modules and method steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. The programs corresponding to the software modules and method steps can be placed in random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium well-known in the technical field. To clearly illustrate the interchangeability of electronic hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0103] Next, refer to Figure 7 , which shows a schematic structural diagram of a computer system of a server for implementing the method, system, and device embodiments of the present application.Figure 7 The server shown is only an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.

[0104] As Figure 7 shown, the computer system includes a central processing unit (CPU, Central Processing Unit) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM, Read Only Memory) 702 or the program loaded from the storage section 708 into the random access memory (RAM, Random Access Memory) 703. In the RAM 703, various programs and data required for system operation are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The input / output (I / O, Input / Output) interface 705 is also connected to the bus 704.

[0105] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT, Cathode Ray Tube), a liquid crystal display (LCD, Liquid Crystal Display), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as required. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as required so that a computer program read from it can be installed into the storage section 708 as required.

[0106] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, the computer program including program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above-described functions defined in the methods of the present application are performed. It should be noted that the computer-readable medium described above in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0107] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0108] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0109] The terms "first", "second", etc. are used to distinguish similar objects, rather than to describe or represent a specific order or sequence.

[0110] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or device / apparatus that comprises a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to those processes, methods, articles, or device / apparatus.

[0111] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. A roadside object detection method for multi-modal feature alignment and fusion, characterized in that, The method includes: Step S10: Obtain the image to be subjected to 3D object detection and its corresponding point cloud data; extract the multi-scale feature maps of the input image; extract multiple point feature sets of the point cloud data; Step S20: Map the point features in each of the multiple point feature sets onto the multi-scale feature maps respectively to obtain corresponding coordinates; obtain regional values according to the coordinates, and obtain the image features within multiple regions with the regional values added on the multi-scale feature maps as the first fused image features; Step S30: Fuse the image features of the multi-scale feature maps and each of the first fused image features with a pre-constructed attention weight matrix respectively, and splice the image features of the fused multi-scale feature maps with the fused first fused image features respectively to obtain second fused image features, and splice the second fused image features with the corresponding point feature sets to obtain enhanced point cloud features; The method for obtaining the enhanced point cloud features is: Step S31, perform weighted processing on the point cloud feature F p , image feature F i and the first fused image feature F' i through multiple learnable weight matrices to obtain attention weight matrices W attention , W' attention : W attention = sigmoid(W4tanh(W2F p +W3F i )); W′ attention = sigmoid(W4tanh(W2F p + W3F′ i )); wherein, W2, W3, and W4 are learnable weight matrices; Step S32, according to the said W attention , the said W' attention , the said F i , the said F' i obtain the enhanced point cloud feature F' p ; F′ p = C(C(W attention F i ∪ W′ attention F′ i ), F p ); wherein, C represents the feature splicing operation, and ∪ represents the union operation; Step S40: Perform deconvolution processing on the image features of each multi-scale feature map, and splice the image features after the deconvolution processing to obtain spliced image features, fuse the spliced image with the enhanced point cloud features to obtain multi-modal feature fused point clouds; generate 3D bounding boxes and classification scores through a detection head for the multi-modal feature fused point clouds, and output them as 3D object detection results; the detection head is constructed based on convolutional layers.

2. The roadside target detection method for multi-modal feature alignment and fusion according to claim 1, characterized in that The method for obtaining the first fused image features is: Step S21: Obtain the point cloud feature F in the point feature set p At the coordinate point P in the lidar coordinates, map the P on the multi-scale feature as the key point coordinate P c , using the P c as a pointer to search on the multi-scale feature map, and obtain the image feature F of the corresponding multi-scale feature map i ; Step S22, add the regional value P c around the said P offset , and use the sum of the said regional value P offset and the said P c as a pointer to conduct a search to obtain the first fused image feature F′ i : P offset = R n × sigmoid(W1F p ) P′ c = P c + P offset ; Among them, the R n is the neighborhood range parameter, the W1 is the learnable weight matrix, and P' c represents the regional value P offset and P c sum.

3. A roadside target detection method for multimodal feature alignment and fusion according to claim 1, characterized in that The spliced image feature F w , and the acquisition method thereof is as follows: F w = C(F1 + ∑Deconv(F n )); Among them, F1 is the image feature of the largest scale, and F n is the image feature other than F1, and Deconv represents the deconvolution operation.

4. The roadside target detection method for multimodal feature alignment and fusion according to claim 1, characterized in that, The loss function of the model corresponding to the roadside object detection method for multi-modal feature alignment and fusion includes classification loss, regression loss, and forced consistency loss during the training process; wherein, the classification loss adopts the Focal loss function; the regression parameters are optimized by the Smooth L1 loss function, and regression losses are added to the three parameters in the x-axis, y-axis, and z-axis directions.

5. A roadside target detection method for multimodal feature alignment and fusion according to claim 1, characterized in that, The method for extracting multiple point feature sets of the point cloud data includes the PointNet++ method.

6. A roadside target detection system for multi-modal feature alignment and fusion, based on the multi-modal feature alignment and fusion roadside target detection method according to any one of claims 1-5, characterized in that, The system includes: An extraction module configured to obtain the image to be subjected to 3D object detection and its corresponding point cloud data; extract the multi-scale feature maps of the input image; extract multiple point feature sets of the point cloud data; A first fusion module configured to map the point features in each of the multiple point feature sets onto the multi-scale feature maps respectively to obtain corresponding coordinates; obtain regional values according to the coordinates, and obtain the image features within multiple regions with the regional values added on the multi-scale feature maps as the first fused image features; A second fusion module configured to fuse the image features of the multi-scale feature maps and each of the first fused image features with a pre-constructed attention weight matrix respectively, and splice the image features of the fused multi-scale feature maps with the fused first fused image features respectively to obtain second fused image features, and splice the second fused image features with the corresponding point feature sets to obtain enhanced point cloud features; The method for obtaining the enhanced point cloud features is as follows: The point cloud feature F p in the point feature set, the image feature F i and the first fused image feature F′ i are weighted by multiple learnable weight matrices to obtain the attention weight matrices W attention and W′ attention : W attention = sigmoid(W4tanh(W2F p +W3F i )); W′ attention = sigmoid(W4tanh(W2F p + W3F′ i )); Among them, W2, W3, and W4 are learnable weight matrices; According to the said W attention 、the said W′ attention 、the said F i 、the said F′ i obtain the enhanced point cloud feature F′ p ; F′ p = C(C(W attention F i ∪ W′ attention F′ i ), F p ); Among them, C represents the feature concatenation operation, and ∪ represents the union operation; The result output module is configured to perform deconvolution processing on the image features of each multi-scale feature map, concatenate the deconvolved image features to obtain concatenated image features, fuse the concatenated image with the enhanced point cloud features to obtain a multi-modal feature fused point cloud; generate a 3D bounding box and a classification score from the multi-modal feature fused point cloud through a detection head, and output them as the 3D object detection result; the detection head is constructed based on convolutional layers.

7. A storage device in which multiple programs are stored, characterized in that, The program is applicable to be loaded and executed by a processor to implement a roadside object detection method for multi-modal feature alignment and fusion according to any one of claims 1-5.

8. A processing device includes a processor and a storage device; the processor is adapted to execute each program; the storage device is adapted to store multiple programs; characterized in that, The program is applicable to be loaded and executed by a processor to implement a roadside object detection method for multi-modal feature alignment and fusion according to any one of claims 1-5.

Citation Information

Patent Citations

  • Three-dimensional target detection method based on point cloud and image data fusion

    CN114092780A

  • Three-dimensional target detection and identification method based on multi-source fusion in vehicle-road cooperation scene

    CN114332494A