Efficient multi-modal fusion detection method and device based on sparse instance guidance
The sparse instance-guided method improves the accuracy and real-time performance of multi-modal fusion by converting and sampling features for efficient 3D target detection in complex traffic scenarios, overcoming the inefficiencies of complex network structures.
Patent Information
- Application Number
- CN202510285124.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-15
AI Technical Summary
The multimodal fusion detection method in the prior art relies on complex network structures with complex structures and large parameters, resulting in high computational complexity and low fusion process efficiency, reducing fusion effect, accuracy and real-timeness.
The sparse instance guidance method is adopted to convert the features of the lidar point cloud data and camera image data into bird's eye features. Through the sparse instance point cloud features as guidance, the key positions in the image bird's eye feature are sampled, sparse image features are obtained, and fused with the sparse instance point clouds to generate the target sparse fusion feature input to the sparse detection head network to generate three-dimensional object detection results.
It improves the accuracy and real-time nature of fusion detection, reduces the computational complexity, and improves the efficiency of the fusion process.
Smart Images

Figure CN120318786A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of intelligent driving, and particularly relates to an efficient multi-modal fusion detection method and device based on sparse instance guidance. Background Art
[0002] With the continuous expansion of the application scenarios of autonomous driving in intelligent networked vehicles, how to improve the performance of three-dimensional target detection systems in complex traffic scenarios has become an urgent problem to be solved. Due to its good accuracy and robustness, multi-modal sensor fusion is at the forefront of research in this field, and among them, the data fusion of camera images and lidar point clouds is particularly important.
[0003] In related technologies, through a method of directly fusing the point cloud voxel features and image perspective features output by the backbone network in space, or by choosing to transform the above-mentioned point cloud and image features into a unified two-dimensional bird's-eye view space, and then performing multi-modal information fusion in this unified space. And during the cross-modal fusion process, appropriate features are found through global query to achieve the fusion between different modal data.
[0004] However, the fusion detection methods in related technologies rely on complex network structures with complex structures and large numbers of parameters, resulting in high computational complexity and low efficiency in the fusion process, leading to poor real-time performance of the fusion perception model, and reducing the fusion effect, accuracy and real-time performance, which urgently need to be solved. Summary of the Invention
[0005] The present application provides an efficient multi-modal fusion detection method and device based on sparse instance guidance to solve the problems in related technologies that the fusion detection methods rely on complex network structures with complex structures and large numbers of parameters, resulting in high computational complexity, low efficiency in the fusion process, poor real-time performance of the fusion perception model, and reducing the fusion effect, accuracy and real-time performance.
[0006] The first aspect embodiment of the present application provides an efficient multi-modal fusion detection method based on sparse instance guidance, including the following steps: extracting at least one point cloud voxel feature from the lidar point cloud data in the target scene and at least one image perspective feature from the target camera image data; converting the at least one point cloud voxel feature into at least one target point cloud bird's-eye view feature, and converting the at least one image perspective feature into at least one target image bird's-eye view feature; sampling key positions in the at least one target point cloud bird's-eye view feature to obtain sparse instance point cloud features of the at least one target point cloud bird's-eye view feature; guiding by the sparse instance point cloud features, sampling key positions in the at least one target image bird's-eye view feature to obtain sparse image features, fusing the sparse instance point cloud features and the sparse image features to generate target sparse fusion features, and inputting the target sparse fusion features into a sparse detection head network to generate a three-dimensional object detection result of the target scene.
[0007] Optionally, in an embodiment of the present application, the converting the at least one image perspective feature into at least one target image bird's-eye view feature includes: encoding the at least one image perspective feature into a depth distribution feature enhanced by point cloud, and splicing the depth distribution feature and the at least one image perspective feature to generate a first spliced image feature; performing target processing on the spliced first spliced image feature to generate the depth distribution of each image pixel point in the target camera image data; using target voxel pooling to synthesize the at least one image perspective feature and the depth distribution of each image pixel point to generate the at least one target image bird's-eye view feature.
[0008] Optionally, in an embodiment of the present application, the sampling key positions in the at least one target point cloud bird's-eye view feature to obtain sparse instance point cloud features of the at least one target point cloud bird's-eye view feature includes: obtaining an importance-encoded image feature of the at least one target image bird's-eye view feature, and splicing the importance-encoded image feature and the at least one target point cloud bird's-eye view feature to generate a second spliced image feature; using the second spliced image feature to determine an instance heat map to generate the position and category of sparse instances according to the instance heat map; determining the sparse instance point cloud features based on the point cloud bird's-eye view feature corresponding to the position of the sparse instance and the category encoding feature corresponding to the category of the sparse instance.
[0009] Optionally, in an embodiment of the present application, the sampling of key positions in the at least one target image bird's-eye view feature guided by the sparse instance point cloud feature to obtain the sparse image feature of the at least one target image bird's-eye view feature includes: using the sparse instance point cloud feature as a query feature and using the image feature after importance encoding as a value feature; based on the query feature and the value feature, using a deformable attention mechanism to sample a preset number of feature points within a preset range of a target reference point to determine the sparse image feature in the at least one target image bird's-eye view feature.
[0010] Optionally, in an embodiment of the present application, the target sparse fusion feature is represented as:
[0011]
[0012] Wherein, is the target sparse fusion feature, is the sparse instance query feature, is the sparse image feature, ξ fus is the sparse feature fusion encoder.
[0013] An embodiment of the second aspect of the present application provides an efficient multi-modal fusion detection device based on sparse instance guidance, including: an extraction module for extracting at least one point cloud voxel feature in the lidar point cloud data and at least one image perspective feature in the target camera image data in a target scene; a conversion module for converting the at least one point cloud voxel feature into at least one target point cloud bird's-eye view feature and converting the at least one image perspective feature into at least one target image bird's-eye view feature; a sampling module for sampling key positions in the at least one target point cloud bird's-eye view feature to obtain the sparse instance point cloud feature of the at least one target point cloud bird's-eye view feature; a detection module for sampling key positions in the at least one target image bird's-eye view feature guided by the sparse instance point cloud feature to obtain a sparse image feature, fusing the sparse instance point cloud feature and the sparse image feature to generate a target sparse fusion feature, and inputting the target sparse fusion feature into a sparse detection head network to generate a three-dimensional target detection result of the target scene.
[0014] Optionally, in an embodiment of the present application, the conversion module includes: a splicing unit, configured to encode the at least one image perspective feature into a depth distribution feature enhanced by point cloud, and splice the depth distribution feature and the at least one image perspective feature to generate a first spliced image feature; a processing unit, configured to perform target processing on the spliced first spliced image feature to generate the depth distribution of each image pixel point in the target camera image data; a synthesis unit, configured to synthesize the at least one image perspective feature and the depth distribution of each image pixel point by using a target voxel pooling to generate the at least one target image bird's-eye view feature.
[0015] Optionally, in an embodiment of the present application, the sampling module includes: an acquisition unit, configured to acquire the importance-encoded image feature of the at least one target image bird's-eye view feature, and splice the importance-encoded image feature and the at least one target point cloud bird's-eye view feature to generate a second spliced image feature; a first determination unit, configured to use the second spliced image feature to determine an instance heat map, and generate the position and category of sparse instances according to the instance heat map; a second determination unit, configured to determine the sparse instance point cloud feature based on the point cloud bird's-eye view feature corresponding to the position of the sparse instance and the category-encoded feature corresponding to the category of the sparse instance.
[0016] Optionally, in an embodiment of the present application, the sampling module includes: a third determination unit, configured to use the sparse instance point cloud feature as a query feature and the importance-encoded image feature as a value feature; a sampling unit, configured to sample a preset number of feature points within a preset range of a target reference point based on the query feature and the value feature by using a deformable attention mechanism to determine the sparse image feature in the at least one target image bird's-eye view feature.
[0017] Optionally, in an embodiment of the present application, the target sparse fusion feature is expressed as:
[0018]
[0019] Wherein, is the target sparse fusion feature, is the sparse instance query feature, is the sparse image feature, ξ fus is the sparse feature fusion encoder.
[0020] An embodiment of the third aspect of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the efficient multi-modal fusion detection method based on sparse instance guidance as described in the above embodiments.
[0021] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the above-mentioned efficient multi-modal fusion detection method based on sparse instance guidance.
[0022] In the fifth aspect of the embodiments of the present application, a computer program product is provided, including a computer program, and when the computer program is executed, it is used to implement the above-mentioned efficient multi-modal fusion detection method based on sparse instance guidance.
[0023] The embodiments of the present application can convert the point cloud voxel features in the target scene into point cloud bird's-eye view features, and convert the image perspective features into image bird's-eye view features. Then, guided by the sparse instance point cloud features, key positions in the image bird's-eye view features are sampled to obtain sparse image features, which are fused with the sparse instance point cloud. The generated target sparse fusion features are input into the sparse detection head network to generate the three-dimensional target detection results of the target scene, effectively improving the fusion accuracy and real-time performance. Thus, it solves the problems in the related fusion detection methods that rely on complex network structures with complex structures and large numbers of parameters, resulting in high computational complexity, low efficiency in the fusion process, and reduced fusion effect, accuracy, and real-time performance.
[0024] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. Description of the Drawings
[0025] The above-mentioned and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0026] Figure 1 It is a flowchart of an efficient multi-modal fusion detection method based on sparse instance guidance provided according to an embodiment of the present application;
[0027] Figure 2 It is a schematic diagram of the principle of converting the enhanced image perspective features of the point cloud into the bird's-eye view space in a specific embodiment of the present application;
[0028] Figure 3 It is a schematic diagram of the principle of importance encoding of coordinate perception in a specific embodiment of the present application;
[0029] Figure 4 It is a schematic diagram of the principle of the instance-guided cross-modal adaptive fusion module in a specific embodiment of the present application;
[0030] Figure 5 It is a schematic structural diagram of an efficient multi-modal fusion detection device based on sparse instance guidance provided according to an embodiment of the present application;
[0031] Figure 6 Schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. Specific implementation manners
[0032] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation of the present application.
[0033] An efficient multi-modal fusion detection method and device based on sparse instance guidance according to an embodiment of the present application will be described below with reference to the accompanying drawings. For the problems in the related art fusion detection method mentioned in the above background technology that rely on a complex network structure with a complex structure and a large number of parameters, resulting in high computational complexity, low efficiency in the fusion process, and reduced fusion effect, accuracy, and real-time performance, etc., the present application provides an efficient multi-modal fusion detection method based on sparse instance guidance. In this method, the voxel features of the point cloud in the target scene can be converted into bird's-eye view features of the point cloud, and the perspective features of the image can be converted into bird's-eye view features of the image. Then, guided by the sparse instance point cloud features, key positions in the bird's-eye view features of the image are sampled to obtain sparse image features, which are fused with the sparse instance point cloud. The generated target sparse fusion features are input into the sparse detection head network to generate three-dimensional object detection results for the target scene, effectively improving the fusion accuracy and real-time performance. Thus, the problems in the related art fusion detection method that rely on a complex network structure with a complex structure and a large number of parameters, resulting in high computational complexity, low efficiency in the fusion process, and reduced fusion effect, accuracy, and real-time performance, etc., are solved.
[0034] Specifically, Figure 1 Flow chart of an efficient multi-modal fusion detection method based on sparse instance guidance provided by an embodiment of the present application.
[0035] As Figure 1 shown, the efficient multi-modal fusion detection method based on sparse instance guidance includes the following steps:
[0036] In step S101, at least one voxel feature of the lidar point cloud data in the target scene and at least one perspective feature of the target camera image data are extracted.
[0037] In an embodiment of the present application, the target scene is the scene of the current driving area of the vehicle.
[0038] It can be understood that, first, the embodiments of the present application can collect lidar point cloud data and camera image data in the vehicle driving area. Then, at least one point cloud voxel feature in the lidar point cloud data and at least one image perspective feature in the target camera image data are extracted through their respective backbone networks. Additionally, multiple point cloud voxel features and multiple image perspective features can also be extracted to increase the accuracy of the data, thereby effectively improving the executability of multi-modal fusion.
[0039] For example, let the lidar point cloud data be where is the number of laser points in the point cloud The element in the set represents a laser point, which includes the three-dimensional coordinates and reflection value of the corresponding laser point. Let the above lidar point cloud data be processed by the backbone network to generate the point cloud voxel feature
[0040] Next, let the camera image data be where is the number of multi-view surround cameras deployed on the autonomous vehicle. The element in the set is the image corresponding to the i-th camera, which includes the RGB three-channel pixel information with height H and width W. Let the above camera image data be processed by the backbone network to generate the image perspective feature where the element is the perspective feature corresponding to the i-th camera image.
[0041] In step S102, at least one point cloud voxel feature is converted into at least one target point cloud bird's-eye view feature, and at least one image perspective feature is converted into at least one target image bird's-eye view feature.
[0042] It can be understood that the embodiments of the present application can, based on the feature space normalization module for point cloud enhancement, convert at least one point cloud voxel feature obtained in the above steps into at least one target point cloud bird's-eye view feature, and convert at least one image perspective feature into at least one target image bird's-eye view feature. That is to say, the features of two heterogeneous modalities, namely the point cloud voxel feature and the image perspective feature, are transferred to a unified bird's-eye view space, realizing the reduction of the feature dimension to be fused from three-dimensional space to two-dimensional plane, effectively improving the multi-modal information fusion efficiency.
[0043] It should be noted that due to the inherent three-dimensional space information of the lidar point cloud, the point cloud voxel feature can be very naturally converted from the voxel space to the bird's-eye view space. Specifically, the present application uses a voxel encoding network to implement the view conversion of the point cloud voxel feature from the voxel space to the bird's-eye view space, which will not be elaborated here.
[0044] Among them, in an embodiment of the present application, converting at least one image perspective feature into at least one target image bird's-eye view feature includes: encoding at least one image perspective feature into a point cloud-enhanced depth distribution feature, and splicing the depth distribution feature with at least one image perspective feature to generate a first spliced image feature; performing target processing on the spliced first spliced image feature to generate the depth distribution of each image pixel point in the target camera image data; using a target voxel pooling to synthesize at least one image perspective feature and the depth distribution of each image pixel point to generate at least one target image bird's-eye view feature.
[0045] In the actual execution process, different from the fact that the voxel features of lidar point clouds are distributed in a three-dimensional space, the image perspective features of the camera in the embodiment of the present application are distributed in a two-dimensional plane space, and the key to converting it into a bird's-eye view space consistent with the point cloud features is to complete the conversion of the image perspective feature from a two-dimensional pixel coordinate system to a three-dimensional point cloud coordinate system. Specifically, according to the camera imaging principle, the core of the conversion between the above two coordinate systems lies in compensating for the lost depth information in camera imaging, such as Figure 2 As shown, different from the existing method that directly predicts the pixel depth based on image features, the embodiment of the present application can add cross-modal point cloud data to the prediction of the pixel depth distribution, aiming to use the accurate depth information of the point cloud to enhance the accuracy of the process of converting the image perspective feature into a bird's-eye view space.
[0046] Among them, for the image perspective feature corresponding to the i-th camera, the present application comprehensively considers the internal and external parameters and the point cloud within the image range, and encodes it into a point cloud-enhanced depth distribution feature based on formula (1) That is:
[0047]
[0048] Among them, ξ pe is a point cloud-enhanced depth feature encoder, which is specifically composed of various operators such as perspective projection, pooling, one-hot encoding, and convolution; is the lidar point cloud data; is the external parameter matrix of the i-th camera; is the internal parameter matrix of the i-th camera.
[0049] Then, the depth distribution feature is spliced with the image perspective feature and processed by a decoder δ pd composed of operators such as convolution and batch normalization as shown in formula (2), that is:
[0050]
[0051] Among them, is the depth distribution of each pixel in the image, is the perspective feature corresponding to the i-th camera image, is the depth distribution feature. Furthermore, the depth distribution of each pixel in the image is generated This depth distribution aggregates the accurate depth information provided by the point cloud data and the rich semantic information provided by the image features.
[0052] Finally, this application generates the bird's-eye view feature of the image based on voxel pooling by integrating the image perspective feature and the depth distribution corresponding to each image pixel point. Among them, to avoid the video memory occupation caused by constructing large-scale frustum point cloud features for display, this application uses the internal and external camera parameters to pre-index voxels and the corresponding frustum indices one by one, and accesses the image features and depth distributions based on the frustum indices, and then parallelly implements the multiplication and summation of the feature values and depth values within the same voxel index, thereby generating the bird's-eye view feature of the image, that is:
[0053]
[0054] Among them, is the bird's-eye view feature of the image, is the set of internal parameter matrices of the surround-view cameras, is the set of external parameter matrices of the surround-view cameras.
[0055] In step S103, sample the key positions in at least one target point cloud bird's-eye view feature to obtain the sparse instance point cloud feature of at least one target point cloud bird's-eye view feature.
[0056] It can be understood that the embodiments of this application are based on the point cloud bird's-eye view feature and the image bird's-eye view feature in the unified feature space. Use the sparse instance query feature sampling module to sample the key positions in the point cloud bird's-eye view feature in the following steps to obtain the sparse instance point cloud feature of the point cloud bird's-eye view feature. In the sampling process of the embodiments of this application, coordinate key encoding can be performed on multi-modal features, making full use of the differential advantages of heterogeneous sensor data in complex traffic environments, realizing the adaptive weight adjustment of multi-modal features during sampling, thereby realizing the dimensionality reduction of the feature size to be fused from dense two-dimensional to instance one-dimensional, and improving the model efficiency.
[0057] Optionally, in an embodiment of the present application, sampling is performed on key positions in at least one target point cloud bird's-eye view feature to obtain sparse instance point cloud features of at least one target point cloud bird's-eye view feature, including: obtaining importance-encoded image features of at least one target image bird's-eye view feature, and splicing the importance-encoded image features and at least one target point cloud bird's-eye view feature to generate second spliced image features; using the second spliced image features to determine an instance heat map, so as to generate the positions and categories of sparse instances according to the instance heat map; based on the point cloud bird's-eye view features corresponding to the positions of the sparse instances and the category encoding features corresponding to the categories of the sparse instances, determining sparse instance point cloud features.
[0058] As a possible implementation manner, an embodiment of the present application proposes a coordinate-aware feature importance encoding module, and its mathematical expression is shown in Equation (4):
[0059]
[0060] where is the image bird's-eye view feature after coordinate-aware encoding, is the grid coordinate matrix of the bird's-eye view space.
[0061] Based on the coordinate-aware feature importance encoding module, the weights of the modalities are dynamically and adaptively adjusted during the sampling process, so as to give full play to the advantages of different modality sensors, and can more comprehensively cover the targets to be detected in complex traffic scenes during the "dense -> sparse" dimensionality reduction process.
[0062] where, as Figure 3 shown, the importance encoder ξ ce includes operators such as average pooling, convolution, batch normalization, and rectified linear unit, and its calculation formula is shown in Equations (5)-(10):
[0063]
[0064]
[0065] where is the average pooling output corresponding to the feature of the w-th column, H b is the height of the bird's-eye view grid, (i, w) is the position coordinate of the i-th row and w-th column in the grid coordinate matrix x c (i, w) is the feature value of the corresponding position of the c-th channel in the image bird's-eye view feature ; is the average pooling output corresponding to the feature of the h-th row, W b is the width of the bird's-eye view grid, (h, j) is the position coordinate of the h-th row and j-th column in the grid coordinate matrix x c(h,j) is the bird's-eye view feature of the image The eigenvalue at the corresponding position of the c-th channel in
[0066] Based on the above average pooling output, the present application can construct an intermediate feature encoding the bird's-eye view spatial height and width position information, as shown in Equation (7):
[0067] f = ReLU(Conv(Cat(z h ,z w ))) (7)
[0068] Where z h and z w are the features aggregated along the channel dimension respectively, and f is the intermediate feature processed by concatenation, convolution and rectified linear activation function. and
[0069] Based on the intermediate feature f, the present application splits it into a height-direction independent tensor f h and a width-direction independent tensor f w along the spatial dimension, and generates the corresponding height-direction coordinate weight value g h and width-direction coordinate weight value g w through convolution and SigMoid activation function processing, that is:
[0070] g h = Sigmoid(Conv(f h )) (8)
[0071] g w = Sigmoid(Conv(f w )) (9)
[0072] Next, based on the height and width direction weight values of each coordinate, the present application generates the importance-encoded eigenvalue corresponding to each position in the feature map, as shown in Equation (10):
[0073]
[0074] Where is the weight value corresponding to the feature of the i-th row and c-th channel in in g h ; is the weight value corresponding to the feature of the j-th column and c-th channel in in g h ; y c (i,j) is the importance-encoded eigenvalue of x c (i,j), corresponding to the feature value of the i-th row, j-th column and c-th channel in the feature map.
[0075] Secondly, the image features after importance encoding are concatenated with the point cloud bird's-eye view features, and an instance heatmap is obtained through the processing of the instance feature encoding module: where N cls is the number of categories, and then the positions and categories of the sampled sparse instances are generated. Based on the point cloud bird's-eye view features corresponding to the positions of the sparse instances and the category encoding features corresponding to the categories of the sparse instances, the sparse instance point cloud features are determined, where:
[0076]
[0077]
[0078] where ξ ins is the instance heatmap encoder, which consists of convolutional and batch normalization operators; δ ins is the corresponding decoder, which consists of max pooling and sorting operators; is the set of sparse instance positions, and its value is {p q ∣q = 1, 2, …, N qy}, the size N qy of the set is the number of sparse instances, and the element p q in the set is the position index corresponding to the q-th instance; is the set of sparse instance categories, and its value is {c q ∣q = 1, 2, …, N qy}, and the element c q in the set is the category index corresponding to the q-th instance.
[0079] Then, the query feature of each instance is obtained by summing the point cloud bird's-eye view features corresponding to the instance position
[0080]
[0081] where is the category encoder, which consists of one-hot encoding and convolutional operators; c q is the category index corresponding to the q-th instance.
[0082] Finally, the query features corresponding to all instances jointly form the sparse instance query feature set where the number N qy of query features is much smaller than the size H b ×W b of the bird's-eye view network, thus realizing effective dimensionality reduction of the aggregation range between the two modalities, and then improving the efficiency of multimodal fusion for instances in a targeted manner.
[0083] In step S104, guided by the sparse instance point cloud features, key positions in at least one target image bird's-eye view feature are sampled to obtain sparse image features. The sparse instance point cloud features and the sparse image features are fused to generate target sparse fusion features, and the target sparse fusion features are input into a sparse detection head network to generate a 3D object detection result for the target scene.
[0084] It can be understood that the embodiments of the present application can sample key positions in the image bird's-eye view feature in the following steps guided by the sparse instance point cloud features. For example, the embodiments of the present application can sample the adaptive sparse image features obtained in the image bird's-eye view feature based on the deformable attention mechanism guided by the sparse instance point cloud features. Then, the sparse instance query features in the point cloud bird's-eye view feature and the adaptive sparse image features sampled in the image bird's-eye view feature based on the deformable attention mechanism are input into a sparse feature fusion encoder to generate cross-modal sparse fusion features. Secondly, the cross-modal sparse fusion features are input into a sparse detection head network constructed based on a feed-forward neural network to generate a 3D object detection result for the target scene, effectively reducing the computational complexity, improving the fusion efficiency, and enhancing the robustness and real-time performance of multi-modal fusion detection.
[0085] Among them, in an embodiment of the present application, sampling key positions in at least one target image bird's-eye view feature guided by the sparse instance point cloud features to obtain sparse image features of at least one target image bird's-eye view feature includes: using the sparse instance point cloud features as query features and the image features encoded with importance as value features; based on the query features and the value features, using the deformable attention mechanism to sample a preset number of feature points within a preset range of the target reference point to determine the sparse image features in at least one target image bird's-eye view feature.
[0086] In the embodiments of the present application, the preset range is the range near the reference point, which is specifically set by those skilled in the art and will not be specifically limited here.
[0087] In some embodiments, the embodiments of the present application can propose an instance-guided cross-modal adaptive fusion module based on the sampled sparse instance point cloud features and the corresponding bird's-eye view grid position coordinates, specifically as Figure 4 shown. For the q-th sparse instance query feature in the adaptive fusion module uses its position in the bird's-eye view grid space as the reference point and the linearly transformed image bird's-eye view feature as the dense value feature, and performs sparse sampling based on the deformable attention near the reference point, and then adaptively selects appropriate image features for fusion, so as to achieve efficient and robust aggregation across modalities.
[0088] Specifically, through operations such as linear transformation on , the sampling position offsets and corresponding attention weights of different attention heads and sampling points can be obtained. Based on the former, sparse features are sampled from the dense value features, and multiplied and summed with the latter attention weights to obtain the sparse sampling value features corresponding to different attention heads, thereby generating adaptive sparse image features, and finally generating sparse fusion features after fusion coding with sparse query features.
[0089] First, the embodiment of the present application can be guided by sparse instance point cloud features, and the sparse instance point cloud features are used as query features, and the image bird's-eye view features after importance coding are used as value features. Based on the deformable attention mechanism, K feature points are further sampled near the reference point p q to adaptively extract the required sparse image features Among them, the present application uses the position of the sparse instance as the reference point, and the specific process is shown in formula (14):
[0090]
[0091] Among them, V img is the image bird's-eye view feature after importance coding, the value feature after linear transformation, M is the number of layers of multi-head attention, W m is the weight matrix corresponding to the m-th layer attention head, k is the index of the sampling point in the deformable attention mechanism, K is the total number of sampling points, and respectively represent the attention weight and sampling position offset corresponding to the k-th sampling point of the q-th query feature in the m-th layer attention head, and both are obtained by linear transformation; is the feature value at the position img in V .
[0092] Then, the present application inputs the sparse instance query features in the point cloud bird's-eye view feature and the adaptive features sampled from the image bird's-eye view feature based on the deformable attention mechanism into the sparse feature fusion encoder ξ fus to generate cross-modal sparse fusion features The specific process is shown in formula (15):
[0093]
[0094] Among them, is the cross-modal sparse fusion feature, that is, the target sparse fusion feature; For sparse instance query features; For sparse image features; ξ fus For a sparse feature fusion encoder, which is composed of operators such as convolution, batch normalization, and rectified linear activation function.
[0095] Finally, this application combines the sparse instance query feature set Each instance feature in Performs the adaptive fusion process shown in equations (13) and (14) to obtain its corresponding cross-modal sparse fusion feature And then forms a sparse fusion feature set
[0096] In summary, the embodiment of this application uses the sparse instance point cloud features obtained by sampling as the query guidance, adaptively searches for the required features in the image features for fusion, and alleviates the impact of feature misalignment on fusion detection. In particular, different from the global cross-modal fusion strategy of existing methods, this application uses the sparse instance position as a reference point and only performs deformable sampling near the reference point, realizing an effective dimensionality reduction of the cross-modal query range from "sparse-dense" to "sparse-sparse", and further improving the efficiency of the fusion process.
[0097] Furthermore, the embodiment of this application can construct a sparse detection head network based on a feedforward neural network, and use the sparse fusion feature set As the input, output the target category and its three-dimensional bounding box in the surrounding environment.
[0098] Specifically, for the q-th sparse feature The target category classification head outputs the probability distribution corresponding to N cls Categories Similarly, the three-dimensional bounding box parameter regression head outputs regression parameters Which includes the center point position offset Relative ground height Size Yaw angle vector Speed Therefore, the three-dimensional bounding box information corresponding to each sparse feature point can be encoded as a 10-tuple as shown in equation (16), that is:
[0099]
[0100] During the inference process, the above bounding box regression parameters Can be post-processed to generate the final three-dimensional bounding box information
[0101] Finally, based on the target category classification head and the three-dimensional bounding box regression head, the detection results corresponding to the sparse features Are generated Furthermore, by inputting each sparse feature in into the above head network for parallel processing, the detection result corresponding to the current frame is finally obtained. Where the number N of targets in the detection result det is the same as the number N of query features in qy which is equal.
[0102] According to the efficient multi-modal fusion detection method based on sparse instance guidance proposed in the embodiments of the present application, the voxel features of the point cloud in the target scene can be converted into bird's-eye view features of the point cloud, and the perspective features of the image can be converted into bird's-eye view features of the image. Then, guided by the sparse instance point cloud features, key positions in the bird's-eye view features of the image are sampled to obtain sparse image features, which are fused with the sparse instance point cloud. The generated target sparse fusion features are input into the sparse detection head network to generate the three-dimensional target detection result of the target scene, effectively improving the fusion accuracy and real-time performance. Thus, the problems in the related fusion detection methods that rely on complex network structures with complex structures and large numbers of parameters, resulting in high computational complexity, low efficiency in the fusion process, and reduced fusion effect, accuracy, and real-time performance are solved.
[0103] Secondly, a high-efficiency multi-modal fusion detection device based on sparse instance guidance proposed in the embodiments of the present application is described with reference to the accompanying drawings.
[0104] Figure 5 is a block diagram of the high-efficiency multi-modal fusion detection device based on sparse instance guidance according to the embodiments of the present application.
[0105] As Figure 5 shown, the high-efficiency multi-modal fusion detection device 10 based on sparse instance guidance includes: an extraction module 100, a conversion module 200, a sampling module 300, and a detection module 400.
[0106] Specifically, the extraction module 100 is configured to extract at least one voxel feature of the lidar point cloud data and at least one perspective feature of the target camera image data in the target scene.
[0107] The conversion module 200 is configured to convert at least one voxel feature into at least one target bird's-eye view feature of the point cloud, and convert at least one perspective feature into at least one target bird's-eye view feature of the image.
[0108] The sampling module 300 is configured to sample key positions in at least one target bird's-eye view feature of the point cloud to obtain sparse instance point cloud features of at least one target bird's-eye view feature of the point cloud.
[0109] The detection module 400 is configured to sample key positions in at least one target image bird's-eye view feature guided by sparse instance point cloud features to obtain sparse image features, fuse the sparse instance point cloud features and the sparse image features to generate target sparse fusion features, and input the target sparse fusion features into a sparse detection head network to generate 3D object detection results of the target scene.
[0110] Optionally, in an embodiment of the present application, the conversion module 200 includes: a splicing unit, a processing unit, and an integration unit.
[0111] Among them, the splicing unit is configured to encode at least one image perspective feature into a depth distribution feature enhanced by point cloud, and splice the depth distribution feature and at least one image perspective feature to generate a first spliced image feature.
[0112] The processing unit is configured to perform target processing on the spliced first spliced image feature to generate the depth distribution of each image pixel point in the target camera image data.
[0113] The integration unit is configured to integrate at least one image perspective feature and the depth distribution of each image pixel point by using target voxel pooling to generate at least one target image bird's-eye view feature.
[0114] Optionally, in an embodiment of the present application, the sampling module 300 includes: an acquisition unit, a first determination unit, and a second determination unit.
[0115] Among them, the acquisition unit is configured to acquire the importance-encoded image features of at least one target image bird's-eye view feature, and splice the importance-encoded image features and at least one target point cloud bird's-eye view feature to generate a second spliced image feature.
[0116] The first determination unit is configured to determine an instance heat map by using the second spliced image feature, and generate the position and category of sparse instances according to the instance heat map.
[0117] The second determination unit is configured to determine sparse instance point cloud features based on the point cloud bird's-eye view features corresponding to the positions of the sparse instances and the category encoding features corresponding to the categories of the sparse instances.
[0118] Optionally, in an embodiment of the present application, the sampling module 300 includes: a third determination unit and a sampling unit.
[0119] Among them, the third determination unit is configured to use the sparse instance point cloud features as query features and the importance-encoded image features as value features.
[0120] A sampling unit, configured to sample a preset number of feature points within a preset range of a target reference point based on a query feature and a value feature by using a deformable attention mechanism, so as to determine sparse image features in at least one target image bird's-eye view feature.
[0121] Optionally, in an embodiment of the present application, the target sparse fusion feature is represented as:
[0122]
[0123] Wherein, is the target sparse fusion feature, is the sparse instance query feature, is the sparse image feature, ξ fus is the sparse feature fusion encoder.
[0124] It should be noted that the foregoing explanation of the embodiment of the efficient multi-modal fusion detection method based on sparse instance guidance is also applicable to the efficient multi-modal fusion detection device based on sparse instance guidance in this embodiment, and will not be elaborated here.
[0125] The efficient multi-modal fusion detection device based on sparse instance guidance proposed according to the embodiment of the present application can convert the point cloud voxel features in the target scene into point cloud bird's-eye view features, and convert the image perspective features into image bird's-eye view features. Then, guided by the sparse instance point cloud features, sample the key positions in the image bird's-eye view features to obtain sparse image features, and fuse them with the sparse instance point cloud. The generated target sparse fusion features are input into the sparse detection head network to generate the three-dimensional target detection results of the target scene, effectively improving the fusion accuracy and real-time performance. Thus, it solves the problems in the related fusion detection methods that rely on complex network structures with complex structures and large numbers of parameters, resulting in high computational complexity, low efficiency in the fusion process, and reduced fusion effect, accuracy and real-time performance.
[0126] Figure 6 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device may include:
[0127] A memory 601, a processor 602, and a computer program stored on the memory 601 and executable on the processor 602.
[0128] When the processor 602 executes the program, it implements the efficient multi-modal fusion detection method based on sparse instance guidance provided in the above embodiment.
[0129] Further, the electronic device further includes:
[0130] A communication interface 603, configured for communication between the memory 601 and the processor 602.
[0131] A memory 601 for storing a computer program that can run on a processor 602.
[0132] The memory 601 may include a high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0133] If the memory 601, the processor 602, and the communication interface 603 are implemented independently, the communication interface 603, the memory 601, and the processor 602 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 6 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0134] Optionally, in a specific implementation, if the memory 601, the processor 602, and the communication interface 603 are integrated on a single chip, the memory 601, the processor 602, and the communication interface 603 can communicate with each other through an internal interface.
[0135] The processor 602 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0136] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the efficient multi-modal fusion detection method based on sparse instance guidance as described above is implemented.
[0137] This embodiment also provides a computer program product, including a computer program, and when the computer program is executed, it is used to implement the efficient multi-modal fusion detection method based on sparse instance guidance as described above.
[0138] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0139] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0140] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or N executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a manner that is not in the order shown or discussed, including in a substantially simultaneous manner according to the functions involved or in the reverse order, which should be understood by those skilled in the art to which the embodiments of this application belong.
[0141] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered a definitional sequence list of executable instructions for implementing logical functions, and can be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion (electronic device) having one or N wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0142] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0143] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0144] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0145] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. An efficient multimodal fusion detection method based on sparse instance guidance, characterized in that Including the following steps: Extracting at least one point cloud voxel feature from the lidar point cloud data and at least one image perspective feature from the target camera image data in the target scene; Converting the at least one point cloud voxel feature into at least one target point cloud bird's-eye view feature, and converting the at least one image perspective feature into at least one target image bird's-eye view feature; Sampling key positions in the at least one target point cloud bird's-eye view feature to obtain sparse instance point cloud features of the at least one target point cloud bird's-eye view feature; Guided by the sparse instance point cloud features, sampling key positions in the at least one target image bird's-eye view feature to obtain sparse image features, fusing the sparse instance point cloud features and the sparse image features to generate target sparse fusion features, and inputting the target sparse fusion features into a sparse detection head network to generate a 3D object detection result of the target scene.
2. The method according to claim 1, wherein The converting the at least one image perspective feature into at least one target image bird's-eye view feature includes: Encoding the at least one image perspective feature into a depth distribution feature enhanced by point clouds, and splicing the depth distribution feature and the at least one image perspective feature to generate a first spliced image feature; Performing target processing on the spliced first spliced image feature to generate the depth distribution of each image pixel point in the target camera image data; Using target voxel pooling to synthesize the at least one image perspective feature and the depth distribution of each image pixel point to generate the at least one target image bird's-eye view feature.
3. The method according to claim 1, wherein The sampling key positions in the at least one target point cloud bird's-eye view feature to obtain sparse instance point cloud features of the at least one target point cloud bird's-eye view feature includes: Obtaining an importance-encoded image feature of the at least one target image bird's-eye view feature, and splicing the importance-encoded image feature and the at least one target point cloud bird's-eye view feature to generate a second spliced image feature; Using the second spliced image feature to determine an instance heatmap to generate the positions and categories of sparse instances according to the instance heatmap; Based on the point cloud bird's-eye view feature corresponding to the position of the sparse instance and the category encoding feature corresponding to the category of the sparse instance, determining the sparse instance point cloud feature.
4. The method according to claim 3, characterized in that, The guiding by the sparse instance point cloud features to sample key positions in the at least one target image bird's-eye view feature to obtain sparse image features of the at least one target image bird's-eye view feature includes: Using the sparse instance point cloud feature as a query feature and the importance-encoded image feature as a value feature; Based on the query feature and the value feature, using a deformable attention mechanism to sample a preset number of feature points within a preset range of a target reference point to determine the sparse image features in the at least one target image bird's-eye view feature.
5. The method according to claim 1, wherein The target sparse fusion feature is represented as: Among them, is the target sparse fusion feature, is the sparse instance query feature, is the sparse image feature, ξ fus is the sparse feature fusion encoder.
6. An efficient multi-modal fusion detection device based on sparse instance guidance, characterized in that, Including: An extraction module for extracting at least one point cloud voxel feature from the lidar point cloud data and at least one image perspective feature from the target camera image data in the target scene; A conversion module, configured to convert the at least one point cloud voxel feature into at least one target point cloud bird's-eye view feature, and convert the at least one image perspective feature into at least one target image bird's-eye view feature; A sampling module, configured to sample key positions in the at least one target point cloud bird's-eye view feature to obtain sparse instance point cloud features of the at least one target point cloud bird's-eye view feature; A detection module, configured to sample key positions in the at least one target image bird's-eye view feature guided by the sparse instance point cloud features to obtain sparse image features, fuse the sparse instance point cloud features and the sparse image features to generate target sparse fusion features, and input the target sparse fusion features into a sparse detection head network to generate 3D object detection results of the target scene.
7. The device according to claim 6, characterized in that, The conversion module includes: A splicing unit, configured to encode the at least one image perspective feature into a depth distribution feature enhanced by point clouds, and splice the depth distribution feature and the at least one image perspective feature to generate a first spliced image feature; A processing unit, configured to perform target processing on the spliced first spliced image feature to generate the depth distribution of each image pixel point in the target camera image data; An integration unit, configured to integrate the at least one image perspective feature and the depth distribution of each image pixel point by using target voxel pooling to generate the at least one target image bird's-eye view feature.
8. An electronic device, characterized in that, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the efficient multi-modal fusion detection method based on sparse instance guidance according to any one of claims 1-5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to be used for implementing the efficient multi-modal fusion detection method based on sparse instance guidance according to any one of claims 1-5.
10. A computer program product, comprising a computer program, characterized in that, The computer program is executed by the processor to be used for implementing the efficient multi-modal fusion detection method based on sparse instance guidance according to any one of claims 1-5.