Three-dimensional object detection method based on multi-view and point cloud BEV feature fusion
By adopting the multi-view and point cloud BEV feature fusion method in the three-dimensional object detection technology, and using the cross attention mechanism to perform adaptive fusion of multi-sensor data, the problem of simple information loss and feature fusion in the prior art is solved, and higher detection accuracy and environmental perception capabilities are achieved.
Patent Information
- Application Number
- CN202510258998.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-06
AI Technical Summary
In the existing three-dimensional object detection technology, machine vision-based methods lack depth information, resulting in inaccurate positioning; lidar-based methods lack semantic information, which easily confuses the foreground and background, and the sparseness of point clouds affects the recognition accuracy of remote objects and small objects. The multi-sensor fusion method has information loss and noise introduction, and the feature fusion process is too simple, which hinders the improvement of the accuracy of the detection algorithm.
A three-dimensional object detection method based on the fusion of multi-view and point cloud BEV features is adopted, and point cloud data is acquired through a surround view camera, lidar voxel features and machine vision voxel features are generated, and they are spread to the BEV space for compression and fusion. The multi-source heterogeneous feature adaptive fusion module using the cross attention mechanism is used to realize the adaptive fusion of multi-sensor data.
Through the multi-view rotary voxel encoder and the cross attention mechanism fusion module, the alignment matching and weight ratio analysis of multi-sensor data features are realized, the accuracy and robustness of three-dimensional object detection are improved, and the environmental perception ability of intelligent vehicles is enhanced.
Smart Images

Figure CN119785012B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of environmental perception technology, and in particular to a three-dimensional target detection method based on multi-view and point cloud BEV feature fusion. Background Art
[0002] As the primary task of autonomous driving, environmental perception is crucial. In particular, the accurate perception of the three-dimensional geometric structure of scenes or objects in the environment has become a key technology of the environmental perception system, an indispensable prerequisite for achieving autonomous and safe driving of unmanned vehicles, and also the core focus and difficulty of research.
[0003] The current mainstream 3D target detection technologies can be mainly divided into machine vision-based methods, LiDAR-based methods, and multi-sensor fusion-based methods. Among them, the machine vision-based methods cannot accurately locate 3D targets because the images do not contain depth information, and the detection accuracy is difficult to improve; the point cloud collected by the LiDAR-based method lacks semantic information, so it may confuse the foreground and background with similar structures, and cause false detection, interfering with normal driving. At the same time, the sparsity of the point cloud affects the accuracy of the LiDAR-only method in identifying remote objects and small objects.
[0004] Therefore, methods based on multi-sensor fusion have gradually attracted attention. The current mainstream multi-sensor fusion method is to convert the point cloud data collected by the lidar and the image data collected by the camera into BEV features, and then perform feature stitching. However, this method has the following problems: First, the image data is mostly front view, and the process of converting to BEV features is usually accompanied by information loss and noise introduction; second, the feature fusion process in the form of direct splicing is too simple and crude, and the feature alignment and weight ratio analysis of multiple sensors are not realized, which hinders the further improvement of the accuracy of this fusion detection algorithm. Summary of the invention
[0005] In view of the problems existing in the existing methods, the invention provides a three-dimensional target detection method based on the fusion of multi-view and point cloud BEV features, in order to achieve the adaptive fusion of BEV features of multi-source heterogeneous data in the three-dimensional detection process, thereby ensuring the accuracy of the perceived environment assessment.
[0006] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solution: a three-dimensional target detection method based on multi-view and point cloud BEV feature fusion, comprising the following steps:
[0007] S1, image and point cloud data acquisition: collect multi-view images of the target scene through the surround view camera, and use the laser radar to obtain the point cloud data of the target scene;
[0008] S2, lidar voxel feature generation: convert the lidar point cloud data into a voxel sequence, extract the set of non-empty voxel center points, calculate the features of each voxel, and obtain the lidar voxel features;
[0009] S3, machine vision voxel feature generation: Use the multi-view to voxel encoder to project the center point set of non-empty voxels onto the multi-view image, and use the fully connected layer and deformable Transformer network to extract the center voxel features; at the same time, project the point cloud in the non-empty voxels onto the multi-view image, use the bilinear interpolation method to obtain the point cloud voxel features, and splice the center voxel features with the point cloud voxel features to obtain the machine vision voxel features;
[0010] S4, BEV feature generation and compression: diffuse the machine vision voxel features and the lidar voxel features into the BEV space and compress them to obtain the BEV features of the machine vision and lidar;
[0011] S5, BEV feature fusion: Using the multi-source heterogeneous feature adaptive fusion module based on the cross-attention mechanism, the image weight and radar weight are calculated using the cross-attention mechanism, and the BEV features of machine vision and lidar are weighted and fused at the feature channel level to obtain the fused BEV features;
[0012] S6. 3D object detection: Input the fused BEV features into the detection network to obtain 3D object detection information, including object category, size, position, spatial direction and confidence.
[0013] Preferably, the S3 comprises the following steps: S31: multi-view projection: projecting points in the non-empty voxel center point set onto multiple views of the multi-view image of the frame to obtain a multi-view projection voxel center point set;
[0014] S32: Central voxel feature extraction: Through the fully connected layer and the deformable Transformer network, the central point features of each view are extracted from the set of projected voxel center points of each view, and then the central point features of each view are spliced into the central voxel feature.
[0015] S33: Point cloud voxel feature acquisition: Project the point cloud in non-empty voxels onto the multi-view image and use bilinear interpolation to obtain the point cloud voxel features;
[0016] S34: Feature fusion: The central voxel feature and the point cloud voxel feature are concatenated along the feature channel, and the machine vision voxel feature is obtained through the fully connected layer.
[0017] Preferably, S5 comprises the following steps:
[0018] S51: Initialization and input compression: define the number of iterations K and initialize, compress the machine vision BEV features and the lidar BEV features as the input of the current iteration;
[0019] S52: Map to matrix: Map the machine vision and lidar BEV inputs to the Query, Key, and Value matrices of the image and lidar, respectively;
[0020] S53: Calculate radar weight: Calculate radar attention weight matrix by multiplying image query and radar key, and multiply it with radar value matrix to generate radar weight;
[0021] S54: Calculate image weight: Calculate the image attention weight matrix by multiplying the radar query and the image key, and multiply it with the image value matrix to generate the image weight;
[0022] S55: Update BEV input: multiply the radar weight by the lidar BEV input, and multiply the image weight by the machine vision BEV input to obtain updated machine vision and lidar BEV inputs;
[0023] S56: Iterative update: update the BEV input of the machine vision and the laser radar and return to step S52 until a predetermined number of iterations K is reached;
[0024] S57: Feature fusion and output: Concatenate the updated lidar and machine vision BEV inputs, and generate fusion features through the fully connected layer to obtain the final fused BEV features.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] 1. In order to solve the problems of information loss and noise introduction in the process of converting image features to BEV features, the present invention proposes a multi-view to voxel encoder, which projects the voxel centers of a preset voxel frame into the multi-view image, and encodes the image features into preset voxel features and then converts them into BEV features through interpolation method and deformable Transformer network framework, thereby realizing the alignment and matching of heterogeneous features in the feature conversion process, preserving more useful information, and helping to improve the accuracy of target detection based on multi-sensor fusion.
[0027] 2. The present invention proposes a multi-source heterogeneous feature adaptive fusion module based on the cross-attention mechanism, which uses the cross-attention mechanism network to dynamically estimate the correlation between the two modalities and obtain the weight sparsity between the multi-modal features, so that the network can automatically align the multi-source heterogeneous data, which is conducive to improving the environmental perception ability of intelligent vehicles. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is an overall flow chart of the three-dimensional target detection method based on multi-view and point cloud BEV feature fusion of the present invention;
[0029] Figure 2 FIG. 1 is a framework diagram of a multi-view voxel encoder according to the present invention;
[0030] Figure 3 It is a framework diagram of the multi-source heterogeneous feature adaptive fusion module based on the cross attention mechanism of the present invention;
[0031] Figure 4 This is a diagram showing the machine vision detection effect of the present invention. DETAILED DESCRIPTION
[0032] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
[0033] In this embodiment, a 3D target detection method based on multi-view and point cloud BEV feature fusion is provided. Figure 1 As shown, the following steps are included:
[0034] S1: For Frame, using the surround camera to collect images of the target scene , and spliced into Frame multi-view image , wherein the Frame multi-view image ,in, Indicates Frame multi-view image Height, Indicates Frame multi-view image The width of Frame multi-view image The number of RGB channels; at the same time, the laser radar is used to collect the point cloud data of the target scene and output the Frame LiDAR point cloud sequence , wherein the Frame LiDAR point cloud sequence The dimension is ,in, Indicates the number of point cloud data. The number of information representing each point cloud data; the information includes the center coordinates And the reflection intensity ;
[0035] S2: Using the dynamic voxel method Frame LiDAR point cloud sequence Convert to Frame LiDAR Voxel Sequence , and obtain the Frame LiDAR Voxel Features The center point set of non-empty voxels , then Frame LiDAR Voxel Sequence will be converted into Frame LiDAR Voxel Features ; wherein the Frame LiDAR Voxel Sequence In the vehicle coordinate system The range in the axial direction is , The range in the axial direction is , The range in the axial direction is The partition size in the space is A collection of voxels with dimension ,in, represents the number of non-empty voxels, Represents the maximum number of point clouds in each voxel, Represents the number of features of each point cloud in a voxel. Let the number of point clouds contained in a voxel be ,like , then the point cloud in the voxel is randomly downsampled until the number of point clouds is ,like , then the remaining feature dimensions are set to zero; the non-empty voxel center point set The dimension is , where 3 represents the coordinates of the center point of the non-empty voxel ; said Frame LiDAR Voxel Features It is by Frame LiDAR Voxel Sequence For every non-empty voxel in The feature is averaged and its dimension is ;
[0036] S3: Using multi-view voxel encoder, such as Figure 2 As shown, according to Frame multi-view image And the set of non-empty voxel centers , generating Frame Machine Vision Voxel Features , among which, Frame Machine Vision Voxel Features The dimensions and Frame LiDAR Voxel Features Stay consistent, for ;
[0037] S4: Frame Machine Vision Voxel Features With Frame LiDAR Voxel Features Diffused to the vehicle coordinate system The range in the axial direction is , The range in the axial direction is , The range in the axial direction is The partition size in the space is In the collection of voxels, and in Compression is performed in the axial direction to obtain the Frame Machine Vision BEV Features With Frame LiDAR BEV Features , among which, Frame Machine Vision BEV Features With Frame LiDAR BEV Features The characteristics of ,in is the number of characteristic channels of BEV features;
[0038] S5: Using the multi-source heterogeneous feature adaptive fusion module based on the cross attention mechanism, such as Figure 3 As shown, for Frame Machine Vision BEV Features With Frame LiDAR BEV Features Perform feature fusion to obtain the Frame fusion BEV features , among which, Frame fusion BEV features The dimension is ;
[0039] S6: Frame fusion BEV features Input into the detection head network and obtain the Frame 3D target detection information, where The frame 3D target detection information includes: Frame 3D object category , No. 3D prediction box size of the frame 3D target , No. 3D prediction box position of the frame 3D target , No. The spatial orientation of the 3D prediction box of the 3D target in the frame and the confidence of the network's final prediction ;No. 3D prediction box size of the frame 3D target Includes: Long ,Width and high ;No. 3D prediction box position of the frame 3D target include: , the detection effect diagram is as follows Figure 4 As shown, the red box in the figure represents the final prediction box detected by the network.
[0040] Among them, S3 includes the following steps, such as Figure 2 As shown:
[0041] S31: Set the center points of non-empty voxels All points in the projected Frame multi-view image of On each view, obtain the set of multi-view projection voxel center points , where for the set of non-empty voxel centers Any point in , projected to The corresponding projection points obtained on the view are , the multi-view projection voxel center point set for The set of corresponding projection points on the views has a dimension of ,represent The center of non-empty voxels is The pixel coordinate system coordinates of the corresponding projection point on each view;
[0042] S32: Using multi-view projection voxel center point set and Get the central voxel feature , the steps are as follows: for the First, the non-empty voxel center point set Projection to Get the first The set of projected voxel centers of the views , whose dimensions are ; then the The center point set of the projected voxels of the view is input into the fully connected layer to obtain the The center point query sequence of the view , whose dimensions are ,in For the The center point query sequence of the view The number of feature channels The center point query sequence of the view With Views Input the deformable Transformer network and get the Center point feature of each view , whose dimensions are ; Finally, The center point features of the views are spliced along the feature channel to obtain the Frame center voxel features , whose dimensions are ;
[0043] S33: Get point cloud voxel features , the steps are as follows: for the A view will The point cloud in the non-empty voxels is projected onto the In each view, for any non-empty voxel feature, the projections of all point clouds contained in it are interpolated to the location of the projection point of the voxel center by bilinear interpolation, so as to obtain the first Frame point cloud voxel features , whose dimensions are ;
[0044] S34: Frame center voxel features With Frame point cloud voxel features Splicing along the feature channel, the dimension is , and then through a fully connected layer to obtain the Frame Machine Vision Voxel Features , the dimension is .
[0045] Among them, S5 includes the following steps: Figure 3 As shown:
[0046] S51: Define the current number of iterations as , and initialize =0, let K represent the total number of iterations;
[0047] The first Frame Machine Vision BEV Features With Frame LiDAR BEV Features Compressed to dimensions is No. Iteration 1 of machine vision BEV input and LiDAR BEV input for the first iteration ;
[0048] S52: Iteration 1 of machine vision BEV input Mapping to image Query matrix , image key matrix and the image Value matrix ,Right now:
[0049]
[0050] In formula (1) to formula (3), , , They are all learnable linear transformation matrices;
[0051] The first LiDAR BEV input for the first iteration Mapping to radar query matrix , radar key matrix and radar value matrix ,Right now:
[0052]
[0053] In formula (4) to formula (6), , , They are all learnable linear transformation matrices;
[0054] S53: Image Query matrix With radar Key matrix Multiply and make Processing, obtain radar attention weight matrix ,Right now:
[0055] (7)
[0056] in, Radar attention weight matrix representing the output The number of feature channels.
[0057] Then the radar attention weight matrix and radar value matrix Multiply and pass through a fully connected network and a Network, generate Radar weights for iterations ;
[0058] S54: Radar Query Matrix With Key matrix Multiply and make Processing, obtain the image attention weight matrix ,Right now:
[0059] (8)
[0060] In formula (8), Image attention weight matrix representing the output The number of feature channels, whose value is related to the radar attention weight matrix The number of feature channels is the same.
[0061] Then the image attention weight matrix and the image Value matrix Multiply and pass through a fully connected network and a Network, generate The image weights for the iteration ;
[0062] S55: Radar weights for iterations With LiDAR BEV input for the first iteration Multiply them together to get LiDAR BEV input for the first iteration ; The image weights for the iteration With Iteration 1 of machine vision BEV input Multiply them together to get Iteration 1 of machine vision BEV input ;
[0063] S56: Assign to ,Will Assign to ,Will Assign to After that, return to step S52 and execute sequentially until = K, and get the LiDAR BEV input for the first iteration With Iteration 1 of machine vision BEV input ;
[0064] S57: LiDAR BEV input for the first iteration With Iteration 1 of machine vision BEV input Splicing along the feature channel and passing through a fully connected network, we get a dimension of The fusion feature is expanded in dimension to obtain the Frame fusion BEV features , whose dimensions are .
[0065] Although the embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
Claims
1. A three-dimensional target detection method based on multi-view and point cloud BEV feature fusion, characterized in that: The following steps are involved: S1, collect multi-view images and point cloud data of the target scene; S2, converting the laser radar point cloud data into a voxel sequence, extracting a set of non-empty voxel center points, calculating the features of each voxel, and obtaining the laser radar voxel features; S3, extract the central voxel feature, project the point cloud in the non-empty voxel onto the multi-view image, use the bilinear interpolation method to obtain the point cloud voxel feature, and splice the central voxel feature with the point cloud voxel feature to obtain the machine vision voxel feature; S4, diffusing the machine vision voxel features and the laser radar voxel features into the BEV space and compressing them to obtain the BEV features of the machine vision and the laser radar; S5, weighted fusion of the BEV features of machine vision and lidar at the feature channel level to obtain fused BEV features; The specific steps include: S51, initialization and input compression: define the number of iterations K and initialize, compress the machine vision BEV features and the laser radar BEV features as the input of the current iteration; S52, Map to Matrix: Map the machine vision BEV input to the Query matrix, Key matrix, and Value matrix of the image, and map the laser radar BEV input to the Query matrix, Key matrix, and Value matrix of the radar; S53, calculate radar weight: calculate the radar attention weight matrix by multiplying the image query matrix and the radar key matrix, and multiply it by the radar value matrix to generate the radar weight; S54, calculate image weight: calculate the image attention weight matrix by multiplying the radar query matrix and the image key matrix, and multiply it with the image value matrix to generate the image weight; S55, updating BEV input: multiplying the radar weight by the laser radar BEV input, and multiplying the image weight by the machine vision BEV input, to obtain updated machine vision BEV input and laser radar BEV input; S56, iterative update: update the machine vision BEV input and the laser radar BEV input and return to step S52 until a predetermined number of iterations K is reached; S57, feature fusion and output: concatenate the updated lidar BEV input and the machine vision BEV input, and generate fusion features through the fully connected layer to obtain the final fusion BEV features; S6. Input the fused BEV features into the detection network to obtain three-dimensional target detection information.
2. The three-dimensional target detection method based on multi-view and point cloud BEV feature fusion according to claim 1 is characterized in that: The S3 comprises the following steps: S31, multi-view projection: projecting points in the non-empty voxel center point set onto multiple views of the multi-view image to obtain a multi-view projection voxel center point set; S32, central voxel feature extraction: Through the fully connected layer and the deformable Transformer network, the central point features of each view are extracted from the set of projected voxel centers of each view, and then the central point features of each view are spliced into the central voxel feature; S33, point cloud voxel feature acquisition: project the point cloud in the non-empty voxels onto the multi-view image, and use bilinear interpolation to obtain the point cloud voxel features; S34, feature fusion: The central voxel feature and the point cloud voxel feature are spliced along the feature channel, and the machine vision voxel feature is obtained through the fully connected layer.
3. The three-dimensional target detection method based on multi-view and point cloud BEV feature fusion according to claim 1 is characterized in that: The S51 comprises the following steps: Frame Machine Vision BEV Features With Frame LiDAR BEV Features Compressed to dimensions is No. Iteration 1 of machine vision BEV input and LiDAR BEV input for the first iteration .
4. The three-dimensional target detection method based on multi-view and point cloud BEV feature fusion according to claim 3 is characterized in that: The S52 comprises the following steps: Iteration 1 of machine vision BEV input Mapping to image Query matrix , image key matrix and the image Value matrix : ; ; ; In the formula, , , They are all learnable linear transformation matrices; The first LiDAR BEV input for the first iteration Mapping to radar query matrix , radar key matrix and radar value matrix ,Right now: ; ; ; In the formula, , , Both are learnable linear transformation matrices.
5. The three-dimensional target detection method based on multi-view and point cloud BEV feature fusion according to claim 4 is characterized in that: The S53 comprises the following steps: With radar Key matrix Multiply and make Processing, obtain radar attention weight matrix , and then the radar attention weight matrix and radar value matrix Multiply and pass through a fully connected network and a Network, generate Radar weights for iterations .
6. The three-dimensional target detection method based on multi-view and point cloud BEV feature fusion according to claim 5 is characterized in that: The S54 comprises the following steps: With the image Key matrix Multiply and make Processing, obtain the image attention weight matrix , and then the image attention weight matrix and the image Value matrix Multiply and pass through a fully connected network and a Network, generate The image weights for the iteration .
7. The three-dimensional target detection method based on multi-view and point cloud BEV feature fusion according to claim 6 is characterized in that: The S55 comprises the following steps: Radar weights for iterations With LiDAR BEV input for the first iteration Multiply them together to get LiDAR BEV input for the first iteration ; The image weights for the iteration With Iteration 1 of machine vision BEV input Multiply them together to get Iteration 1 of machine vision BEV input .
8. The three-dimensional target detection method based on multi-view and point cloud BEV feature fusion according to claim 7 is characterized in that: The S56 comprises the following steps: Assign to ,Will Assign to ,Will Assign to After that, return to step S52 and execute sequentially until = K, and get the LiDAR BEV input for the first iteration With Iteration 1 of machine vision BEV input .
Citation Information
Patent Citations
Target detection method based on laser radar and machine vision fusion
CN115032651A
3D target detection method, electronic equipment and storage medium
CN116246119A