Robust three-dimensional target detection method and system based on object-level feature fusion

Through the object-level feature fusion method, foreground key points and pseudo-key points are generated, combined with the internal and external hierarchical grid structure, the spatial information loss, heterogeneity and cross-modal misalignment problems of three-dimensional object detection in complex driving scenarios are solved, achieving more accurate and robust detection effects.

CN120298847AInactive Publication Date: 2025-07-11WUHAN UNIV OF TECH +1

Patent Information

Application Number
CN202510783779.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems in spatial information loss, background interference, heterogeneity of multimodal feature fusion and cross-modal misalignment in complex driving scenarios, resulting in insufficient accuracy and robustness of three-dimensional object detection.

Method used

The object-level feature fusion method is adopted to generate foreground key points and pseudo-key points, combined with the internal and external hierarchical grid structure, differentiated fusion of multimodal features is achieved, and the consistency of point clouds and image features in high-dimensional semantic space is utilized to avoid spatial information loss and cross-modal misalignment.

Benefits of technology

It realizes more accurate and robust three-dimensional object detection in complex driving scenarios, improves detection accuracy and stability, and overcomes the shortcomings of traditional methods in spatial information retention, feature heterogeneity utilization and cross-modal alignment robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298847A_ABST
    Figure CN120298847A_ABST
Patent Text Reader

Abstract

The invention discloses a robust three-dimensional target detection method and system based on object-level feature fusion. The method comprises the following steps: S1, extracting multi-scale image features from a visible light camera image; extracting multi-scale point cloud features from the laser radar point cloud, and generating a rough three-dimensional object candidate frame; s2, generating foreground key points and pseudo key points; s3, extracting original foreground key point features, voxel features and BEV features of foreground key points projected in the aerial view from the multi-scale point cloud features; projecting the pseudo key points to multi-scale image features, and extracting corresponding image features; s4, fusing the features to obtain multi-modal object-level features belonging to the same target instance in the laser radar point space; and S5, inputting the multi-modal object-level features of the same target instance into a classification and regression detection head, and generating a final target detection result. According to the method, challenges caused by weak alignment among multi-modal data are effectively reduced, and more robust and accurate three-dimensional target detection is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent driving vehicle environment perception, and particularly relates to a robust three-dimensional object detection method and system based on object-level feature fusion. Background Art

[0002] In complex real driving scenarios, a key and challenging environment perception task is three-dimensional object detection. In recent years, there has been a trend to fuse data of two modalities, namely, Light Detection and Ranging (LiDAR) and visible light cameras, to enhance the performance of 3D object detection. Existing multi-modal fusion 3D object detection methods mainly perform fusion at the feature level, and among them, point-by-point and Region of Interest (RoI)-wise methods are the most common ways to establish the correspondence between multi-modalities. However, either method will bring a series of complexities.

[0003] Due to the sparsity of point clouds, the point-by-point method using foreground LiDAR points as the cross-modal correspondence basis will waste a large amount of dense image features. To make full use of dense image features, an alternative point-by-point method is to generate pseudo points from image pixels through per-pixel depth estimation, so as to densify foreground points. Depth estimation is usually carried out by obtaining depth from adjacent LiDAR points or through a supervised learning process. However, depth estimation may introduce problems of inaccurate estimation accuracy and generate a large amount of computational overhead.

[0004] RoI-wise methods also attempt to utilize the complementarity of multi-modal features in the fusion process. However, traditional RoI-wise methods use RoIs as the basis for sampling features across different modalities, which will cause loss of spatial information when projecting point clouds into the Bird's Eye View (BEV) space to obtain RoI features, and will be interfered by background pixels when obtaining rectangular RoI features in the Frontal View (FV) space of images. These two aspects jointly lead to the generation of blurred object-level features. To make full use of the spatial features of point clouds and obtain accurate foreground image features, a class of methods attempts to unify the original point clouds and image features into a homogeneous representation, usually voxel features, and then uses 3D RoIs to extract corresponding voxel features for multi-modal feature fusion. However, this RoI-wise method fuses multi-modal features indiscriminately, ignoring the inherent significant heterogeneity of multi-modal features. In addition, similar to the point-by-point method, the RoI-wise method is also vulnerable to the challenges brought by weak alignment.

[0005] It can be seen that the existing technologies mainly have the following disadvantages: (1)Blurred object-level feature problems caused by spatial information loss and background interference: Existing RoI-level fusion methods suffer from spatial detail loss when projecting point clouds into the bird's-eye view (BEV) space to obtain features, and are vulnerable to background pixel interference when using rectangular RoIs to extract image features in the front view (FV) space of images.

[0006] (2)Difficulties in distinguishing heterogeneity in multi-modal feature fusion: Traditional RoI-level fusion methods treat different modal features without discrimination and cannot effectively highlight the respective advantages of LiDAR and image features in localization and classification.

[0007] (3)Cross-modal misalignment (weak alignment) problem: In the case of inaccurate correspondence in the low-dimensional space, existing methods are difficult to achieve robust multi-modal feature fusion. When performing feature densification using point cloud projection or per-pixel depth estimation, it is vulnerable to alignment errors and inaccurate depth estimation, and the computational cost is high, which limits the fusion efficiency and accuracy.

[0008] Therefore, how to utilize the structural characteristics and distribution laws of multi-source sensor data to construct a fusion mechanism that can efficiently fuse multi-modal features is of great significance for achieving accurate and robust 3D object detection in complex driving environments. Summary of the Invention

[0009] The present invention aims to propose a robust 3D object detection method and system based on object-level feature fusion. Through an innovative object-level feature fusion mechanism, a breakthrough is achieved in the feature extraction and fusion framework, overcoming the deficiencies of the existing technology in terms of spatial information retention, feature heterogeneity utilization, and cross-modal alignment robustness, thereby achieving more accurate and robust 3D object detection in complex driving scenarios.

[0010] The technical solution adopted by the present invention is as follows: Provide a robust 3D object detection method based on object-level feature fusion, including the following steps: S1. Extract multi-scale image features from visible light camera images; extract multi-scale point cloud features from lidar point clouds, generate a bird's-eye view feature map based on the multi-scale point cloud features, and then generate rough 3D object candidate boxes based on the bird's-eye view feature map; S2. Generate foreground key points for the lidar point cloud through a hybrid sampling and binary classification supervision strategy, and generate pseudo key points according to the center offset mechanism; S3. Extract the original foreground key point features, voxel features, and BEV features projected by the foreground key points in the bird's-eye view from the multi-scale point cloud features to obtain a lidar point cloud feature set in the lidar point cloud space, project the pseudo key points onto the multi-scale image features, and extract the corresponding image features to obtain an image feature set in the lidar point cloud space; S4. Create an inner and outer hierarchical meshed structure of a certain target object using the rough 3D object candidate box. The meshed structure includes inner vertices and outer vertices. For the outer vertices, aggregate lidar point cloud features from the foreground key points through neighborhood search. For the inner vertices, aggregate image features from the pseudo key points using the same method. Extract the context information of each grid vertex, generate and splice all grid vertex features to obtain the multi-modal object-level features belonging to the same target object in the fused lidar point space. S5. Input the multi-modal object-level features belonging to the same target object into the target classification detection head and the bounding box regression detection head respectively to generate the detection results of the target object, including the object category and the three-dimensional spatial position, size, and orientation.

[0011] Continuing with the above technical solution, the process of generating foreground key points in step S2 is specifically as follows: Combining the spatial distance and the semantic distance, use the farthest point sampling method for hybrid key point sampling. First, select 4N points from the input point cloud using the farthest point sampling of the three-dimensional Euclidean distance to achieve uniform coverage of the space. Second, sample N / 2 points using the feature farthest point sampling method based on the semantic distance. At the same time, sample N / 2 points again using the farthest point sampling based on the three-dimensional Euclidean distance. Finally, obtain a set of N original key points and the features corresponding to the key points. Introduce a binary classification supervision mechanism to calculate the foreground probability score for the original key point features, suppress the background point features, and obtain the final foreground key points.

[0012] Continuing with the above technical solution, the process of generating pseudo key points in step S2 is specifically as follows: Using the three-dimensional coordinates of the original key points as the input, obtain the predicted three-dimensional displacement through a multi-layer perceptron (MLP). If the original key point is a foreground key point, the predicted three-dimensional displacement will be supervised by the regression loss towards the center to ensure that the point moves towards the geometric center of the target object, obtaining the final pseudo key points.

[0013] Continuing with the above technical solution, the specific process of the lidar point cloud feature set in step S3 is as follows: Use the neighborhood search method to obtain the voxel features centered on the foreground key points, generating a voxel feature set containing multiple spatial scales. Project the foreground key points onto the bird's-eye view and extract the features from the bird's-eye view perspective at the foreground key point positions using bilinear interpolation. Aggregate the original foreground key point features, voxel features, and features from the bird's-eye view perspective in the multi-scale point cloud features to obtain the lidar point cloud feature set in the lidar point space.

[0014] Continuing with the above technical solution, the specific process of the image feature set in step S3 is as follows: Project the pseudo key points onto the multi-scale image features, use the bilinear interpolation method to obtain the features of adjacent pixels, and aggregate the multi-scale image features of the sampling points to obtain the image feature set.

[0015] Continuing with the above technical solution, step S4 is specifically as follows: Use the rough 3D object candidate box to create a mesh of the shape of a certain target object, where is the number of divisions along each dimension. This mesh includes internal vertices and external vertices; for the external vertices, through neighborhood search, aggregate the lidar point cloud features from the foreground key points; for the internal vertices, use the same method to aggregate the image features from the pseudo key points; extract the context information of each mesh vertex, generate all mesh vertex features and splice them to form an internal and external hierarchical grid structure to fuse the multi-modal object-level features belonging to the same target object in the lidar point space.

[0016] Continuing with the above technical solution, the detection head in step S5 includes two parallel branches: one is the object classification branch, which is used to evaluate the confidence of the candidate box and judge the category of the target; the other is the bounding box regression branch, which is used to accurately predict the geometric parameters of the candidate box, including length, width, height, center position and orientation.

[0017] Continuing with the above technical solution, step S3 further includes the step of: splicing the two feature sets along the feature dimension and multiplying by the foreground probability score to suppress the background features.

[0018] The present invention also provides a robust 3D object detection system based on object-level feature fusion, including: A feature extraction and object candidate box generation module, which is used to extract multi-scale image features from visible light camera images; extract multi-scale point cloud features from lidar point clouds, generate a bird's-eye view feature map based on the multi-scale point cloud features, and then generate a rough 3D object candidate box based on the bird's-eye view feature map; A foreground key point and pseudo key point generation module, which is used to generate foreground key points for the lidar point cloud image through a hybrid sampling and binary classification supervision strategy, and generate pseudo key points according to the center offset mechanism; A feature unified representation module, which is used to extract the original foreground key point features, voxel features and BEV features projected by the foreground key points in the bird's-eye view from the multi-scale point cloud features according to the foreground key points, obtain the lidar point cloud feature set in the lidar point cloud space, project the pseudo key points onto the multi-scale image features, and extract the corresponding image features to obtain the image feature set in the lidar point cloud space; The cross-modal object-level feature fusion module is used to create a hierarchical grid structure inside and outside a target object by using rough 3D object candidate boxes. The grid structure includes inner vertices and outer vertices. For the outer vertices, lidar point cloud features are aggregated from foreground key points through neighborhood search. For the inner vertices, image features are aggregated from pseudo key points in the same way. The context information of each grid vertex is extracted, and all grid vertex features are generated and concatenated to obtain multi-modal object-level features belonging to the same target object in the fused lidar point space. The target object detection module is used to input the multi-modal object-level features belonging to the same target object into the target classification detection head and the bounding box regression detection head respectively, and generate the detection results of the target object, including the object category, 3D spatial position, size, and orientation.

[0019] The present invention also provides a computer storage medium, which stores a computer program executable by a processor. The computer program executes the robust 3D target detection method based on object-level feature fusion described in any one of the above technical solutions.

[0020] The present invention also provides a vehicle-mounted robust 3D target detection system based on object-level feature fusion, including a data collector, a vehicle-mounted storage and computing platform, and a vehicle controller and a vehicle actuator. The data collector includes a lidar, a visible light camera, and a vehicle data sensor. The vehicle-mounted storage and computing platform is provided with the above computer storage medium. The vehicle controller controls the instructions output by the vehicle-mounted storage and computing platform, and the vehicle actuator performs corresponding actions according to the control instructions.

[0021] The beneficial effects of the present invention are as follows: The core idea of the present invention is to unify the features from lidar and images into a homogeneous representation in the point cloud space, and achieve multi-modal feature aggregation at the object level through a carefully designed hierarchical grid structure with different fusion strategies inside and outside. Specifically, the present invention first generates original foreground key points and pseudo key points by using a point enhancement and center offset mechanism, and associates the foreground point cloud features and image features point by point to avoid the loss of spatial information and feature blur caused by projection and background interference. Subsequently, a cross-modal object-level feature fusion mechanism is proposed, which uses a hierarchical grid structure inside and outside to fuse multi-modal features, and solves the problem that it is difficult for traditional methods to distinguish multi-modal heterogeneity. In addition, the present invention strengthens the consistency of object-level features at the high-dimensional semantic space level, and can still maintain the robustness of detection performance even in the presence of cross-modal misalignment. Through the innovative object-level feature fusion mechanism, the present invention has achieved a breakthrough in the feature extraction and fusion framework, overcome the deficiencies of the prior art in terms of spatial information retention, feature heterogeneity utilization, and cross-modal alignment robustness, and thus realized more accurate and robust 3D target detection in complex driving scenarios.

[0022] Of course, it is not necessary for any product implementing the present invention to achieve all of the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will further introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0024] Figure 1 is a flowchart of a robust 3D object detection method based on object-level feature fusion according to an embodiment of the present invention; Figure 2 is a flowchart of a robust 3D object detection method based on object-level feature fusion according to another embodiment of the present invention; Figure 3 is a specific processing schematic diagram of a robust 3D object detection method based on object-level feature fusion according to another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0026] It should be noted that the illustrations provided in the embodiments of the present invention only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape and size of the components in actual implementation. The type, quantity and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0027] In the present invention, it should also be noted that when terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. appear, the orientation or position relationship indicated is based on the orientation or position relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present application. In addition, when terms such as "first" and "second" appear, they are only used for descriptive and differentiating purposes and cannot be understood as indicating or implying relative importance.

[0028] In addition, it should be noted that the features of various embodiments of the present invention can be partially or wholly combined or integrated, and as can be understood by those skilled in the art, they can interact and operate in different ways. Each embodiment can be implemented independently of each other or in an associated relationship.

[0029] As Figure 1 shown, the robust three-dimensional object detection method based on object-level feature fusion according to the embodiment of the present invention includes the following steps: S1. Extract multi-scale image features from visible light camera images; extract multi-scale point cloud features from lidar point clouds, generate a bird's-eye view feature map based on the multi-scale point cloud features, and then generate a rough three-dimensional object candidate box based on the bird's-eye view feature map; S2. Generate foreground key points for the lidar point cloud through a hybrid sampling and binary classification supervision strategy, and generate pseudo key points according to the center offset mechanism; S3. Extract the original foreground key point features, voxel features, and BEV features projected by the foreground key points in the bird's-eye view from the multi-scale point cloud features to obtain a lidar point cloud feature set in the lidar point cloud space, and project the pseudo key points onto the multi-scale image features to extract the corresponding image features to obtain an image feature set in the lidar point cloud space; S4. Use the rough three-dimensional object candidate box to create an inner and outer layered grid structure for a certain target object, and the grid structure includes inner vertices and outer vertices; for the outer vertices, aggregate the lidar point cloud features from the foreground key points through neighborhood search; for the inner vertices, use the same method to aggregate the image features from the pseudo key points; extract the context information of each grid vertex, generate and splice all grid vertex features to obtain multi-modal object-level features belonging to the same target object in the fused lidar point space; S5. Input the multi-modal object-level features belonging to the same target object into the target classification detection head and the bounding box regression detection head respectively to generate the detection results of the target object, including the object category, three-dimensional spatial position, size, and orientation.

[0030] In step S1, the features of the last scale in the multi-scale point cloud features are pooled along the z-axis to generate a tensor with a fixed spatial size. This tensor not only retains rich feature information but also reflects the spatial layout under the bird's-eye view, so it is called the "bird's-eye view feature map". Based on this feature map, a rough three-dimensional object candidate box is generated by predicting the displacement between the predetermined anchor points and the true bounding boxes.

[0031] Specifically, the process of generating foreground key points in step S2 is as follows: Combining spatial distance and semantic distance, the farthest point sampling method is used for hybrid key point sampling. First, 4N points are selected from the input point cloud using the farthest point sampling of three-dimensional Euclidean distance to achieve uniform coverage of the space. Second, N / 2 points are sampled using the feature farthest point sampling method based on semantic distance; at the same time, N / 2 points are sampled again using the farthest point sampling based on three-dimensional Euclidean distance, and finally a set of N original key points and the features corresponding to the key points are obtained; a binary classification supervision mechanism is introduced into the original key point features to calculate the foreground probability score, suppressing the background point features, and obtaining the final foreground key points.

[0032] Specifically, the process of generating pseudo key points in step S2 is as follows: Taking the three-dimensional coordinates of the original key points as input, the predicted three-dimensional displacement is obtained through a multi-layer perceptron MLP; if the original key point is a foreground key point, the predicted three-dimensional displacement will be supervised by the regression loss towards the center to ensure that the point moves towards the geometric center of the target object, obtaining the final pseudo key points.

[0033] Furthermore, the specific process of the lidar point cloud feature set in step S3 is as follows: The neighborhood search method is used to obtain the voxel features centered on the foreground key points, generating a voxel feature set containing multiple spatial scales; the foreground key points are projected onto the bird's-eye view, and bilinear interpolation is used to extract the features from the bird's-eye view at the positions of the foreground key points; the original foreground key point features, voxel features, and features from the bird's-eye view in the multi-scale point cloud features are aggregated to obtain the lidar point cloud feature set in the lidar point space.

[0034] The specific process of the image feature set in step S3 is as follows: The pseudo key points are projected onto the multi-scale image features, and the bilinear interpolation method is used to obtain the features of adjacent pixels, and the multi-scale image features of the sampled points are aggregated to obtain the image feature set.

[0035] Step S3 also includes the step of: concatenating the two feature sets along the feature dimension and multiplying by the foreground probability score to suppress the background features.

[0036] Step S4 is specifically: Using the rough three-dimensional object candidate box to create a mesh of the shape of a certain target object, where is the number of divisions along each dimension, and this mesh includes internal vertices and external vertices; for the external vertices, through neighborhood search, the lidar point cloud features are aggregated from the foreground key points; for the internal vertices, the same method is used to aggregate the image features from the pseudo key points; the context information of each mesh vertex is extracted, and all mesh vertex features are generated and concatenated to form an internal and external hierarchical grid structure to fuse the multi-modal object-level features belonging to the same target object in the lidar point space.

[0037] Among them, the target classification detection head includes two parallel branches: one is the target classification branch, which is used to evaluate the confidence of the candidate box and judge the category of the target; the other is the bounding box regression branch, which is used to accurately predict the geometric parameters of the candidate box, including length, width, height, center position and orientation.

[0038] It can be seen that this embodiment can construct object-level features without ambiguity: The present invention no longer strictly depends on the precise spatial alignment requirements of point cloud and image features, but realizes the precise sampling of multi-modal features by utilizing the consistency of object-level features in the high-dimensional semantic space. The present invention effectively avoids the problems of spatial information loss and background pixel interference brought by traditional RoI mapping, thereby constructing a clear and unambiguous object-level feature representation.

[0039] Secondly, this embodiment is based on a differential feature fusion mechanism with an internal and external hierarchical architecture: Aiming at the problems of heterogeneous mixing and blurring that are prone to occur in multi-modal feature fusion, the present invention adopts an internal and external hierarchical grid architecture to strictly distinguish and orderly fuse point cloud and image features in the outer layer and the inner layer respectively. This invention ensures that the accurate spatial information from the peripheral LiDAR features and the rich semantics from the foreground image features are combined in an accurate manner, thereby realizing the efficient fusion of multi-modal features at the object level and greatly improving the accuracy and stability of 3D object detection.

[0040] Finally, this embodiment weakens the low-dimensional alignment requirements and strengthens the robustness of high-dimensional semantic drive: Aiming at practical challenges such as cross-modal misalignment and inaccurate depth estimation, the present invention weakens the dependence on low-dimensional spatial correspondence relationships from the high-dimensional semantic level. By making full use of the high-dimensional semantic consistency at the object level, the present invention effectively resists the adverse effects brought by alignment errors and depth deviations, and realizes more robust object detection.

[0041] Another embodiment of the present invention is a robust three-dimensional object detection method based on object-level feature fusion. As shown in FIGS. 2 and 3, the method includes the following steps: Step 1. The Swin-Transformer can be used to extract multi-scale image features in the visible light camera image, and 3D sparse convolution can be used to extract multi-scale point cloud features in the lidar point cloud image, and then a rough three-dimensional object candidate box is generated through the RPN network as the carrier for subsequent object-level feature fusion. The specific steps are as follows: Step 11. Image feature sampling: Add the image features generated after block segmentation processing to the position encoding with the same shape to form a new feature vector; then gradually expand the receptive field and encode information based on the new feature vector through a network composed of a block merging module and a hierarchical structure Transformer module; Step 12. Point cloud feature sampling: Divide the unordered lidar point cloud into regular voxels, and use PointNet to encode the lidar point information in each voxel. Subsequently, in order to reduce the computational complexity, a feature extraction network containing 3D sparse convolution is used to convert the encoded voxel features into high-dimensional feature vectors. The feature extraction network can be composed of multiple 3D convolution stages, and each stage gradually extracts and aggregates higher-level features through a series of convolution operations; the convolution operations can capture the local geometric features of the point cloud and obtain more extensive context information through multi-scale feature aggregation; Step 13. Coarse 3D object candidate box generation: Extract features from the sparse 3D voxel features generated in the last stage of the feature extraction network in Step 12 from the bird's-eye view perspective. Subsequently, use the RPN network, and use the features from the bird's-eye view perspective as input to predict the displacement between the predetermined anchor and the ground truth bounding box, and generate rough 3D object candidate boxes.

[0042] Step 2. Perform point enhancement. Specifically, use hybrid sampling and binary classification supervision to enhance a set of foreground key points as the medium for aggregating lidar point cloud features, and then use the center offset mechanism to obtain a set of pseudo key points from these enhanced key points as the medium for aggregating image features; the specific steps are as follows: Step 21: To enhance the foreground features of the key points, the present invention adopts a hybrid sampling strategy that combines spatial distance and semantic distance. First, use farthest point sampling based on the three-dimensional Euclidean distance to select 4N points from the input point cloud to achieve uniform coverage of the space. Subsequently, extract semantic features through PointNet++, and select N / 2 points based on farthest point sampling of semantic distance features, and at the same time use farthest point sampling based on the three-dimensional Euclidean distance to sample N / 2 points again. Finally, merge the two sampling results to generate N key points and their feature representations .

[0043] Step 22: To suppress the background point features, the present invention introduces an auxiliary semantic supervision mechanism to process the key point features . The specific steps are as follows: First, reduce the number of channels to 256, and then further reduce it to 1 channel through two layers of multi-layer perceptrons (MLPs). Subsequently, apply the Sigmoid function to calculate the probability of each point being a foreground point, and its mathematical expression is as follows:

[0044] This score will be used to adjust the subsequent aggregation features element-wise, and use the binary cross-entropy loss function to calculate the loss.

[0045] Step 23: The center offset mechanism designed in the present invention uses the three-dimensional coordinates of the original key points as input, and obtains a three-dimensional displacement through the MLP layer ; if the key point is a foreground key point, the predicted will be supervised by the regression loss to the center to ensure that the point moves towards the geometric center of the target object. The designed key point offset loss function is as follows:

[0046] In the formula, represents the number of foreground points among the key points. In the training stage, if a key point is within the object bounding box, it is regarded as a foreground point. represents the true offset of the foreground point to the corresponding object center; Step 3. Propose a feature unified representation module, which uses the foreground key points and pseudo key points generated by the point enhancement module to unify the point cloud features and image features into the same homogeneous representation in the lidar point space. The specific steps are as follows: Step 31: On the basis of retaining the original key point features with rich three-dimensional spatial details, adopt the neighborhood search method to obtain a multi-scale voxel feature set centered on the original key points . Then, project the original key points onto the BEV plane and obtain the BEV features through bilinear interpolation. Finally, integrate the original key point features, voxel features and BEV features to form a multi-scale lidar point cloud feature set .

[0047] Step 32: Project the pseudo key points onto the four-scale image features of the corresponding visible light camera image. Considering that the projected pixels may not be exactly aligned with the integer coordinates on the image, use the bilinear interpolation method to obtain the features of adjacent pixels. Then, integrate the multi-scale image features of the sampled points to generate an image feature set .

[0048] Step 33: Concatenate the lidar feature set and the image feature set along the feature dimension, and multiply by the foreground probability score to suppress the background features. Then, fuse the features of different scales and characteristics through the MLP, and its mathematical expression is as follows:

[0049]

[0050] where represents the number of key points, Represents the number of feature channels for aggregated features.

[0051] Step 4. Propose a cross-modal object-level feature fusion module, and construct an internal and external hierarchical grid structure to fuse multi-modal object-level features belonging to the same target instance in the lidar point space, so as to generate unambiguous object-level fusion features. The specific steps are as follows: Step 41: Create a -shaped grid using the rough 3D object candidate box, where is the number of divisions along each dimension. This grid includes internal vertices 3 内部 with the shape of N and external vertices 3 -N 3 内部 with the shape of N .

[0052] Step 42: For the external vertex , aggregate the lidar point cloud features from the foreground key points through neighborhood search; for the internal vertex , aggregate the image features from the pseudo key points using the same method.

[0053] Step 43: Use PointNet to extract the context information of each grid vertex, generating 128-dimensional grid features and . Subsequently, concatenate the features of all grid vertices and input them into a two-layer MLP to obtain a 128-dimensional object-level feature representation, and its mathematical expression is as follows: .

[0054] Step 5. Input the multi-modal object-level features belonging to the same target instance into the target classification detection head and the bounding box regression detection head respectively, so as to generate accurate and robust 3D object detection results, including the object category, its length, width, height, center position and orientation.

[0055] Specifically: Input the object-level features of each target instance into the detection head, which is the core module of the candidate box refinement network, and refine the rough object candidate box. The detection head contains two parallel branches: one is the target classification branch, which is used to evaluate the confidence of the candidate box and judge the category of the target; the other is the bounding box regression branch, which is used to accurately predict the geometric parameters of the candidate box, including length, width, height, center position and orientation. The two branches work together to achieve fine correction of the candidate box, so as to generate accurate and robust 3D object detection results.

[0056] The present invention also provides a robust 3D object detection system based on object-level feature fusion, including: A feature extraction and object candidate box generation module, configured to extract multi-scale image features from visible light camera images; extract multi-scale point cloud features from lidar point cloud images, and generate rough 3D object candidate boxes based on the bird's-eye view feature map; A foreground key point and pseudo key point generation module, configured to generate foreground key points for the lidar point cloud image through a hybrid sampling and binary classification supervision strategy, and generate pseudo key points according to the center offset mechanism; A feature unified representation module, configured to extract the original foreground key point features, voxel features, and BEV features projected by the foreground key points in the bird's-eye view from the multi-scale point cloud features according to the foreground key points, to obtain a lidar point cloud feature set in the lidar point cloud space, project the pseudo key points onto the multi-scale image features, and extract the corresponding image features, to obtain an image feature set in the lidar point cloud space; A cross-modal object-level feature fusion module, configured to create an inner and outer hierarchical grid structure of a certain target object by using the rough 3D object candidate box, where the grid structure includes inner vertices and outer vertices; for the external vertices, aggregate the lidar point cloud features from the foreground key points through neighborhood search; for the internal vertices, aggregate the image features from the pseudo key points in the same way; extract the context information of each grid vertex, generate and splice all grid vertex features, to obtain the multi-modal object-level features belonging to the same target object in the fused lidar point space; An object detection module, configured to input the multi-modal object-level features belonging to the same target object into a target classification detection head and a bounding box regression detection head respectively, and generate the detection result of the target object, including the object category, 3D spatial position, size, and orientation.

[0057] Each module is mainly used to implement the steps of the above method embodiments, which will not be elaborated here one by one.

[0058] The present application also provides a computer-readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disc, a server, an App application mall, etc., on which a computer program is stored, and when the program is executed by a processor, the corresponding functions are implemented. When the computer-readable storage medium of this embodiment is executed by a processor, the robust 3D object detection method based on object-level feature fusion in the method embodiment is implemented.

[0059] Based on the robust 3D object detection method based on object-level feature fusion in the above embodiments, the present invention further constructs a vehicle-mounted robust 3D object detection system based on object-level feature fusion, including sensors (for data acquisition, including lidar, visible light cameras, vehicle data sensors, etc.) and a vehicle-mounted storage and computing platform (memory, positioning and perception computing platform, decision-making and planning computing platform), etc. Among them, the sensors communicate with the vehicle-mounted storage and computing platform through data transmission interfaces (vehicle-mounted Ethernet, USB, CAN). The execution process of this system is as follows: (1) Convert the robust 3D object detection algorithm based on object-level feature fusion proposed by the present invention into instruction codes and deploy them in the memory of the vehicle-mounted computing platform.

[0060] (2) Configure the drivers of the lidar and visible light camera sensors to achieve the parsing and forwarding of sensor data, and the form of the forwarded data matches the instruction codes in (1).

[0061] (3) Based on the instruction codes in (1), perform computational processing on the parsed and forwarded data on the positioning and perception computing platform, obtain the detection results and send them to the memory. The decision-making and planning computing platform reads the real-time detection results from the memory and completes downstream tasks on the decision-making and planning computing platform according to the positioning and perception results obtained by other algorithms.

[0062] (4) The vehicle controller outputs corresponding control instructions to the vehicle actuator according to the downstream tasks for action execution.

[0063] It should be noted that according to the needs of implementation, each step / component described in this application can be split into more steps / components, or two or more steps / components or partial operations of steps / components can be combined into new steps / components to achieve the purpose of the present invention.

[0064] The sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0065] It should be understood that those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A robust three-dimensional object detection method based on object-level feature fusion, characterized in that, It includes the following steps: S1. Extract multi-scale image features from visible light camera images; Extract multi-scale point cloud features from lidar point clouds, generate a bird's-eye view feature map based on the multi-scale point cloud features, and then generate rough 3D object candidate boxes based on the bird's-eye view feature map; S2. Generate foreground key points for the lidar point cloud through a hybrid sampling and binary classification supervision strategy, and generate pseudo key points according to the center offset mechanism; S3. Extract the original foreground key point features, voxel features, and BEV features projected by the foreground key points in the bird's-eye view from the multi-scale point cloud features to obtain a lidar point cloud feature set in the lidar point cloud space, and project the pseudo key points onto the multi-scale image features to extract the corresponding image features to obtain an image feature set in the lidar point cloud space; S4. Use the rough 3D object candidate box to create an inner and outer layered grid structure for a certain target object. The grid structure includes inner vertices and outer vertices; for the outer vertices, aggregate the lidar point cloud features from the foreground key points through neighborhood search; For the inner vertices, use the same method to aggregate image features from the pseudo key points; extract the context information of each grid vertex, generate and splice all grid vertex features to obtain the multi-modal object-level features belonging to the same target object in the fused lidar point space; S5. Input the multi-modal object-level features belonging to the same target object into the target classification detection head and the bounding box regression detection head respectively to generate the detection results of the target object, including the object category and the three-dimensional spatial position, size, and orientation.

2. The robust three-dimensional object detection method based on object-level feature fusion according to claim 1, characterized in that The process of generating foreground key points in step S2 is specifically as follows: Combining the spatial distance and the semantic distance, use the farthest point sampling method for hybrid key point sampling; First, select 4N points from the input point cloud using the farthest point sampling of the three-dimensional Euclidean distance to achieve uniform coverage of the space; Second, use the feature farthest point sampling method based on the semantic distance to sample N / 2 points; At the same time, use the farthest point sampling method of the three-dimensional Euclidean distance again to sample N / 2 points, and finally obtain a set of N original key points and the features corresponding to the key points; Introduce a binary classification supervision mechanism for the original key point features to calculate the foreground probability score and suppress the background point features to obtain the final foreground key points.

3. The robust 3D object detection method based on object-level feature fusion according to claim 2, wherein The process of generating pseudo key points in step S2 is specifically as follows: Taking the three-dimensional coordinates of the original key points as the input, obtain the predicted three-dimensional displacement through a multi-layer perceptron MLP; if the original key point is a foreground key point, the predicted three-dimensional displacement will be supervised by the regression loss to the center to ensure that the point moves towards the geometric center of the target object to obtain the final pseudo key points.

4. The robust three-dimensional object detection method based on object-level feature fusion according to claim 1, wherein, The specific process of obtaining the lidar point cloud feature set in step S3 is as follows: The voxel features centered on the foreground key points are obtained by using the neighborhood search method, generating a voxel feature set containing multiple spatial scales; the foreground key points are projected onto the bird's-eye view, and the features from the bird's-eye view are extracted at the positions of the foreground key points using bilinear interpolation; the original foreground key point features, voxel features, and features from the bird's-eye view in the multi-scale point cloud features are aggregated to obtain the lidar point cloud feature set in the lidar point space.

5. The robust three-dimensional object detection method based on object-level feature fusion according to claim 1, characterized in that The specific process of obtaining the image feature set in step S3 is as follows: The pseudo key points are projected onto the multi-scale image features, and the bilinear interpolation method is used to obtain the features of adjacent pixels, and the multi-scale image features of the sampled points are aggregated to obtain the image feature set.

6. The robust 3D object detection method based on object-level feature fusion according to claim 1, characterized in that Step S4 specifically includes: creating a mesh of a certain shape of a target object by using a rough three-dimensional object candidate box, where is the number of divisions along each dimension, and the mesh includes internal vertices and external vertices; for the external vertices, aggregate lidar point cloud features from foreground key points through neighborhood search; ​ For the internal vertices, the image features are aggregated from the pseudo key points using the same method; the context information of each grid vertex is extracted, and all grid vertex features are generated and concatenated to form an inner-outer hierarchical grid structure to fuse the multi-modal object-level features of the same target object in the lidar point space.

7. The robust three-dimensional object detection method based on object-level feature fusion according to claim 1, characterized in that The detection head contains two parallel branches: one is the object classification branch, which is used to evaluate the confidence of the candidate boxes and determine the category of the object; the other is the bounding box regression branch, which is used to accurately predict the geometric parameters of the candidate boxes, including length, width, height, center position, and orientation.

8. The robust three-dimensional object detection method based on object-level feature fusion according to claim 1, wherein Step S3 further includes the step of concatenating the two feature sets along the feature dimension and multiplying by the foreground probability score to suppress the background features.

9. A robust three-dimensional object detection system based on object-level feature fusion, characterized in that, Including: The feature extraction and object candidate box generation module is used to extract multi-scale image features from the visible light camera image; Extract multi-scale point cloud features from the lidar point cloud, generate a bird's-eye view feature map based on the multi-scale point cloud features, and then generate rough three-dimensional object candidate boxes based on the bird's-eye view feature map; The foreground key point and pseudo key point generation module is used to generate foreground key points for the lidar point cloud through a hybrid sampling and binary classification supervision strategy, and generate pseudo key points according to the center offset mechanism; The feature unified representation module is used to extract the original foreground key point features, voxel features, and BEV features projected by the foreground key points in the bird's-eye view from the multi-scale point cloud features according to the foreground key points to obtain the lidar point cloud feature set in the lidar point cloud space, and project the pseudo key points onto the multi-scale image features to extract the corresponding image features to obtain the image feature set in the lidar point cloud space; The cross-modal object-level feature fusion module is used to create an inner-outer hierarchical grid structure of a certain target object using the rough three-dimensional object candidate boxes, and the grid structure includes inner vertices and outer vertices; for the outer vertices, the lidar point cloud features are aggregated from the foreground key points through neighborhood search; For the internal vertices, the image features are aggregated from the pseudo key points using the same method; the context information of each grid vertex is extracted, and all grid vertex features are generated and concatenated to obtain the multi-modal object-level features of the same target object in the fused lidar point space; The target object detection module is used to input multi-modal object-level features belonging to the same target object into the target classification detection head and the bounding box regression detection head respectively, and generate the detection results of the target object, including the object category, three-dimensional spatial position, size, and orientation.

10. A computer storage medium, characterized in that, It stores a computer program executable by a processor, and the computer program executes the robust three-dimensional target detection method according to any one of claims 1-8 based on object-level feature fusion.

Citation Information

Patent Citations

  • Three-dimensional point cloud target detection method fused with two-dimensional image semantics

    CN116597264A

Cited By

  • 3D target detection method, system, device and medium

    CN120599598A

  • Laser radar position identification method based on point cloud structure characteristics

    CN122049054A