3D target detection method and device for adaptively optimizing BEV features, and computer equipment

Through adaptive optimization of BEV features, the Transformer decoder layer is used for feature optimization, which solves the problems of sparsity and uneven distribution of BEV features in the prior art, and significantly improves the accuracy and efficiency of 3D object detection.

CN120047935AActive Publication Date: 2025-05-27HANGZHOU DIANZI UNIV

Patent Information

Application Number
CN202411950715.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-27
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

The existing 3D object detection method based on camera faces the problems of BEV feature sparsity and uneven distribution density of cone point clouds caused by low confidence in depth estimation, resulting in low detection accuracy.

Method used

Adaptively optimized BEV features are adopted, and BEV features are generated through feature extraction backbone network and view angle conversion module, and feature optimization is used for 6-layer Transformer decoder layer. By uniformly setting the query box to query and fusion in three-dimensional space, adaptive fusion features are obtained.

Benefits of technology

It significantly improves the accuracy of 3D object detection of circumferential images, alleviates the problems of feature sparsity and uneven information distribution, improves the representativeness of the data, and realizes efficient 3D object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047935A_ABST
    Figure CN120047935A_ABST
Patent Text Reader

Abstract

The invention relates to a 3D target detection method and device for adaptively optimizing BEV features, computer equipment and a storable medium. The 3D target detection method for adaptively optimizing the BEV features comprises the following steps: acquiring all-round view image data from a Nuscenes data set; performing feature extraction on the look-around image by using a feature extraction backbone network to obtain image features; raising the dimension of the image features to a BEV space, and compressing a z-axis to obtain BEV features; n 3D query boxes to be queried are uniformly arranged in a BEV space, candidate boxes are projected to a BEV plane to obtain a 2D rotating frame, BEV features in the 2D rotating frame are extracted as prior features, the center points of the query boxes are used as query points, camera parameters are utilized to project the query points to an image feature plane for interpolation to obtain corresponding query features, and the prior features and the query features are fused on a channel level; the method comprises the following steps of: predicting a perd box and a label score by using fusion features, manually generating optimization points at the centers of six surfaces of k pred boxes before the score, adding offset, optimizing and updating BEV features by using the corresponding fusion features, sending the optimized BEV features into a positioning detection head and a classification detection head, and respectively predicting a 3D object bounding box and an object category. The method has the advantage that the accuracy of 3D target detection of the look-around image is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision object detection, and particularly to a 3D object detection method, device, computer device and readable storage medium for adaptively optimizing BEV features. Background Art

[0002] As a key component in 3D perception, 3D object detection can be widely applied in fields such as autonomous driving and robotics. Although many LiDAR-based 3D detection methods have proven to have significant performance, camera-based methods have received increasing attention in recent years. The reasons for this shift are not only the lower deployment cost of cameras, but also the advantages of cameras in long-distance detection and identifying visual road elements. However, different from LiDAR sensors that directly provide accurate depth information, detecting objects relying solely on camera sensor images faces huge challenges. Therefore, how to construct effective BEV features using multi-view images has become a key issue.

[0003] Currently, the mainstream BEV feature construction methods are divided into two branches: one is the method based on Lift-Splat-Shoot, which elevates the 2D surround-view image features to 3D frustum point cloud features through explicit depth estimation and coordinate system transformation, and finally obtains the BEV features through the BEVPool operation. However, due to the characteristics of depth estimation, the feature information obtained by BEV grids with low depth probability confidence is invalid, which will cause the sparsity of BEV features; at the same time, due to the characteristics of camera imaging, the distribution density of the constructed frustum point cloud in the BEV space is uneven, showing a trend of being dense near and sparse far, resulting in uneven distribution of the constructed BEV feature information. The other branch is the method based on Transformer. Such methods do not rely on depth estimation, but interact with 2D surround-view image features through querying to complete the end-to-end 3D object detection task. At the same time, according to the density of the queries, it can be further divided into dense queries and sparse queries, and the typical representatives are BEVFormer and PETR respectively. The former constructs dense BEV grid features in advance, performs a large number of interactions between the queries and the 2D surround-view image features, and finally completes the construction of BEV features. However, due to the dense queries, this method has a slow inference speed and high computational cost. The latter uses sparse queries to lighten the entire network structure, and at the same time makes the query mechanism pay more attention to foreground object information. However, since the process of constructing BEV features is omitted, it will lead to the loss of overall scene information. Summary of the Invention

[0004] The purpose of the present invention is to overcome the shortcomings of the prior art, and provide a 3D object detection method, device, computer device and readable storage medium for adaptively optimizing BEV features, which greatly improves the accuracy of 3D object detection in surround-view images.

[0005] The present invention adopts the following technical solutions to achieve the above object. In the first aspect, the present invention provides a 3D object detection method for adaptively optimizing BEV features, including:

[0006] S1. Obtain panoramic image data from the Nuscenes dataset;

[0007] S2. Use a feature extraction backbone network to extract features from the panoramic image to obtain image features;

[0008] S3. Input the image features into a perspective conversion module to generate BEV features;

[0009] S4. The BEV feature optimizer includes 6 layers of teansformer decoder layer. Query boxes are uniformly set in the three-dimensional space in advance to query and fuse the BEV features and image features respectively to obtain adaptive fused features. The fused features are sent to the localization detection head and the classification detection head to obtain the 3D box attribute pred box and the label score lebel. At the same time, the query box is updated using the pred box and sent to the next layer of decoder layer;

[0010] S5. Select the top N pred boxes with high lebel label scores and use the corresponding fused features to update and optimize the BEV features;

[0011] S6. Generate heatmaps for each category in the Nuscenes dataset and attributes such as the center point height of the target, the length, width, and height of the bounding box, the yaw angle around the z-axis, and the speeds along the x-axis and y-axis through the optimized BEV features;

[0012] Further, step S3 specifically includes:

[0013] S31. Manually generate depth preset values from 1 to 60 with a spacing of 1 meter at the image feature level. At this time, in the pixel space, the coordinates of each depth point can be expressed as (u, v, d). Using the internal and external parameters of the camera, the depth points can be converted from the pixel coordinate system to the ego-vehicle coordinate system to obtain the pseudo-point cloud coordinates (x, y, z) in the BEV space. The conversion formula is:

[0014]

[0015] Among them, represents the camera external parameters, represents the camera internal parameters.

[0016] S32. Send the image features into the semantic feature extractor and the depth generator respectively to obtain semantic features and depth probability confidence. Perform a vector cross product on the two to obtain the pseudo-point cloud semantic features in the BEV space;

[0017] S33. Construct a BEV grid with a length and width range of -51.2 to 51.2 and a spacing of 0.8 m. It can be divided into 128×128 cubic columns with equal bottom areas and infinite heights. This cubic column is called a Pillar. Sum the semantic features of all the pseudo-point clouds in the same Pillar to obtain the BEV Feature;

[0018] Furthermore, step S4 specifically includes:

[0019] S41. Map the query box in the three-dimensional space to the BEV plane, extract the features of all BEV grids within the oriented box, and calculate the mean to obtain the RoI feature;

[0020] S42. All the RoI features interact with each other through self-attention to obtain global information and prevent multiple query boxes from converging to the same RoI feature. The self-attention formula is:

[0021]

[0022] Among them, Softmax() represents the normalization function, d represents the dimension of K, τ is a learnable parameter, and D represents the distance between the center points of pairwise query boxes.

[0023] S43. Send the RoI feature into the linear layer to predict the offset and couple it with the center point of the query box to form a reference point. Project the reference point onto the image feature plane and sample the features using bilinear interpolation. The projection formula is:

[0024]

[0025] Among them, represents the extrinsic camera parameters, represents the intrinsic camera parameters.

[0026] S44. Fuse the RoI feature and the sampled image feature at the channel level as the query feature. After passing through the ffn layer, send it into the localization detection head and the classification detection head to obtain the target attribute pred box and the target class score lebel. Use the pred box to update the query box and use it as the input of the next decoder layer. Iterate a total of 6 times.

[0027] Furthermore, step S5 specifically includes:

[0028] S51. Manually generate the center points P of the query box output by the last decoder layer on 6 faces k , send the RoI features corresponding to the query box into a linear layer to predict the offset, and couple with P k to obtain P' k , sample the planar features using the projection formula of S33 to obtain F k ;

[0029] S52. Map P' k to the BEV plane to obtain two-dimensional coordinates (i k , j k ), and its corresponding BEV grid index is G k , fuse and F k to obtain the optimized features The formula is:

[0030]

[0031] where n is the number of P' k falling within the current BEV grid.

[0032] In a second aspect, the present invention provides a 3D object detection device for panoramic images, which is used to implement the 3D object detection method for adaptively optimizing BEV features as described above. The detection device includes:

[0033] A panoramic image dataset loading module, specifically used for:

[0034] Obtain panoramic images according to the scene sequence;

[0035] Obtain the 3D object box information in the global coordinate system corresponding to the scene sequence, and transform the 3D object box positioning information to the front camera coordinate system according to the transformation matrix. Taking this as the reference coordinate system, calculate the yaw angle and the speed along the xy axis in this coordinate system at the same time;

[0036] Filter out the objects in the 3D object box without lidar or ladar point clouds to prevent the interference of visually unreasonable objects on the training of the network model;

[0037] A perspective conversion module, which is used to select a model for constructing BEV features for BEV perception for perspective conversion, specifically used for:

[0038] Select Lift-Splat as the model architecture for perspective conversion;

[0039] An adaptive BEV feature optimization module, which is used to select a method based on the transformer query mechanism to optimize BEV features, specifically used for:

[0040] Select SparseBEV as the model architecture for adaptive BEV feature optimization;

[0041] A training module, which is used to select a neural network model for 3D object detection of panoramic images for training, specifically for:

[0042] Select ResNet50 and FPN as the feature extraction backbone network, and the CenterPoint model as the 3D object detection trainer;

[0043] Input the image features into the above-mentioned perspective conversion module and adaptive BEV feature optimization module to obtain BEV features;

[0044] Input the BEV features and train the CenterPoint model;

[0045] A performance evaluation module, which is used for model performance evaluation, specifically for:

[0046] Detect the input panoramic image in the BEV space and use non-maximum suppression merging to obtain the final prediction result;

[0047] Use the test images of the Nuscenes panoramic image dataset as the evaluation dataset, and calculate mAP and NDS as the model performance evaluation metrics;

[0048] Input the evaluation dataset to obtain the evaluation results.

[0049] In a third aspect, the present invention provides a computer device, including a memory, the memory stores program instructions, and when the program instructions run, they execute the 3D object detection method for adaptively optimizing BEV features as described above in the claims.

[0050] The beneficial effects of the present invention are as follows:

[0051] By projecting the pixels in the panoramic image into the three-dimensional space and further mapping them to the BEV plane, the present invention effectively retains the deep semantic information of the image, providing a richer semantic basis for subsequent processing.

[0052] The present invention introduces an advanced sparse query mechanism to adaptively optimize the BEV features, greatly enriching the feature expression, effectively alleviating the problems of feature sparsity and uneven information distribution existing in traditional methods, and significantly improving the representativeness of data. At the same time, the present invention introduces an offset at the fixed point of the region of interest during adaptive optimization, enabling the network to focus more precisely on target instances rather than general 3D boxes.

[0053] Based on the neural network model trained with the surround-view image dataset, the present invention can accurately identify targets in complex autonomous driving scenarios, achieving efficient and stable 3D object detection for surround-view images and demonstrating excellent scene understanding capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 FIG. is a flowchart of a 3D object detection method for adaptively optimizing BEV features provided by an embodiment of the present invention;

[0055] Figure 2 FIG. is a structural block diagram of a 3D object detection device for surround-view images provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0057] Embodiment 1

[0058] The present invention utilizes the Nusenens dataset to construct and optimize BEV features for training a 3D object detection model of a convolutional neural network. As Figure 1 shown, the 3D object detection method for adaptively optimizing BEV features, which can also be referred to as the 3D object detection method for adaptively optimizing BEV features, specifically includes the following steps:

[0059] Step 1: Obtain surround-view image data from the Nuscenes dataset. The specific steps are as follows:

[0060] 1) Read surround-view image data from the Nuscenes dataset. This dataset contains rich 3D object annotation data and surround-view images, which are suitable for autonomous driving tasks. For each scene sequence, extract the corresponding surround-view images and obtain the 3D object bounding box information in the global coordinate system of the scene.

[0061] 2) Convert the 3D object bounding box information from the global coordinate system to the front-view camera coordinate system and use this coordinate system as the reference coordinate system for subsequent processing.

[0062] 3) Filter out the targets that do not contain Lidar or Ladar point clouds to ensure that only valid target data is used and to avoid invalid targets that interfere with training.

[0063] Step 2: Use the feature extraction backbone network to extract features from the surround-view images to obtain image features. The specific steps are as follows:

[0064] 1) Use the pre-trained ResNet50 and FPN to extract features from the surround-view images to obtain the feature maps of the surround-view images;

[0065] Step 3: Input the image features into the perspective conversion module to generate BEV features. The specific steps are as follows:

[0066] 1) Manually generate depth preset values from 1 to 60 with a spacing of 1 meter at the image feature level. At this time, in the pixel space, the coordinates of each depth point can be expressed as (u, v, d). Using the internal and external parameters of the camera, the depth points can be converted from the pixel coordinate system to the ego-vehicle coordinate system to obtain the pseudo-point cloud coordinates (x, y, z) in the BEV space. The conversion formula is:

[0067]

[0068] Among them, represents the external parameters of the camera, represents the internal parameters of the camera.

[0069] 2) Send the image features into the semantic feature extractor and the depth generator respectively to obtain semantic features and depth probability confidence. Perform a vector cross product on the two to obtain the semantic features of the pseudo-point cloud in the BEV space;

[0070] 3) Construct a BEV grid with a length and width range of -51.2 to 51.2 and a spacing of 0.8 m, which can be divided into 128×128 cubic columns with equal bottom areas and infinite heights. This cubic column is called a Pillar. Sum the semantic features of all the pseudo-point clouds in the same Pillar to obtain the BEV Feature;

[0071] Step 4: Optimize and update the BEV features. The specific steps are as follows:

[0072] 1) Uniformly set 900 query boxes (x, y, z, w, l, h, sinθ, cosθ, v x , v y ) in the three-dimensional space in advance;

[0073] 2) Map the query boxes in the three-dimensional space to the BEV plane and extract the features of all the BEV grids within the oriented box, and calculate the mean value to obtain the RoI features;

[0074] 3) All the RoI features interact with each other through self-attention to obtain global information and avoid multiple query boxes converging to the same RoI feature. The self-attention formula is:

[0075]

[0076] Among them, Softmax() represents the normalization function, d represents the dimension of K, τ is a learnable parameter, and D represents the distance between the centers of pairwise query boxes.

[0077] 4) Send the RoI feature to the linear layer to predict the offset and couple the query box center point to form a reference point. Project the reference point to the image feature plane and use bilinear interpolation to sample the features. The projection formula is:

[0078]

[0079] in, represents the camera extrinsic parameters, Indicates the camera internal parameters.

[0080] 5) At the channel level, the RoI features and the sampled image features are fused as query features, and sent to the positioning detection head and the classification detection head after passing through the FFN layer to obtain the target attribute pred box and the target category score lebel. The pred box is used to update the query box and serve as the input of the next decoder layer, with a total of 6 iterations.

[0081] 6) Manually generate the query box output by the last decoder layer at the center point P of the six faces k , the RoI feature corresponding to the query box is sent to the linear layer to predict the offset, and the coupling P k , and we get P′ k , use the projection formula of S33 to sample the plane features F k ;

[0082] 7) P′ k Mapping to the BEV plane yields the two-dimensional coordinates (i k ,j k ), whose corresponding BEV grid index is G k , fusion and F k , and get the optimized features The formula is:

[0083]

[0084] Where n is the P′ falling on the current BEV grid k The number of

[0085] Step 5: Select the convolutional neural network model for 3D object detection in surround view images for training. The specific steps are as follows:

[0086] 1) Select a basic convolutional neural network model, and the present invention uses CenterPoint to describe the training process and performance evaluation;

[0087] 2) Input the optimized BEV features and train the model;

[0088] Step 6: Performance evaluation of the model:

[0089] 1) During prediction, the input panoramic image will be detected in the BEV space and merged using non-maximum suppression (NMS) to obtain the final prediction result;

[0090] 2) Use the test set of the Nuscenes dataset as the evaluation dataset, and use mAP and NDS as the model performance evaluation metrics;

[0091] 3) Input the evaluation panoramic image and count the evaluation results;

[0092] The present invention first constructs BEV features using panoramic images and designs a method for adaptively optimizing BEV features, thereby obtaining a BEV feature with accurate scene and target information, thus alleviating the problems of feature sparsity and uneven information distribution existing in traditional methods. For the convolutional network model used for 3D object detection, training with the Nuscenes dataset can obtain a 3D object detection model with good performance, complete the 3D object detection task of panoramic images, and can be applied to many application fields such as robot vision, consumer electronics, security, autonomous driving, human-computer interaction, image retrieval, intelligent monitoring, augmented reality, virtual reality, etc., and also provides a basis for the realization of more complex tasks such as semantic segmentation and scene understanding.

[0093] Embodiment 2

[0094] As Figure 2 described, a 3D object detection device for panoramic images is used to implement the 3D object detection method for adaptively optimizing BEV features as in Embodiment 1. The device includes a panoramic image dataset loading module, a perspective conversion module, an adaptive BEV feature optimization module, a training module, a performance evaluation module, and an input evaluation dataset.

[0095] The panoramic image dataset loading module is used to obtain panoramic images according to the scene sequence;

[0096] Specifically, obtain the 3D object box information in the global coordinate system corresponding to the scene sequence, and transform the 3D object box positioning information to the front view camera coordinate system according to the transformation matrix. Taking this as the reference coordinate system, calculate the yaw angle and the speed along the xy axis in this coordinate system at the same time;

[0097] Filter out the targets in the 3D object box without lidar or ladar point clouds to prevent interference from visually unreasonable targets to the training of the network model.

[0098] The perspective transformation module is used to select a BEV feature construction model for BEV perception for perspective transformation, specifically for: selecting Lift-Splat as the model architecture for perspective transformation.

[0099] The adaptive BEV feature optimization module is used to select a method based on the transformer query mechanism to optimize BEV features, specifically for: selecting SparseBEV as the model architecture for adaptive BEV feature optimization.

[0100] The training module is used to select a neural network model for 3D object detection of panoramic images for training, specifically for:

[0101] Select ResNet50 and FPN as the feature extraction backbone network, and the CenterPoint model as the 3D object detection trainer; secondly, input the image features into the above perspective transformation module and adaptive BEV feature optimization module to obtain BEV features; and input the BEV features to train the CenterPoint model.

[0102] The performance evaluation module is used for the performance evaluation of the model, specifically for: detecting the input panoramic image in the BEV space and using non-maximum suppression for merging to obtain the final prediction result; and using the test images of the Nuscenes panoramic image dataset as the evaluation dataset, and calculating mAP and NDS as the model performance evaluation metrics.

[0103] Input the evaluation dataset to obtain the evaluation result.

[0104] Embodiment III

[0105] This embodiment provides a computer device, including a memory, and the memory stores program instructions. When the program instructions run, they execute the 3D object detection method for adaptively optimizing BEV features as in Embodiment I. Those of ordinary skill in the art can understand that all or part of the processes in the above embodiment methods can be completed by program instructions instructing relevant hardware.

[0106] Embodiment IV

[0107] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, it implements the 3D object detection method for adaptively optimizing BEV features as in Embodiment I.

[0108] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in the relevant field. Any changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.

Claims

1. A 3D object detection method for adaptively optimizing BEV features, characterized in that: The steps include: S1. Obtain surround image data from the Nuscenes dataset; S2, extracting features from the surround view image in S1 using a feature extraction backbone network to obtain image features; S3, inputting the image features in S2 into a perspective conversion module to generate BEV features; S4, BEV feature optimizer includes 6 layers of teasformer decoder layer, which pre-sets query boxes evenly in three-dimensional space to query and fuse the BEV features and image features respectively, obtains adaptive fusion features, and sends the fusion features to the positioning detection head and the classification detection head to obtain the 3D box attribute pred box and label score lebel, and uses the pred box to update the query box and send it to the next decoder layer; S5, take the top N pred boxes with high Lebel label scores in S4, and use their corresponding fusion features to update and optimize the BEV features; S6. Generate a heat map of each category in the Nuscenes dataset and the attributes of the target through the optimized BEV features, wherein the attributes include the center point height, the length, width, and height of the bounding box, the yaw angle around the z-axis, and the speed attributes along the x-axis and y-axis.

2. The 3D target detection method of adaptively optimizing BEV features according to claim 1, characterized in that: The S3 specifically includes: S31. Manually generate 1 to 60 depth preset values ​​with a spacing of 1 meter at the image feature level. At this time, in the pixel space, the coordinates of each depth point can be expressed as (u, v, d). The depth point can be converted from the pixel coordinate system to the vehicle coordinate system using the internal and external parameters of the camera to obtain the pseudo point cloud coordinates (x, y, z) in the BEV space. The conversion formula is: in, represents the camera extrinsic parameters, Indicates the camera internal parameters. S32, sending the image features to the semantic feature extractor and the depth generator respectively, obtaining the semantic features and the depth probability confidence, and performing vector cross multiplication on the two to obtain the semantic features of the pseudo point cloud in the BEV space; S33. Construct a BEV grid with a length and width range of -51.2 to 51.2 and a spacing of 0.8m. It can be divided into 128×128 cubic columns with equal bottom area and infinite height. The cubic column is called a Pillar. Sum the semantic features of all pseudo point clouds in the same Pillar to obtain the BEV Feature.

3. The 3D target detection method of adaptively optimizing BEV features according to claim 1, characterized in that: Step S4 specifically includes: S41, map the query box in the three-dimensional space to the BEV plane to extract the features of all BEV grids in the directed box, and average them to obtain the RoI features; S42. All RoI features interact with each other through self-attention to obtain global information and avoid multiple query boxes converging to the same RoI feature. The self-attention formula is: Among them, Softmax() represents the normalization function, d represents the dimension of K, τ is a learnable parameter, and D represents the distance between the center points of each query box. S43, send the RoI feature to the linear layer to predict the offset and couple the query box center point to form a reference point, project the reference point to the image feature plane and use bilinear interpolation to sample the features. The projection formula is: in, represents the camera extrinsic parameters, Indicates the camera internal parameters. S44. At the channel level, the RoI features and the sampled image features are fused as query features, and sent to the positioning detection head and the classification detection head after the FFN layer to obtain the target attribute pred box and the target category score lebel. The pred box is used to update the query box and used as the input of the next decoder layer, with a total of 6 iterations.

4. The 3D target detection method of adaptively optimizing BEV features according to claim 1, characterized in that: Step S5 specifically includes: S51. Manually generate the query box output by the last decoder layer at the center point P of the six faces. k , the RoI feature corresponding to the query box is sent to the linear layer to predict the offset, and the coupling P k ,get Sample the planar features F using the projection formula of S33 k ; S52, will Mapping to the BEV plane yields the two-dimensional coordinates (i k ,j k ), whose corresponding BEV grid index is G k , fusion and F k , and get the optimized features The formula is: Where n is the P′ falling on the current BEV grid k The number of 5. A 3D target detection device for surround image, used to implement the 3D target detection method for adaptively optimizing BEV features as described in any one of claims 1 to 4, characterized in that: The detection device includes: Surround view image dataset loading module, specifically used for: Acquire surround view images according to scene sequences; Get the 3D target frame information in the global coordinate system corresponding to the scene sequence, and transform the 3D target frame positioning information to the forward-looking camera coordinate system according to the transformation matrix, using this as the reference coordinate system, and calculate the yaw angle and speed along the xy axis in this coordinate system; Filter out targets without lidar or ladar point clouds in the 3D target box to prevent visually unreasonable targets from interfering with the network model training; The perspective conversion module is used to select the BEV features perceived by the BEV to build a model for perspective conversion, specifically for: Select Lift-Splat as the model architecture for perspective transformation; The adaptive BEV feature optimization module is used to select a method based on the transformer query mechanism to optimize BEV features. Specifically, it is used to: SparseBEV was selected as the model architecture for adaptive BEV feature optimization; The training module is used to select a neural network model for 3D object detection in surround images for training. Specifically, it is used to: Select ResNet50 and FPN as the feature extraction backbone network, and CenterPoint model as the 3D object detection trainer; The image features are input into the above-mentioned view conversion module and the adaptive BEV feature optimization module to obtain BEV features; Input BEV features to train the CenterPoint model; The performance evaluation module is used to evaluate the performance of the model, specifically for: The input surround view image is detected in the BEV space and merged using non-maximum suppression to obtain the final prediction result; The test images of the Nuscenes surround image dataset are used as the evaluation dataset, and mAP and NDS are calculated as the model performance evaluation indicators; Input the evaluation data set and obtain the evaluation results.

6. A computer device comprising a memory storing program instructions, characterized in that: When the program instructions are executed, the 3D target detection method for adaptively optimizing BEV features as described in any one of claims 1 to 4 is executed.

Citation Information

Patent Citations

  • 3D target detection method based on point cloud data and multi-view image data fusion

    CN115512132A

  • BEV visual perception method based on multiple cameras

    CN115512326A

  • Multi-modal weak supervised learning 3D target detection method based on 3D point labeling

    CN117953205A

  • Cross-view fusion three-dimensional target detection method based on cross attention

    CN118351404A

  • Multi-source heterogeneous sensing information fusion target detection and tracking network model and method

    CN118823308A

Cited By

  • Image processing method and device, electronic equipment and readable storage medium

    CN121438251A