Automatic driving perception method based on 3D detection

By extracting the feature of projected 3D points on the image plane and parallel 3D variables, and using the 3D enclosing frame memory queue embedding parameter information, combined with the interactive attention mechanism for deep interaction between objects and scenes, the existing monocular 3D detection methods are solved, and more efficient autonomous driving perception is achieved.

CN119964125APending Publication Date: 2025-05-09BEIJING TIEMUNIU INTELLIGENT MACHINERY TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510067120.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing monocular 3D detection method relies on the ROI generated by 2D detection, resulting in increased noise, large calculation amount, decreased detection accuracy, and difficult to understand the scene-level 3D spatial structure.

Method used

Using an autonomous driving perception method based on 3D detection, the feature extraction of projected 3D points on the image plane and parallel 3D variables are obtained, the key point features and 3D enclosing frame features are obtained, the 3D enclosing frame memory queue embedding parameter information is used, and the deep interaction between objects and scenes is performed through the interactive attention mechanism.

Benefits of technology

It improves detection accuracy, reduces training time, enhances perception of surrounding obstacles, simplifies model structure, and reduces the amount of parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964125A_ABST
    Figure CN119964125A_ABST
Patent Text Reader

Abstract

The invention provides an automatic driving perception method based on 3D detection. The automatic driving perception method comprises the following steps: carrying out feature extraction on projected 3D points and parallel 3D variables on an image plane to obtain key point features and 3D bounding box features; embedding parameter information by using a 3D bounding box memory queue, encoding 3D bounding box features through the parameter information, and inputting the encoded 3D bounding box features into a 3D bounding box guide decoder for decoding; and in a decoding stage, performing feature fusion on the 3D bounding box features and the target query through a query interaction attention mechanism to obtain a learnable object query, and performing deep interaction between the object and the scene by combining the learnable object query and the key point features with the cross attention mechanism. According to the automatic driving perception method based on 3D detection, the detection accuracy is improved, and the training time is greatly shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and in particular relates to an autonomous driving perception method based on 3D detection. Background Art

[0002] Autonomous driving technology is one of the hottest research directions in the field of artificial intelligence. Accurate and reliable 3D object detection is the key to achieving autonomous driving. Autonomous vehicles need to accurately perceive the surrounding environment, including vehicles, pedestrians, obstacles, etc. 3D object detection can provide the target's precise position, size and orientation information, and provide more comprehensive information for vehicle decision-making planning. Estimating the 3D direction and translation of objects is crucial for autonomous navigation and driving without infrastructure.

[0003] Compared with other sensors, monocular cameras are cheaper and easier to deploy. At the same time, monocular cameras are smaller and easier to integrate into the vehicle body.

[0004] At present, monocular 3D detection still has the following limitations for complex road scenes with dynamic objects such as pedestrians, vehicles, obstacles, and factors such as lighting changes and weather effects:

[0005] 3D monocular detection relies heavily on the ROI (Region of Interest) generated by 2D detection. The acquired ROI is used to predict the 3D object pose. In this way, the 2D detection network is redundant and will introduce non-negligible noise into the 3D detection.

[0006] If the 3Dbbox (3D bounding box) is directly encoded and embedded, the amount of calculation in the training phase will be huge, resulting in slower network convergence and reduced detection accuracy.

[0007] Most existing methods follow the traditional 2D detector, first localizing the object center and then predicting the 3D properties through neighboring features. However, using only local visual features is not enough to understand the scene-level 3D spatial structure and ignores the deep relationship between distant objects. Summary of the invention

[0008] The purpose of the present invention is to solve the problems in the prior art and propose an autonomous driving perception method based on 3D detection, which not only improves the detection accuracy but also greatly reduces the training time.

[0009] In order to achieve the above object, the present invention adopts the following technical scheme.

[0010] The autonomous driving perception method based on 3D detection performs feature extraction on the projected 3D points on the image plane and the parallel 3D variables to obtain key point features and 3D bounding box features;

[0011] Embedding parameter information using a 3D bounding box memory queue, encoding 3D bounding box features using the parameter information, and then inputting the 3D bounding box to guide a decoder for decoding;

[0012] In the decoding stage, the 3D bounding box features are fused with the target query through the query interaction attention mechanism to obtain a learnable object query, and the learnable object query is combined with the key point features through the cross attention mechanism to perform deep interaction between objects and scenes.

[0013] Furthermore, the step of extracting features of the projected 3D points on the image plane and the parallel 3D variables to obtain key point features and 3D bounding box features includes:

[0014] A single object is represented by a single specific keypoint;

[0015] Define the key point as the projected 3D center of the object on the image plane;

[0016] Projected keypoints use camera parameters to recover the 3D position of individual objects;

[0017] Use [xyz] T Represents the 3D center of a single object in the camera frame;

[0018] 3D point to point [x c y c ] T The projection of is obtained in homogeneous form through the camera intrinsic matrix K:

[0019] For a single ground key point, a Gaussian kernel is used to calculate and distribute its corresponding down-sampled position on the feature map;

[0020] Each 3D box on the image consists of 8 2D points [x b,1-8 y b,1-8 ] T express.

[0021] Furthermore, it also includes:

[0022] Perform 3D bounding box extraction, including the regression head predicting the basic variables required to construct the 3D bounding box of a single keypoint on the heat map. The 3D information is encoded as a 7-tuple: δ z represents the depth offset, is the discretization offset due to downsampling, δ h δ w δ1 represents the residual dimension, α is the vector representation of the object observation angle;

[0023] The input variables are encoded as residual representations, and the feature map size of the regression result is

[0024] For a single object, its depth z is obtained by predefined scale and translation parameters μzσ z recover:

[0025] z=μ z +δ z δ z

[0026] According to the object depth z, by using the discretized projection centroid [x c y c ] T and downsampling offset Recover the position of a single object in the camera frame:

[0027]

[0028] Use precomputed average dimensions by category Retrieve object dimensions [hwl] T , which is calculated over the entire dataset;

[0029] The individual object dimensions are offset by using the residual dimension [δ h δ w δ l ] T recover:

[0030]

[0031] Using the observation angle α z And the object position to obtain the yaw angle θ:

[0032]

[0033] Using the yaw rotation matrix R θ Object size [hwl] T and position [xyz] T Construct the 8 corner points of the 3D bounding box in the camera frame:

[0034]

[0035] Furthermore, the embedding of parameter information using the 3D bounding box memory queue includes:

[0036] Design an N×K memory queue, where N is the number of frames stored and K is the number of objects stored per frame;

[0037] Information elements for 3D frames To store, the information elements corresponding to the foreground objects are selected and pushed into a memory queue.

[0038] Further, encoding the 3D bounding box feature by using the parameter information and then inputting the encoding into the 3D bounding box to guide the decoder for decoding includes:

[0039] Obtain key point feature information and 3D bounding box feature information to encode key points and 3D bounding boxes, and design two encoders to generate scene-level embeddings respectively;

[0040] Three encoder blocks are set for the 3D bounding box encoder and one encoder block is set for the key points. A single encoder block includes a self-attention layer and a feed-forward neural network.

[0041] Through the global self-attention mechanism, the 3D bounding box encoder explores the dependencies of 3D boxes from different regions and provides non-local geometric clues in the stereo space;

[0042] The input image is encoded from two perspectives by decoupling the 3D bounding box encoder and the keypoint encoder.

[0043] Further, the feature fusion of the 3D bounding box feature and the target query through the query interaction attention mechanism to obtain a learnable object query, and the deep interaction of the object and the scene by combining the learnable object query with the key point feature through the cross attention mechanism includes:

[0044] Perform 3D bounding box guided decoding, including based on non-local information Leverage a set of learnable object queries q∈R N×C Detect 3D objects via a depth-guided decoder, where N represents the predefined maximum number of objects in the input image;

[0045] A single decoder block contains a deep criss-cross attention layer, an inter-query self-attention layer, a visual criss-cross attention layer, and a feed-forward neural network;

[0046] Informative deep features are captured via a 3D bounding box criss-cross attention layer, where object queries and deep embeddings are linearly transformed into queries, keys, and values:

[0047]

[0048] Where Q q ∈R N×C , K D ,V D ∈R N×C ;

[0049] Compute query deep attention graph And aggregate the information depth features weighted by AD to generate the depth-aware query q', the formula is as follows:

[0050]

[0051] q′=Linear(A D V D ).

[0052] The 3D-aware query is fed into an inter-query self-attention layer for feature interaction between objects and a keypoint cross-attention layer for Collect key point semantics.

[0053] Furthermore, it also includes:

[0054] Match a single query to an object, compute the loss for a single query-label pair, and find the global optimal match using the Hungarian algorithm;

[0055] For a single query-label pair, the loss is divided into 2D attribute loss and 3D attribute loss;

[0056] The 2D attribute loss and 3D attribute loss are added together and denoted as L 2D and L 3D ;

[0057] Use L 2D as the matching cost for matching a single query-label pair.

[0058] Furthermore, it also includes,

[0059] Calculate the Overall loss function, including matching, to obtain N from N queries gt valid pairs, where N gt Represents the number of real objects;

[0060] The overall loss formula for training images is:

[0061]

[0062] Where L damp Focal Loss representing the classified foreground depth map.

[0063] In order to achieve the above-mentioned purpose, the present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a program running on the processor, and when the processor runs the program, the steps of the autonomous driving perception method based on 3D detection as described above are executed.

[0064] In order to achieve the above-mentioned objectives, the present invention also provides a computer-readable storage medium on which computer instructions are stored, and when the computer instructions are executed, the steps of the autonomous driving perception method based on 3D detection as described above are executed.

[0065] The present invention proposes an autonomous driving perception method based on 3D detection, which has the following beneficial effects:

[0066] The present invention can improve the algorithm's scene-level comprehension ability and enhance the perception of surrounding obstacles, while reducing the number of model parameters, discarding redundant 2D information, and accelerating the application of the model in the training and deployment stages.

[0067] Other features and advantages of the present invention will be set forth in the description which follows, and in part will be apparent from the description, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0069] Figure 1 This is a flow chart of an autonomous driving perception method based on 3D detection according to the present invention;

[0070] Figure 2 A schematic diagram of an autonomous driving perception model architecture based on 3D detection according to the present invention. DETAILED DESCRIPTION

[0071] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0072] Example 1

[0073] Figure 1 The flowchart of the autonomous driving perception method based on 3D detection according to the present invention is as follows. Figure 1 , the 3D detection-based autonomous driving perception method of the present invention is described in detail.

[0074] In step 101 , feature extraction is performed on the projected 3D points on the image plane and the parallel 3D variables.

[0075] Optionally, in terms of feature extraction, the backbone network uses the hierarchical fusion network DLA-34 as the backbone for extracting features because it can aggregate information across different layers. Following the same structure as objects as points, all hierarchical aggregation connections are replaced by deformable convolutional networks (DCNs), replacing redundant 2D information.

[0076] In this embodiment, the output feature map is downsampled 4 times relative to the original image. In order to extract more refined features, the input feature map is downsampled so that the resolution of the feature map output by the network is one-fourth of the resolution of the input feature map. Compared with the original implementation, this method uses GroupNorm (GN) to replace all BatchNorm (BN) operations, because BatchNorm has been shown to be less sensitive to batch size and more robust to training noise. This adjustment not only improves the detection accuracy, but also greatly reduces the training time.

[0077] Optionally, keypoint extraction is performed and a keypoint estimation network is defined similar to the objects as points network, where each object is represented by a specific keypoint. Instead of identifying the center of a 2D bounding box, a keypoint is defined as the projected 3D center of the object on the image plane. The projected keypoints can fully recover the 3D position of each object using the camera parameters. Let [xyz] T represents the 3D center of each object in the camera frame. The 3D point to the image plane point [x c y c ] T The projection of can be obtained in homogeneous form through the camera intrinsic matrix K:

[0078]

[0079] Optionally, for each ground truth keypoint, compute and distribute its corresponding downsampled position on the feature map using a Gaussian kernel;

[0080] Each 3D box on the image consists of 8 2D points [x b,1-8 y b,1-8 ] T express.

[0081] Optionally, 3D bounding box extraction is performed. The regression head predicts the basic variables needed to construct the 3D bounding box of each keypoint on the heatmap. The 3D information is encoded as a 7-tuple: δ z represents the depth offset, is the discretization offset due to downsampling, δ h δ w δ1 represents the residual dimension, and α is the vector representation of the object's observation angle. All variables to be input are encoded as residual representations, that is, only the difference is focused on to reduce the learning interval and simplify the training task. The feature map size of the regression result is For each object, its depth z can be obtained by predefined scale and translation parameters μzσ z Recovery, that is:

[0082] z=μ z+δ z δ z

[0083] Alternatively, given the object depth z, the projection centroid [x c y c ] T and downsampling offset To recover the position of each object in the camera frame:

[0084]

[0085] Optionally, to retrieve the object dimensions [hwl] T , using a precomputed per-category average dimension This dimension is calculated over the entire dataset. Each object dimension can be calculated by

[0086] Use the residual dimension shift [δ h δ w δ l ] T To restore:

[0087]

[0088] The yaw angle θ can be calculated using the observation angle α z And the object position is obtained:

[0089]

[0090] Optionally, use the yaw rotation matrix R θ Object size [hwl] T and position [xyz] T Construct the 8 corner points of the 3D bounding box in the camera frame:

[0091]

[0092] Optionally, a memory queue embedding of the 3D bounding box is performed, including designing an N×K memory queue to achieve effective information embedding. N is the number of frames stored, and K is the number of objects stored per frame. By embedding the 7 information elements of the 3D box For storage, the above information corresponding to the foreground object (with the top K highest classification scores) is selected and pushed into the memory queue. The entry and exit of the memory queue follow the first-in-first-out (FIFO) rule. When information from a new frame is added to the memory queue, the oldest frame will be discarded. The proposed memory queue is highly flexible and customizable, and the maximum memory size can be freely controlled during training and inference.

[0093] In step 102, during the encoding phase, a 3Dbox memory queue embedding is designed for additional parameter information to reduce the amount of network training parameters.

[0094] Optionally, a keypoint and 3D bounding box encoder is performed. Given the keypoint and 3D box information, two Transformer encoders are specially designed to generate scene-level embeddings with global receptive fields. Three blocks are set for the 3D box encoder and only one block is set for the keypoints, because the discrete keypoint information is easier to encode than the rich 3D box information. Each encoder block consists of a self-attention layer and a feed-forward neural network (FFN). Through the global self-attention mechanism, the 3D box encoder explores the dependencies of 3D boxes from different regions, thereby providing non-local geometric clues in the stereo space. In addition, the decoupling of the 3D box encoder and the keypoint encoder enables them to better learn their own features to encode the input image from two perspectives.

[0095] Optionally, a 3D bounding box guided decoder is performed, including one based on non-local information Leverage a set of learnable object queries q∈R N×C 3D objects are detected via a depth-guided decoder, where N represents the predefined maximum number of objects in the input image. Each decoder block sequentially contains a depth criss-cross attention layer, an inter-query self-attention layer, a visual criss-cross attention layer, and a FFN. Specifically, the query first passes through a 3D box criss-cross attention layer to capture informative deep features, where the object query and depth embedding are linearly transformed into query, key, and value:

[0096]

[0097] Where Q q ∈R N × C , K D ,V D ∈R N × C Then, the query depth attention map is calculated And aggregate the informative deep features weighted by AD to generate a depth-aware query q', as follows:

[0098]

[0099] q′=Linear(A D V D ).

[0100] This mechanism enables each object query to adaptively capture spatial cues from the depth-guided regions of 3D information, leading to a better understanding of the scene-level space. The 3D-aware query is then fed into an inter-query self-attention layer for feature interactions between objects, and a keypoint cross-attention layer to extract features from the keypoints. Gathering keypoint semantics. We stack three decoder blocks to fully fuse scene-level depth cues into object queries.

[0101] In step 103, in the decoding stage, since the local query cannot adaptively estimate its 3D properties based on the depth-guided region on the image, an attention mechanism for query interaction is designed to make the 3D object a learnable query, and perform deep object-scene interaction through 3D box and key point information combined with a cross-attention mechanism.

[0102] Optionally, after the decoder, depth-aware queries are fed into a series of MLP-based heads for 3D property prediction, including object category, 2D size, projected 3D center, depth, 3D size, and orientation. For inference, these perspective properties are converted to 3D space bounding boxes using camera parameters without NMS post-processing or predefined anchors. For training, unordered queries are matched with true labels and the loss for paired queries is calculated.

[0103] Optionally, bipartite matching is performed, which involves computing the loss for each query-label pair in order to correctly match each query to a true object, and finding the global optimal match using the Hungarian algorithm. For each pair, we combine the losses of the six attributes into two groups. The first group contains object category, 2D size, and 3D center of projection, as these attributes mainly involve the 2D visual appearance of the image. The second group includes depth, 3D size, and orientation, which are 3D spatial properties of the object.

[0104] Optionally, sum the losses of the two groups separately and denote it as L 2D and L 3D Since the network usually predicts 3D attributes less accurately than 2D attributes, especially at the beginning of training, L 3D The value of is unstable and will interfere with the matching process. We only use L 2D as the matching cost for each query-label pair.

[0105] Optionally, calculate the Overall loss function, including matching, to obtain N from N queries gt valid pairs, where N gt represents the number of true objects. Then, the overall loss formula for training images is:

[0106]

[0107] Where L damp Focal Loss representing the classified foreground depth map.

[0108] The present invention proposes an autonomous driving perception method based on 3D detection, which adopts 3D target detection of the detr architecture to improve the scene-level 3D understanding ability. In terms of feature extraction, feature extraction is directly performed on the projected 3D points on the image plane and the parallel 3D variables, replacing the redundant 2D information. In the encoding stage, a 3Dbox memory queue embedding is designed for additional parameter information to reduce the amount of network training parameters. In the decoding stage, a local query cannot adaptively estimate its 3D properties according to the depth-guided area on the image, so a method is designed to formulate the 3D object as a learnable query, and a deep object-scene interaction is performed through the 3Dbox and key point information combined with the cross-attention mechanism.

[0109] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a program running on the processor, and the processor executes the above-mentioned steps of autonomous driving perception based on 3D detection when running the program.

[0110] The present invention also provides a computer-readable storage medium having computer instructions stored thereon, and the computer instructions, when executed, execute the above-mentioned autonomous driving perception based on 3D detection. The autonomous driving perception based on 3D detection is described in the aforementioned part and will not be repeated here.

[0111] Those skilled in the art can understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention is described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions recorded in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An autonomous driving perception method based on 3D detection, characterized in that: include: Feature extraction is performed on the projected 3D points on the image plane and the parallel 3D variables to obtain key point features and 3D bounding box features; Embedding parameter information using a 3D bounding box memory queue, encoding 3D bounding box features using the parameter information, and then inputting the 3D bounding box to guide a decoder for decoding; In the decoding stage, the 3D bounding box features are fused with the target query through the query interaction attention mechanism to obtain a learnable object query, and the learnable object query is combined with the key point features through the cross attention mechanism to perform deep interaction between objects and scenes.

2. The autonomous driving perception method based on 3D detection according to claim 1, characterized in that: The feature extraction of the projected 3D points on the image plane and the parallel 3D variables to obtain key point features and 3D bounding box features includes: A single object is represented by a single specific keypoint; Define the key point as the projected 3D center of the object on the image plane; Projected keypoints use camera parameters to recover the 3D position of individual objects; Use [xyz] T Represents the 3D center of a single object in the camera frame; 3D point to point [x c y c ] T The projection of is obtained in homogeneous form through the camera intrinsic matrix K: For a single ground key point, a Gaussian kernel is used to calculate and distribute its corresponding down-sampled position on the feature map; Each 3D box on the image consists of 8 2D points [x b,1-8 y b,1-8 ] T express.

3. The autonomous driving perception method based on 3D detection according to claim 2, characterized in that: Also includes, Perform 3D bounding box extraction, including the regression head predicting the basic variables required to construct the 3D bounding box of a single keypoint on the heat map. The 3D information is encoded as a 7-tuple: δ z represents the depth offset, is the discretization offset due to downsampling, δ h δ w δ1 represents the residual dimension, α is the vector representation of the object observation angle; The input variables are encoded as residual representations, and the feature map size of the regression result is For a single object, its depth z is obtained by predefined scale and translation parameters μzσ z recover: z=μ z +d z s z According to the object depth z, by using the discretized projection centroid [x c y c ] T and downsampling offset Recover the position of a single object in the camera frame: Use precomputed average dimensions by category Retrieve object dimensions [hwl] T , which is calculated over the entire dataset; The individual object dimensions are offset by using the residual dimension [δ h δ w δ l ] T recover: Using the observation angle α z And the object position to obtain the yaw angle θ: Using the yaw rotation matrix R θ Object size [hwl] T and position [xyz] T Construct the 8 corner points of the 3D bounding box in the camera frame:

4. The autonomous driving perception method based on 3D detection according to claim 1, characterized in that: The embedding of parameter information using the 3D bounding box memory queue includes: Design an N×K memory queue, where N is the number of frames stored and K is the number of objects stored per frame; Information elements for 3D frames To store, the information elements corresponding to the foreground objects are selected and pushed into a memory queue.

5. The autonomous driving perception method based on 3D detection according to claim 1, characterized in that: The encoding of the 3D bounding box feature by the parameter information and then inputting the encoding into the 3D bounding box to guide the decoder for decoding comprises: Obtain key point feature information and 3D bounding box feature information to encode key points and 3D bounding boxes, and design two encoders to generate scene-level embeddings respectively; Three encoder blocks are set for the 3D bounding box encoder and one encoder block is set for the key points. A single encoder block includes a self-attention layer and a feed-forward neural network. Through the global self-attention mechanism, the 3D bounding box encoder explores the dependencies of 3D boxes from different regions and provides non-local geometric clues in the stereo space; The input image is encoded from two perspectives by decoupling the 3D bounding box encoder and the keypoint encoder.

6. The 3D detection-based autonomous driving perception method according to claim 1, characterized in that: The method of fusing the 3D bounding box feature with the target query through the query interaction attention mechanism to obtain a learnable object query, and combining the learnable object query with the key point feature with the cross attention mechanism to perform deep interaction between objects and scenes includes: Perform 3D bounding box guided decoding, including based on non-local information Leverage a set of learnable object queries q∈R N×C Detect 3D objects via a depth-guided decoder, where N represents the predefined maximum number of objects in the input image; A single decoder block contains a deep criss-cross attention layer, an inter-query self-attention layer, a visual criss-cross attention layer, and a feed-forward neural network; Informative deep features are captured via a 3D bounding box criss-cross attention layer, where object queries and deep embeddings are linearly transformed into queries, keys, and values: Where Q q ∈R N×C , K D ,V D ∈R N×C ; Compute query deep attention graph And aggregate the information depth features weighted by AD to generate the depth-aware query q', the formula is as follows: q′=Linear(A D V D ). The 3D-aware query is fed into an inter-query self-attention layer for feature interaction between objects and a keypoint cross-attention layer for Collect key point semantics.

7. The autonomous driving perception method based on 3D detection according to claim 1, characterized in that: Also includes, Match a single query to an object, compute the loss for a single query-label pair, and find the global optimal match using the Hungarian algorithm; For a single query-label pair, the loss is divided into 2D attribute loss and 3D attribute loss; The 2D attribute loss and 3D attribute loss are added together and denoted as L 2D and L 3D ; Use L 2D as the matching cost for matching a single query-label pair.

8. The autonomous driving perception method based on 3D detection according to claim 7, characterized in that: Also includes, Calculate the Overall loss function, including matching, and obtain Ngt valid pairs from N queries, where N gt Represents the number of real objects; The overall loss formula for training images is: Where L damp Focal Loss representing the classified foreground depth map.

9. An electronic device, characterized in that: It includes a memory and a processor, wherein the memory stores a program running on the processor, and when the processor runs the program, it executes an autonomous driving perception method based on 3D detection as described in any one of claims 1-8.

10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed, an autonomous driving perception method based on 3D detection as described in any one of claims 1 to 8 is executed.