Three-dimensional sparse target detection method and system based on depth restoration

By introducing matching augmentation modules and consistent modules into the three-dimensional sparse object detection network, the candidate box collection and deep information transmission are optimized, and the problems of high computing cost and low detection accuracy in the prior art are solved, and efficient three-dimensional object detection is achieved.

CN120356200APending Publication Date: 2025-07-22TIANJIN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510420433.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-04
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing three-dimensional object detection methods have insufficient calculation cost and detection accuracy, especially the BEV framework has large memory consumption and long training time, the detection speed of the sparse detection paradigm is limited, and the geometric consistency between the two-dimensional detection results and the three-dimensional truth value is difficult to guarantee, resulting in limited detection accuracy and speed.

Method used

By introducing a matching augmentation module and a matching consistency module, combining geometric perturbation and effectiveness filtering strategies, optimizing the candidate box set, and building a joint constraint mechanism for IOU projection through two-dimensional and three-dimensional parameter spatial mapping constraints to ensure the correct transmission of depth information and improve detection accuracy and efficiency.

Benefits of technology

It significantly reduces the calculation cost, improves the accuracy and efficiency of three-dimensional object detection, can effectively avoid missed targets in complex scenarios, and improves the robustness and speed of the detector.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356200A_ABST
    Figure CN120356200A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional sparse target detection method and system based on depth restoration. The method comprises the following steps: sequentially obtaining multi-view images; obtaining an initial detection frame set under each view angle; generating a two-dimensional detection set and an augmented record under each view angle; identifying and processing high-quality detection objects in the two-dimensional detection set to further screen out a detection frame with a correct matching relationship to generate a matching sequence, and optimizing fusion features based on two-dimensional and three-dimensional parameter space mapping constraints; the optimized fusion features and the two-dimensional detection set are transmitted to a three-dimensional feature decoding module through a Query generator, and decoding information is obtained; the decoding information is converted into target three-dimensional information through a detection head; according to the method, a matching augmentation module and a matching consistency module are introduced on the basis of an MV2D network, so that compared with a typical BEV detector and other sparse normal form detectors, the calculation cost is greatly reduced, the algorithm detection rate is effectively ensured, and the three-dimensional target detection precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image detection, and particularly relates to a three-dimensional sparse object detection method and system based on depth repair. Background Art

[0002] Current mainstream three-dimensional object detection methods are mostly based on the BEV framework, while the detection system based on the sparse paradigm (i.e., non-BEV architecture) has not been fully explored. The traditional BEV framework consumes a large amount of memory when constructing the feature space, significantly prolongs the training time, and results in a relatively high proportion of low-confidence detection boxes. In contrast, the sparse detection paradigm realizes more efficient feature expression by directly establishing an interaction mechanism between image features and three-dimensional queries, combining a carefully designed two-dimensional spatial position encoding and an attention mask mechanism in the three-dimensional decoding process.

[0003] Li Z et al. proposed a method using a spatio-temporal transformer to learn a unified BEV representation, which utilizes spatial and temporal information by interacting with spatial and temporal features through BEV queries with a predefined grid shape. However, due to the significant increase in complexity and computational amount brought by the establishment of the BEV space and query attention interaction, and the superposition of features at different depths on a single BEV query, this method will bring more redundant results and limited detection speed.

[0004] Li Y et al. proposed a depth estimation module (BEVDepth), which uses the depth supervision of point clouds to guide deep learning; as the first method to comprehensively analyze how depth quality affects the entire system, it innovatively proposed encoding the camera internal and external into the deep learning module to make the detector robust to various camera settings. In addition, the network further introduced a depth refinement module to refine the learned depth; however, the depth error introduced by BEVDepth relying on the monocular depth estimation network will be exponentially amplified in the BEV space, resulting in a sharp drop in the three-dimensional positioning accuracy of distant objects, posing a safety hazard to high-speed autonomous driving scenarios.

[0005] Liu Y et al. proposed a new three-dimensional position-aware representation (PETR), which realizes position-aware object representation by embedding three-dimensional coordinate information in image features. Object queries can be updated by interacting with three-dimensional position-aware features and generating three-dimensional predictions; although PETR eliminates the traditional LSS explicit depth estimation module and the number of model parameters is reduced by more than half compared with BEVDepth, due to PETR relying on the network to autonomously learn three-dimensional geometric relationships, it is difficult to accurately predict the scale size of three-dimensional instances and the separation relationship between instances.

[0006] Wang Z et al. proposed an object query based on rich image semantics generated by a two-dimensional detector, which provides valuable prior information for multi-view learning tasks and established a strategy for dynamically generating three-dimensional queries (hereinafter referred to as the MV2D network); in the context of multi-view learning, the MV2D network effectively overcomes the limitations of single-view methods in complex scenarios, such as view occlusion and partial information loss, by fusing information from different views, thus significantly improving the performance and robustness of the model. Compared with traditional BEV frameworks and other sparse paradigm detection systems, although the MV2D network introduces two-dimensional detection results as prior inputs for generating three-dimensional queries, effectively improving the initial positioning accuracy of three-dimensional targets; however, this scheme has two key problems: firstly, it is difficult to guarantee the geometric consistency between the two-dimensional detection results and the final three-dimensional ground truth; at the same time, the dynamic fluctuation of the number of two-dimensional anchor boxes causes continuous oscillation of the model parameters, severely limiting the convergence speed and detection accuracy; secondly, in the feature interaction of multiple (usually 6 layers) decoding layers, due to the unclear supervision object of the target, the queries are affected by nearby objects and gradually deviate from the learning objects assigned by the two-dimensional detection part.

[0007] Based on this, in order to solve the problem of the constantly fluctuating assigned objects, it is necessary to design a corresponding matching algorithm based on a three-dimensional object detector guided by two-dimensional detection to ensure efficient inference speed while improving accuracy. Summary of the Invention

[0008] The purpose of the present invention is to provide a three-dimensional sparse object detection method based on depth repair to solve the above technical problems.

[0009] Another purpose of the present invention is to provide a training method for a three-dimensional sparse object detection network based on depth repair.

[0010] Another purpose of the present invention is to provide a three-dimensional sparse object detection system based on depth repair.

[0011] To this end, the technical solution of the present invention is as follows:

[0012] A three-dimensional sparse object detection method based on depth repair, the steps include:

[0013] S1. Obtain multi-view images through cameras set at different view positions,

[0014] S2. Input the multi-view images into the two-dimensional detection network of the three-dimensional sparse object detection network to obtain a set of initial detection boxes for each view;

[0015] S3. Input the initial detection box sets from each perspective into the matching augmentation module of the 3D sparse object detection network. By deleting redundant and incomplete detection boxes in the initial detection box sets and combining with the 2D effective noisy detection box sets obtained by the geometric perturbation enhancement strategy from the 2D ground truth annotation boxes, generate the 2D detection sets from each perspective and the augmentation records for distinguishing the sources of the detection boxes in the 2D detection sets;

[0016] S4. Input the initial detection box sets from each perspective into the matching consistency module of the 3D sparse object detection network, screen out high-quality detection boxes with correct matching relationships to generate a matching sequence; based on the 2D and 3D parameter space mapping constraints, optimize the fused features of the ROI features and the camera intrinsics with the maximum IOU value sequence and the matching sequence;

[0017] S5. After initializing the 3D space query of the optimized fused features obtained in step S4 and the 2D detection sets obtained in step S3 through the Query generator, transmit them to the 3D feature decoding module to obtain the decoding information.

[0018] Further, in step S3, the method for deleting redundant and incomplete detection boxes in the initial detection box sets is: perform triple-parameter constraint screening of confidence threshold filtering, non-maximum suppression, and maximum retention quantity control on the initial detection box sets in sequence to obtain an effective candidate box set.

[0019] Further, in step S3, the method for generating the 2D effective noisy detection box sets is:

[0020] Step 1. Make G copies of the 2D ground truth annotation boxes in the dataset and perform uniform distribution perturbation on the parameters of each candidate box according to the preset scaling factor s, and its expression is:

[0021] b i =(c x ,c y ,w,h), Δb i =(c x ·Δc x ,c y ·Δc y ,w·Δw,h·Δh)·s,

[0022] b i ′ =b i +Δb i ,

[0023] where Δc x , Δc y , Δw, Δh are perturbation terms randomly drawn from the uniform distribution: Δc x , Δc y, Δw, Δh ~ U(-1, 1); Further, a set of two-dimensional noise-added candidate boxes is obtained.

[0024] Step 2: Adopt a geometric constraint enhancement strategy to sequentially perform image boundary cropping on the set of noise-added candidate boxes and remove invalid candidate boxes that do not match the three-dimensional ground truth spatial position.

[0025] Furthermore, in step S4, the recognition and processing process of the detection boxes with high quality in the two-dimensional detection set and having a correct matching relationship is as follows:

[0026] 1) Construct a cross-modal IOU correlation matrix based on the IOU relationship between the two-dimensional ground truth annotation box and the corresponding two-dimensional detection box in the two-dimensional detection set;

[0027] 2) Select the matching serial number corresponding to the maximum IOU value of each detection box and the ground truth box in the correlation matrix, and add the detection boxes with IOU value ≥ IOU screening threshold to the pairing set;

[0028] 3) According to the M 2t3 projection dictionary, determine the matching relationship between the detection boxes in the pairing set and the three-dimensional ground truth box;

[0029] 4) According to the augmented record, remove the noise-added detection boxes in the MCM pairing set that cannot match the two-dimensional ground truth box and the three-dimensional ground truth box;

[0030] 5) Consider the detection boxes not added to the pairing set and the detection boxes deleted from the pairing set as illegal detection boxes, and uniformly modify their assignments in the augmented record S MAM to negative values to ensure that illegal detection boxes do not participate in subsequent processing;

[0031] 6) Construct a new sequence after reassigning S MAM as the matching sequence S MCM , and the sequence of the maximum IOU value corresponding to each detection box S IOU .

[0032] Furthermore, in step S4, the processing process of optimizing the fused feature of the ROI feature and the camera internal parameter based on the two-dimensional and three-dimensional parameter space mapping constraints with the maximum IOU value sequence and the matching sequence is as follows:

[0033] 1) Fuse the ROI feature and the camera internal parameter, and its expression is:

[0034]

[0035] In the formula, I 2D represents the two-dimensional feature after ROI processing, T i represents the camera internal parameter, C represents the feature splicing operation, is the feature processing module, I2DF It is the processing result of the fusion internal reference.

[0036] 2) Perform 0-1 coding conversion on the matching sequence S MCM to distinguish legal detection frames from illegal detection frames; and then fuse the matching sequence S MCM and the IOU value sequence S IOU , and its expression is:

[0037]

[0038] In the formula, represents the conversion module of the 0-1 coding of S MCM , represents a simple linear fusion conversion matrix, I DTHI is the optimized fusion feature.

[0039] Furthermore, the three-dimensional sparse object detection method based on depth repair further includes step S6 of inputting the decoded information into the detection head to output the three-dimensional position, size, angle, and speed information of the object.

[0040] A training method for a three-dimensional sparse object detection network based on depth repair, the training method includes: inputting a multi-view image set, the internal and external parameters of the corresponding camera, and a 2D_to_3D projection dictionary into the three-dimensional sparse object detection network for training, and applying matching loss supervision based on the matching sequence during the training process;

[0041] The matching loss function has the following expression:

[0042]

[0043] In the formula, represents the matching consistency loss between two dimensions and three dimensions, and λ 3d is the weight coefficient of the three-dimensional part, is the original MV2D loss, is the two-dimensional detection loss part, is the three-dimensional detection loss part;

[0044] has the following expression:

[0045]

[0046] In the formula, the σ function represents three-dimensional ground truth matching according to the matching sequence, and L reg is a simple L1 loss function.

[0047] A three-dimensional sparse object detection system based on depth repair, including:

[0048] A two-dimensional detection network for converting multi-view images into a set of initial detection boxes for each view;

[0049] A matching augmentation module for deleting redundant and incomplete detection boxes in the set of initial detection boxes, and combining a set of two-dimensional effective noisy detection boxes obtained by a geometric perturbation enhancement strategy from two-dimensional ground truth bounding boxes to generate a two-dimensional detection set for each view and an augmentation record for distinguishing the sources of detection boxes in the two-dimensional detection set;

[0050] A matching consistency module for identifying and processing detection boxes with correct matching relationships among high-quality detection boxes in the two-dimensional detection set to generate a matching sequence; and optimizing the fused features of ROI features and camera intrinsics through the maximum IOU value sequence and the matching sequence based on two-dimensional and three-dimensional parameter space mapping constraints;

[0051] A Query generator for initializing three-dimensional space queries for the optimized fused features and the two-dimensional detection set;

[0052] A three-dimensional feature decoding module for obtaining decoding information;

[0053] A detection head for transforming the decoding information and outputting three-dimensional detection results of the target.

[0054] Compared with the prior art, the three-dimensional sparse object detection method based on depth repair is implemented based on a three-dimensional sparse object detection network based on depth repair. The three-dimensional sparse object detection network based on depth repair adds a proposed matching augmentation module on the basis of the MV2D network to multiply the spatial coverage rate of candidate boxes through a geometric perturbation and effectiveness screening strategy, effectively alleviating the problem of target missed detection in dynamic scenes; and by adding a matching consistency module to construct a joint constraint mechanism for IOU projection, enabling the correct transmission of depth information contained in the two-dimensional space and verifying the feasibility and effectiveness of cross-modal geometric correlation; furthermore, compared with typical BEV detectors and other sparse paradigm detectors, this method not only significantly reduces the computational cost relative to the BEV detector, effectively ensuring the detection rate of the algorithm, but also effectively utilizes the detection priors of two-dimensional images to improve the accuracy of three-dimensional object detection. Brief Description of the Drawings

[0055] Figure 1 It is a flowchart of the three-dimensional sparse object detection method based on depth repair of the present invention;

[0056] Figure 2 It is a structural schematic diagram of the three-dimensional object detection network based on two-dimensional and three-dimensional matching consistency of the present invention;

[0057] Figure 3The processing flow chart of the matching augmentation module of the 3D object detection network based on the consistency of 2D and 3D matching of the present invention;

[0058] Figure 4 The processing flow chart of the matching consistency module of the 3D object detection network based on the consistency of 2D and 3D matching of the present invention;

[0059] Figure 5 The visualization schematic diagram of the output result after the trained 3D object detection network based on the consistency of 2D and 3D matching of the present invention inputs multi-angle images;

[0060] Figure 6 The visualization schematic diagram of the processing results of the trained 3D object detection network based on the consistency of 2D and 3D matching of the present invention for different scenes;

[0061] Figure 7 The bird's-eye view comparison diagram of the trained 3D object detection network based on the consistency of 2D and 3D matching of the present invention and the baseline detector MV2D. Detailed implementation manners

[0062] The following further describes the present invention in conjunction with the accompanying drawings and specific embodiments, but the following embodiments are by no means any limitation to the present invention.

[0063] See Figure 1 , and the specific implementation steps of the 3D sparse object detection method based on depth repair are described as follows.

[0064] Step 1, construct a 3D object detection network based on the consistency of 2D and 3D matching to make full use of the prior information of 2D images while ensuring the efficiency of the sparse object detector.

[0065] See Figure 2 , the 3D object detection network based on the consistency of 2D and 3D matching is improved based on the existing MV2D network, simply referred to as the MV2DM network, which specifically includes a 2D detection network, a 3D feature extraction module, a matching augmentation module (MAM), a matching consistency module (MCM), a Query generator, a 3D feature decoding module, and a detection head; the 2D detection network, the 3D feature extraction module, the Query generator, the 3D feature decoding module, and the detection head are all the original structures of the MV2D network, and the matching augmentation module (MAM) and the matching consistency module (MCM) are new modules.

[0066] The 2D detection network adopts the Faster R-CNN architecture, which includes a backbone network, a feature pyramid network (FPN), and a 2D detection head connected in sequence; specifically, the input image group is denoted as I = {I1, I2,..., I N}, where N represents the number of cameras, and the input image group consists of multi-view images collected by cameras at different installation positions; in the 2D detection network, the input image group first passes through the backbone network and the feature pyramid network to extract multi-scale features F = {F1, F2,..., F N}; then, the multi-scale features F are in the 2D detection head, and after confidence threshold screening and non-maximum suppression, an initial detection box set B init = {B1, B2,..., B N} is obtained, where B i = {b1, b2,..., b mi} represents that the i-th view retains mi detection boxes.

[0067] The input end of the 3D feature extraction module is connected to the output end of the feature pyramid network of the 2D detection network and shares weights with the 2D detection head; the 3D feature extraction module is used to extract the features of all possible 2D detection boxes, and then align and match according to the features of the 2D detection boxes (2D PT) output by the 2D detection head to obtain candidate region features (ROI features), which are respectively input into the Query generator and the 3D feature decoding module.

[0068] In this 3D object detection network, the role of the 2D detection network is to give full play to its advantages in the balance of speed and accuracy, generate high-quality priors for the 3D part, and use the 2D ground truth of spatial mapping for supervised training.

[0069] However, in the 3D object detection task, the effective utilization of context information is crucial for object spatial localization. The 2D detection network in the MV2D network has a two-stage detection mechanism, which shows excellent performance on 2D benchmark datasets such as PASCAL VOC and MS COCO, but still faces significant challenges in 3D scene applications. For example, problems such as complex environmental occlusion, image blur, and high proportion of small targets will all lead to redundant or incomplete candidate box sets output by the 2D detector. Based on this, the present invention further adds a matching augmentation module behind the 2D detection network to solve the above technical problems.

[0070] The input end of the matching augmentation module (Matching Augmentation Module, hereinafter referred to as the MAM module) is connected to the output end of the 2D detection network, so that the processing flow of the MAM module starts from the initial candidate box set B of each angle output by the 2D detector i = {b1, b2,..., b mi}; The role of this MAM module in the 3D object detection network is as follows: applying geometric consistency constraints during the 2D data augmentation process, obtaining high-quality 2D prediction boxes through screening, and performing noise-added fusion in combination with the ground truth to generalize the initial generated depth of the model; at the same time, accelerating the training efficiency of the depth module to improve the robustness in occlusion and small-sample scenarios.

[0071] See Figure 3 , and match the augmentation module to the initial detection box set B at each perspective i The specific processing process is described as follows.

[0072] 1) For the detection boxes in the initial detection box set (i.e., the 2D PT in Figure 3 ), they are successively screened by triple parameter constraints of confidence threshold filtering (thr), non-maximum suppression (nms), and maximum retention number control (maximg) to obtain an effective candidate box set.

[0073] This processing step is used to preliminarily delete the excessive initial candidate boxes in the initial detection box set, and improve the quality of the initial candidate box set by removing redundant or incomplete detection boxes; among them, confidence threshold filtering is used to filter low-quality predictions, non-maximum suppression threshold is used to eliminate redundant detections, and the maximum retention number is used to control the computational complexity.

[0074] However, after this step, the number of remaining candidate boxes is generally less than 100, showing a magnitude difference from the standard query number (300) of the DETR series models and the PETR scheme (900). This sparsity of the candidate space will directly lead to a decrease in the coverage density in the 3D query initialization stage, and further affect the detection recall rate; therefore, it is also necessary to introduce a geometric perturbation enhancement strategy to increase the number of effective candidate boxes by several times while maintaining the original detection accuracy, significantly improving the initial spatial perception ability of the 3D detector.

[0075] 2) Apply controllable perturbations to the 2D ground truth annotation boxes in the dataset (i.e., the 2D GT in Figure 3 ) to construct a 2D noise-added detection box set.

[0076] The specific processing process of step 2) is as follows:

[0077] Make G copies of the 2D ground truth annotation boxes in the dataset to obtain a 2D repeated detection box set (i.e., the 2D GT repeated group in Figure 3 );

[0078] Apply uniform distribution perturbations to the parameters of each repeated detection box according to the preset scaling factor s:

[0079] b i =(c x ,cy , w, h), Δb i = (c x ·Δc x , c y ·Δc y , w·Δw, h·Δh)·s,

[0080] b i ′ = b i + Δb i ,

[0081] where Δc x , Δc y , Δw, Δh are perturbation terms randomly drawn from a uniform distribution:

[0082] Δc x , Δc y , Δw, Δh ~ U(-1, 1),

[0083] Furthermore, a two-dimensional noisy candidate box set (i.e., the 2D noisy set in Figure 3 ) is obtained.

[0084] 3) Adopt a geometric constraint enhancement strategy to post-process the noisy candidate box set to obtain a two-dimensional effective noisy detection box set (i.e., the 2D effective noisy set in Figure 3 ).

[0085] The specific processing procedure of step 3) is as follows:

[0086] 3.1) Crop the image boundary according to the candidate box coordinates to ensure that the detected box after adding noise is completely located within the effective imaging area;

[0087] 3.2) Eliminate invalid candidate boxes that do not match the three-dimensional ground truth spatial position according to the projection error between the two-dimensional ground truth annotation box (2D GT) and the three-dimensional ground truth annotation box (3D GT);

[0088] 3.3) Combine the noisy candidate boxes after being processed in step 3.2) with the filtered candidate box set obtained in step 1) to construct a two-dimensional detection set B MAM , which includes the coordinates of all detected boxes in the set;

[0089] 3.4) Form an augmentation record S MAM according to the source of the detected boxes in the two-dimensional detection set to distinguish whether the detected box comes from the initial candidate box set or the two-dimensional effective noisy detection box set; this augmentation record S MAM is used to construct a dynamic mask matrix during the attention calculation stage to suppress the wrong feature interaction between the noisy samples and the original samples.

[0090] The processing step of step 3) realizes cross-sample batch parallel processing through tensor operations, and the output results are: two-dimensional detection set B MAM and augmented record S MAM .

[0091] This MAM module improves the detector's tolerance to noise by constructing a joint noise and geometry constraint mechanism, thereby enhancing its performance in complex environments.

[0092] However, in 3D object detection, accurate 2D-3D geometric correlation is the basis for establishing spatial perception. Experiments show that the traditional Hungarian matching mechanism has inherent defects in maintaining cross-modal consistency, and the loss function exhibits periodic oscillations during the early and middle stages of training, making it difficult to converge. This instability stems from the dual noise coupling effect. Firstly, the inherent ambiguity of the monocular depth estimation network makes it difficult to accurately predict the 3D centroid depth, and the low-quality 3D initial queries gradually accumulate errors during the decoding process and deviate from the optimal solution. Secondly, the 2D projections of the 3D ground truth mostly contain completely or partially occluded parts (such as people occluded by cars or trucks), while the 2D detection results often only provide the visible foreground parts. This asymmetry in geometric information is exponentially amplified in the attention mechanism of the decoder, triggering gradient conflicts and causing oscillations in the optimization direction. Based on this, this 3D object detection network further adds a matching consistency module to construct an efficient 2D-to-3D binding mechanism based on the matching consistency strategy, and uses the IOU screening matrix and matching consistency sequence to optimize feature fusion to improve the detection accuracy of 3D objects. Finally, combined with the matching consistent 3D supervision, the training convergence is further accelerated, and efficient inference is achieved on a single 3090 GPU.

[0093] The input end of the Matching Consistency Module (hereinafter referred to as the MCM module) is connected to the output end of the matching augmentation module, so that the processing flow of the MCM module starts from the two-dimensional detection set B output by the matching augmentation module MAM and augmented record S MAM . The role of this MCM module in the 3D object detection network is: by establishing the mapping constraint between the 2D and 3D parameter spaces, ensuring that the high-quality 2D detection boxes are geometrically aligned with the corresponding 3D results, so that the depth information contained in the 2D space is correctly transmitted.

[0094] The MCM module aims to enhance the correlation between the candidate set and the 3D geometric ground truth, and its inputs include: the two-dimensional detection set B obtained by the MAM module MAM , as well as the two-dimensional true annotation boxes (2DGT) and the 2D_to_3D projection dictionary (the mapping relationship between the 2D and 3D spaces, M 2t3 ).

[0095] SeeFigure 4 , the specific processing procedure of the matching consistency module is described as follows.

[0096] 1) Through the internal parameter projection and external parameter projection of the 3D ground truth bounding box (3D GT), obtain the multi-view 2D ground truth bounding boxes (2D GT); record the IOU relationship between the 2D ground truth bounding boxes (2D GT) and the corresponding 2D detection boxes in the 2D detection set B MAM to construct the cross-modal IOU correlation matrix M IOU ;

[0097] 2) To select candidate matching detection boxes with higher quality, select the matching serial number corresponding to the maximum IOU value between each detection box and the ground truth box in M IOU , and add the detection boxes with IOU value ≥ the IOU screening threshold IOU match to the pairing set of MCM;

[0098] In this embodiment, the IOU screening threshold IOU match is set to 0.6, that is, if the IOU value ≥ 0.6, the candidate box meets the requirements; in addition, it should be noted that since there are multiple ground truth boxes for the specified object in the dataset, each detection box will correspond to a ground truth box that satisfies the maximum IOU value.

[0099] 3) After determining the matching relationship between the detection boxes and the 2D ground truth boxes in the MCM pairing set, obtain the matching relationship between the detection boxes in the MCM pairing set and the 3D ground truth boxes according to the M 2t3 projection dictionary;

[0100] 4) According to the augmented record S MAM , eliminate the noisy detection boxes in the MCM pairing set that cannot be matched with the 2D ground truth boxes and the 3D ground truth boxes; that is, further select by S MAM so that the detection boxes generated by adding noise to a certain 2D true value must locate and track the original 2D true value target and the transformed 3D true value target;

[0101] 5) Regard the detection boxes not added to the MCM pairing set in step 2) and the detection boxes deleted from the pairing set in step 4) as illegal detection boxes, and uniformly modify their assignments in the augmented record S MAM to -1 as an invalid matching identifier to ensure that they do not participate in subsequent processing;

[0102] The expressions of the above processing steps 1) to 5) are:

[0103]

[0104] In the formula, b pred,i and b pred,i represent the i-th detection box and the j-th ground truth box respectively, IOUmatch represents the IOU screening threshold;

[0105] 6) After the above steps 1) to 5), construct the new sequence after reassigning values to S MAM to obtain the matching sequence S MCM , and the sequence of the maximum IOU values corresponding to the detection boxes (i.e., S IOU ), which are used for subsequent optimization and training.

[0106] 7) Fuse the ROI features and the camera internal parameters, and its expression is:

[0107]

[0108] In the formula, I 2D represents the two-dimensional features after ROI processing, T i represents the camera internal parameters, C represents the feature splicing operation, is the feature processing module, and I 2DF is the processing result of the fused internal parameters.

[0109] 8) In the way of encoding legal detection boxes as 1 and illegal detection boxes as 0, perform 0-1 encoding conversion on S MCM , and then fuse the matching sequence S MCM and the IOU value sequence S IOU , and its expression is:

[0110]

[0111] In the formula, represents the conversion module for 0-1 encoding of S MCM , represents a simple linear fusion conversion matrix, and I DTHI is the optimized fused feature, that is, the final output result of the MCM module.

[0112] The optimized fused feature output by the MCM module and the two-dimensional detection set output by the MAM module are input into the Query generator, and then both are input into the three-dimensional feature decoding module.

[0113] The three-dimensional feature decoding module still adopts the interactive feature fusion architecture of the MV2D network. By introducing the three-dimensional position encoding (3D PE) of the embedded two-dimensional features, the position perception of the image features is enhanced; through the camera frustum space coordinate transformation, a three-dimensional space mask matrix is constructed, and combined with the deformable cross-attention mechanism, the spatial constraint in the feature interaction process is realized to obtain the decoded information.

[0114] The decoded information output by the three-dimensional feature decoding module is input into the detection head. After dimensional transformation, the corresponding three-dimensional detection results are finally output, including the three-dimensional position (x, y, z), size (w, l, h), angle (yaw), and speed information (v x , v y ).

[0115] Step 2: Substitute the dataset into the three-dimensional object detection network constructed in Step 1 for training, and apply the matching loss supervision based on S MCM to maintain the two-stage joint optimization objective of this network.

[0116] It should be noted that although the detector obtained by MAM data augmentation has been subjected to noise perturbation, it is a result that is close to the two-dimensional ground truth without learning cost. Before the detection set is sent to the decoder (after the query generator), we need to save the noise-added part according to S MAM for backup to ensure subsequent supervision services. After that, the two-dimensional noise-added augmentation part belonging to the three-dimensional query will be deleted.

[0117] In this Step 2, in order to maintain the supervision system of the two-stage joint optimization objective of this three-dimensional object detection network, during the network training process, it is necessary to introduce geometric consistency constraints on the original basis, and construct matching loss supervision to reduce the matching oscillation between two dimensions and three dimensions and accelerate the convergence speed of the model.

[0118] Specifically, the matching loss function has the following expression:

[0119]

[0120] In the formula, represents the matching consistency loss between two dimensions and three dimensions, and λ 3d is the weight coefficient of the three-dimensional part, is the original MV2D loss, is the two-dimensional detection loss part, is the three-dimensional detection loss part.

[0121] Among them, has the same weight as the three-dimensional loss part of MV2D. Usually, the ground truth correspondence is matched through Hungarian matching in the three-dimensional part. However, in the method based on the matching consistency between two dimensions and three dimensions, it is forced to bind according to the consistency matching sequence S MCM , and this method discards the focal-loss loss corresponding to the label part to further emphasize the importance of the two-dimensional to three-dimensional localization conversion.

[0122] Specifically, has the following expression:

[0123]

[0124] In the formula, the σ function represents according to S MCM to perform three-dimensional true value matching, L reg is a simple L1 loss function.

[0125] Furthermore, by using a trained three-dimensional object detection network to process multi-angle images, three-dimensional sparse object detection based on depth repair can be achieved. This method can not only retain the efficiency of the three-dimensional sparse detector, but also effectively utilize the detection prior of two-dimensional images to improve the accuracy of the three-dimensional object detector.

[0126] Such as Figure 5 shown, when the user needs to upload the captured panoramic image and upload it to the trained three-dimensional object detection network according to the specific image angle order, the corresponding three-dimensional detection result can be output. After passing through the visualization software, the three-dimensional detection result can display the object annotation effect diagram or the bird's-eye view effect diagram as shown.

[0127] In practical applications, the three-dimensional detection result interacts with downstream front-end web components or electronic sensing elements, so as to achieve efficient object detection in a three-dimensional scene and be applied to the non-end-to-end algorithm processes of entertainment, social products, and autonomous driving technologies.

[0128] According to the matching consistency strategy we designed, using the MV2D network as the basic detector, the performance of the three-dimensional object detection network of the present invention is evaluated.

[0129] The method of the present invention (hereinafter simply referred to as MV2DM) and other advanced algorithms on the nuScenes [1] The comparative evaluation results on the validation set are shown in Table 1 below.

[0130] Table 1:

[0131]

[0132] It can be seen from the results in Table 1 that on the widely used three-dimensional object detection dataset nuScenes, the method of the present invention has obvious accuracy advantages in two evaluation indicators of the common three-dimensional object detection indicators NDS (nuScenes Detection Score) and mAP. Among them, NDS is composed of mAP and five true positive indicators of mATE, mASE, mAOE, mAVE, and mAAE. In this evaluation experiment, the noise addition amplitude of the augmentation part of the matching augmentation module in the MV2DM network of the present invention is set to 0.1, and the number of noise addition groups is set to 10; in addition, the IOU matching threshold in the matching consistency module is set to 0.6.

[0133] Specifically, in Table 1, the optimal results of each column are in bold. The representative model is initialized from the FCOS3D backbone. The representative model is pre-trained on nuImages, and * indicates that the model uses additional modal data (such as point clouds) during training. It can be seen that the accuracy of our algorithm is far ahead of other algorithms. For fairness, all algorithms are implemented in our same hardware environment, and the graphics processor used for inference is the RTX 3090 GPU.

[0134] In addition, to verify the effectiveness of the proposed module, we also conducted a comparative study on the improvement of the model accuracy by two newly added modules. The specific results are shown in Table 2 below.

[0135] Table 2:

[0136]

[0137] To more intuitively represent the role of the proposed algorithm, a visual comparison experiment with the baseline model MV2D is used to visually demonstrate the performance advantages of the improved model of the method of the present invention; among them, each example includes "2D GT", "3D GT", the three-dimensional results of MV2D and the method of the present invention, specifically as Figure 6 shown.

[0138] To show the robustness of the model to different scenarios, the following shows three groups of representative results; for easy observation, the top 20 results of the three-dimensional part detection results are extracted and projected onto each perspective, and ★ is used on the graph to more clearly show the comparison results.

[0139] Case (1) shows that in an occlusion scenario, when the truck is partially occluded by the vehicle in front, the detection result of the baseline model MV2D is aligned with the 2D annotation box, but it can only capture the visible part of the target, while the method of the present invention restores the three-dimensional structure of the occluded target completely through geometric consistency constraints, and at the same time successfully detects the obstacles missed by MV2D.

[0140] Case (2) shows that for medium-distance targets, the method of the present invention shows more accurate angle prediction ability, and the spatial orientation of its three-dimensional bounding box is highly consistent with the real road direction.

[0141] Case (3) shows that in the scenario of long-distance multi-target stacking, the method of the present invention not only significantly improves the detection recall rate of small targets, but also maintains the accuracy of three-dimensional positioning, effectively distinguishing multiple objects with similar spatial positions. The visualization results show that the matching consistency strategy we introduced enables the model to more completely reconstruct the three-dimensional structure of the target in complex scenarios and improves the perception ability of the target's spatial attitude and dense distribution by strengthening the geometric correlation constraints between two-dimensional and three-dimensional.

[0142] In addition, as Figure 7 shown, we also compare the method of the present invention with the BEV (bird's-eye view) of the baseline detector MV2D, intuitively reflecting the significant advantages of the improved method of the present invention in depth perception and direction angle prediction; among them, the green box (GT) in the figure represents the ground truth of the three-dimensional target, and the blue box (PT) represents the three-dimensional detection result. For easy observation, the results with a confidence level exceeding 0.15 are extracted from the three-dimensional detection results for comparison. It can be observed from the figure that in multiple scenarios, the detection results of the method of the present invention are closer to the ground truth of the three-dimensional target, and at the same time, the direction perception of the target is more accurate.

[0143] References:

[0144] [1] Caesar H, Bankiti V, Lang A H, et al. nuScenes: A multimodal dataset for autonomous driving[C]. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, 2020: 11621-11631.

[0145] [2] Wang Y, Guizilini V C, Zhang T, et al. Detr3d: 3D object detection from multi-view images via 3D-to-2D queries[C]. In Conference on Robot Learning, 2022: 180-191.

[0146] [3]Li Y, Ge Z, Yu G, et al. Bevdepth: Acquisition of reliable depth for multi-view 3d object detection[C]. In Proceedings of the AAAI Conference on Artificial Intelligence, 2023: 1477-1485.

[0147] [4]Liu Y, Wang T, Zhang X, et al. Petr: Position embedding transformation for multi-view 3d object detection[C]. In European Conference on Computer Vision, 2022: 531-548.

[0148] [5]Wang Z, Huang Z, Fu J, et al. Object as query: Lifting any 2d object detector to 3d detection[C]. In Proceedings of the IEEE / CVF International Conference on Computer Vision, 2023: 3791-3800.

[0149] [6]Wang T, Zhu X, Pang J, et al. Fcos3d: Fully convolutional one-stage monocular 3d object detection[C]. In Proceedings of the IEEE / CVF International Conference on Computer Vision, 2021: 913-922.

[0150] [7]Wang T, Xinge Z, Pang J, et al. Probabilistic and geometric depth: Detecting objects in perspective[C]. In Conference on Robot Learning, 2022: 1475-1485.

[0151] [8]Li Z, Wang W, Li H, et al. BEVFormer: Learning Bird’s-Eye-View Representation From LiDAR-Camera via Spatiotemporal Transformers[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, 47(3): 2020-2036.

Claims

1. A three-dimensional sparse object detection method based on depth repair, characterized in that the steps Including: S1. Obtain multi-view images through cameras set at different viewing positions; S2. Input the multi-view images into the 2D detection network of the 3D sparse object detection network to obtain an initial detection box set for each view; S3. Input the initial detection box sets for each view into the matching augmentation module of the 3D sparse object detection network. By deleting redundant and incomplete detection boxes in the initial detection box sets and combining with the 2D effective noisy detection box set obtained by the geometric perturbation enhancement strategy for 2D ground truth bounding boxes, generate the 2D detection sets for each view and the augmentation records for distinguishing the sources of detection boxes in the 2D detection sets; S4. Input the initial detection box sets for each view into the matching consistency module of the 3D sparse object detection network to screen out high-quality detection boxes with correct matching relationships to generate a matching sequence; Based on the 2D and 3D parameter space mapping constraints, optimize the fused features of the ROI features and the camera intrinsics with the maximum IOU value sequence and the matching sequence; S5. After initializing the 3D space query of the optimized fused features obtained in step S4 and the 2D detection sets obtained in step S3 by the Query generator, transmit them to the 3D feature decoding module to obtain decoding information.

2. The three-dimensional sparse object detection method based on depth repair according to claim 1, characterized in that In step S3, the method for deleting redundant and incomplete detection boxes in the initial detection box set is: successively perform triple parameter constraint screening of confidence threshold filtering, non-maximum suppression, and maximum retention quantity control on the initial detection box set to obtain an effective candidate box set.

3. The three-dimensional sparse object detection method based on depth repair according to claim 1, characterized in that In step S3, the method for generating the 2D effective noisy detection box set is: Step 1. Make G copies of the 2D ground truth bounding boxes in the dataset and perform uniform distribution perturbation on the parameters of each candidate box according to the preset scaling factor s, and its expression is: b i = (c x , c y , w, h), Δb i = (c x ·Δc x , c y ·Δc y , w·Δw, h·Δh)·s, b′ i = b i + Δb i , where Δc x , Δc y , Δw, Δh are perturbation terms randomly sampled from a uniform distribution: Δc x , Δc y , Δw, Δh ∼ U(-1, 1); furthermore, a set of two-dimensional noisy candidate bounding boxes is obtained. Step 2. Adopt the geometric constraint enhancement strategy to successively perform image boundary cropping on the noisy candidate box set and remove the invalid candidate boxes that do not match the 3D ground truth spatial position.

4. The three-dimensional sparse target detection method based on depth repair according to claim 1, wherein In step S4, the recognition and processing process of low-quality detection boxes and detection boxes with incorrect matching relationships in the 2D detection set is: 1) Construct a cross-modal IOU correlation matrix according to the intersection over union (IOU) relationship between the 2D ground truth bounding boxes and the corresponding 2D detection boxes in the 2D detection set; 2) Select the matching serial number corresponding to the maximum IOU value of each detection box and the ground truth box in the correlation matrix, and add the detection boxes with IOU value ≥ IOU screening threshold to the pairing set; 3) According to M 2t3 Projection dictionary, determine the matching relationship between the detection box and the three-dimensional ground truth box in the pairing set; 4) According to the augmentation record, remove the noisy detection boxes in the MCM pairing set that cannot be matched with the 2D ground truth box and the 3D ground truth box; 5) Consider the detection boxes that are not added to the pairing set and the detection boxes deleted from the pairing set as illegal detection boxes, and uniformly modify their assignments in the augmentation record S MAM to negative values to ensure that illegal detection boxes do not participate in subsequent processing; 6) Construct for S MAM The new sequence after re - assignment is used as the matching sequence S MCM , and the sequence S of the maximum IOU value corresponding to each detection box IOU .

5. The three-dimensional sparse target detection method based on depth repair according to claim 4, characterized in that, In step S4, the processing process of optimizing the fused features of the ROI features and the camera intrinsics with the maximum IOU value sequence and the matching sequence based on the 2D and 3D parameter space mapping constraints is: 1) Fuse the ROI features and the camera intrinsics, and its expression is: Where, I 2D represents the two-dimensional feature after ROI processing, T i represents the camera internal parameters, C represents the feature stitching operation, is the feature processing module, I 2DF is the processing result of the fused internal parameters. 2) Perform 0-1 coding conversion on the matching sequence S MCM to distinguish between legal detection boxes and illegal detection boxes; and then fuse the matching sequence S MCM and the IOU value sequence S IOU , and its expression is: In the formula, represents S MCM a conversion module for 0-1 coding, represents a simple linear fusion conversion matrix, I DTHI is the optimized fusion feature.

6. The three-dimensional sparse object detection method based on depth repair according to claim 1, wherein It further includes step S6. Input the decoding information into the detection head to output the 3D position, size, angle, and speed information of the target.

7. A training method for a three-dimensional sparse object detection network based on depth repair, characterized in that, The training method includes: inputting a multi-view image set, the internal and external parameters of the corresponding camera, and a 2D_to_3D projection dictionary into a three-dimensional sparse object detection network for training, and applying matching loss supervision based on the matching sequence during the training process; Matching loss function The expression is as follows: In the formula, represents the matching consistency loss between two dimensions and three dimensions, and λ 3d is the weight coefficient of the three-dimensional part, is the original MV2D loss, is the two-dimensional detection loss part, is the three-dimensional detection loss part; The expression is: where the σ function represents three-dimensional truth matching based on the matching sequence, and L reg is a simple L1 loss function.

8. A three-dimensional sparse object detection system based on depth repair, characterized in that, It includes: A two-dimensional detection network, which is used to convert multi-view images into a set of initial detection boxes for each view; A matching augmentation module, which is used to delete redundant and incomplete detection boxes in the set of initial detection boxes, and combine a set of two-dimensional effective noisy detection boxes obtained by a geometric perturbation enhancement strategy from two-dimensional ground truth annotation boxes to generate a two-dimensional detection set for each view and an augmentation record for distinguishing the source of the detection boxes in the two-dimensional detection set; A matching consistency module, which is used to identify and process low-quality detection boxes and detection boxes with incorrect matching relationships in the two-dimensional detection set to generate a matching sequence; and optimize the fused feature of the ROI feature and the camera internal parameter through the maximum IOU value sequence and the matching sequence based on the two-dimensional and three-dimensional parameter space mapping constraints; A Query generator, which is used to initialize the three-dimensional space query for the optimized fused feature and the two-dimensional detection set; A three-dimensional feature decoding module, which is used to obtain decoding information; A detection head, which is used to transform the decoding information and output the three-dimensional detection results of the target.

Citation Information

Cited By

  • Three-dimensional target detection method based on foreground feature extraction

    CN121033558A

  • Three-dimensional target detection method and device based on language reasoning enhancement

    CN121438297A