Roadside long-distance three-dimensional target detection method based on prior geometric information guidance
By using technical means such as vanishing point cropping, three-dimensional query generation and depth range position embedding in long-distance three-dimensional target detection on the roadside, the problems of feature imbalance, difficulty in depth estimation and difficulty in matching three-dimensional query and feature in long-distance target detection are solved, and high-precision long-distance target detection is achieved.
Patent Information
- Application Number
- CN202510166985.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art faces the problems of feature imbalance, difficulty in depth estimation and difficulty in matching three-dimensional query and feature in long-distance object detection, resulting in insufficient detection accuracy and efficiency.
The roadside long-distance three-dimensional object detection method is used based on prior geometric information guidance, and accurate 3D query is generated through technical means such as vanishing point cropping, three-dimensional query generation and depth range position embedding, and the detection accuracy of long-distance small targets is improved through multi-view clipping strategies and spatial attention mechanisms.
Maintain high detection accuracy within the range of 0-200m, effectively improve the detectability of long-distance targets, and achieve accurate detection and positioning of multiple categories of targets on the roadside.
Smart Images

Figure CN120107902A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision perception technology for autonomous driving and intelligent transportation, and in particular to a roadside long-distance three-dimensional target detection method guided by priori geometric information. Background Art
[0002] With the continuous development of intelligent transportation and vehicle-road cooperative technology, the use of roadside perception equipment (such as roadside cameras) to monitor long-distance traffic targets has become an important demand. Traditional monocular 3D target detection research focuses on vehicle-mounted camera scenarios, and the detection distance is usually only in the range of 0-100m. However, roadside cameras have the characteristics of wide field of view and long coverage distance, and can be used to detect targets at farther distances (for example, 100m-200m or even farther). In this scenario, existing methods face several challenges:
[0003] 1. Feature imbalance caused by the increase in field of view: At long distances, targets are often extremely small and blurry in the image, making feature extraction more difficult; while nearby targets are clearer, resulting in insufficient representation of distant targets by the model.
[0004] 2. Depth estimation becomes more difficult: The pixel difference of distant targets on the image plane is extremely small. The traditional monocular depth estimation error will be magnified at a distance, which has a significant impact on the 3D detection accuracy.
[0005] 3. Difficulty in matching 3D queries with features: For feature-based BEV methods, long-distance sparse and deformed features can lead to inaccurate feature maps. For methods based on sparse queries, if 3D queries are generated randomly or uniformly over a large range, it is easy to make it difficult to converge during the training phase, or difficult to effectively align the location of the real target. Summary of the invention
[0006] In order to overcome the shortcomings of existing methods in terms of accuracy and efficiency in long-distance target detection, the purpose of the present invention is to provide a roadside long-distance three-dimensional target detection method (PGDetection) guided by prior geometric information, which uses prior geometry (such as camera extrinsics, ground priors, depth priors, etc.) to generate accurate 3D queries, and improves the detection accuracy of small targets at a distance through a multi-view cropping strategy and spatial attention mechanism adapted to long distances. By introducing a vanishing point-based image cropping module, a three-dimensional query generation module, and a depth range position embedding module, the present invention can maintain a high detection accuracy within the range of 0-200m, and effectively improve the detectability of long-distance targets (such as vehicles, pedestrians, cyclists, etc.).
[0007] The technical solution of the present invention is specifically described as follows.
[0008] The present invention provides a roadside long-distance three-dimensional target detection method based on prior geometric information guidance, comprising the following steps:
[0009] Step S1: Vanishing Point Cropping
[0010] After acquiring the roadside video or image frame, at least one vanishing point is determined by detecting the intersection of the road extension direction lines; a cropping area is delineated around the vanishing point and enlarged or affine transformed to obtain a cropped image I crop ; The original image I orig With cropped image I crop The subsequent networks are fed in parallel to enhance the feature recognition of distant target areas while ensuring the overall field of view;
[0011] Step S2: 3D Query Generator
[0012] Using a two-dimensional detector or existing prior information, obtain the lower edge center point (u, v) corresponding to the target on the image plane as the reference point in contact with the ground; combine the camera internal and external parameter matrix and the ground plane equation to project the reference point into the three-dimensional world coordinate system to obtain the initial three-dimensional coordinates (x w ,y w , z w ); During the training process, each query is aligned in three-dimensional space according to the distance from the reference point to form a high-quality three-dimensional query Q 3D , providing prior constraints for subsequent decoding;
[0013] Step S3: Range Positional Embedding
[0014] The depth prediction module discretizes the depth of the input image or ROI area to obtain a series of depth candidate values {d n} and its probability distribution {P n}, adaptively shrink or expand the most likely depth range according to the uncertainty; associate each pixel (u, v) with the depth d n Back-projecting to a three-dimensional coordinate system to form a viewing cone with depth prior; and superimposing the depth probability distribution and the three-dimensional position encoding in the feature map to obtain an embedded representation with both semantic and geometric information;
[0015] Step S4: spatial Transformer decoder;
[0016] The multi-view image obtained in step S1, i.e., the original image I orig With cropped image Icrop After the features are extracted by the convolutional backbone network, they are combined with the deep embedding in step S3 as the key and value of the decoder; the three-dimensional query Q in step S2 is 3D It is input into the decoder as a query, and predicts the category and bounding box of the target in three-dimensional space through multi-layer self-attention and cross-attention modules. The spatial attention mask is introduced to make the query only focus on the reference coordinate point closest to itself, so as to improve the focus on distant targets.
[0017] In the present invention, in step S1, when determining the vanishing point, if multiple road lines with different directions are included, the intersection average of the multiple straight lines is taken or the least squares method is used for iterative fitting to obtain the precise vanishing point.
[0018] In the present invention, in step S1, the cropped area is enlarged and translated using an affine transformation matrix, and the high resolution is preferentially retained for the area near the vanishing point to enhance the detectability of distant targets.
[0019] In the present invention, in step S2, the projection process of the three-dimensional query generation satisfies the following relationship:
[0020]
[0021] Where K is the camera internal parameter, T is the external parameter, d i is the constant term in the ground plane equation, λ is the depth scaling factor; (x′, y′) is the normalized direction vector, (u, v) is the image coordinate, (x w ,y w , z w ) is the three-dimensional reference point in the world coordinate system.
[0022] In the present invention, in step S3, the depth range position embedding includes the following sub-steps:
[0023] For depth candidate Perform iterative updates and statistically predict the mean value U p and variance U D , making U D As the depth range for the next update;
[0024] The ROI features of the region of interest are divided into multiple scales, and each pixel (u, v) is associated with a discrete depth d n After back-projection to three-dimensional space, use the three-dimensional coordinates (x n ,y n , z n ) to do position encoding and convert the depth probability P n Element-wise fusion with feature maps;
[0025] The fused features are converted into Key and Value forms that can be used by the Transformer decoder using a multi-layer perceptron (MLP) or convolution method.
[0026] In the present invention, in step S4, the Transformer decoder performs a three-dimensional query Q 3D When performing multi-head attention operations with Key and Value, an attention mask based on three-dimensional Euclidean distance is introduced, which only allows high-weight matching of the query with several of its closest three-dimensional reference points to reduce mismatches caused by sparse long-distance features.
[0027] In the present invention, in step S4, the network training includes two losses: 2D detection and 3D detection. The 2D loss uses cross entropy or Focal Loss to constrain target classification. The 3D loss uses the Hungarian matching algorithm to perform a one-to-one correspondence between the prediction box and the true value box, and the successfully matched target is subjected to L1 regression loss and Focal Loss classification loss to ensure that the three-dimensional bounding box and category are accurately regressed.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] The present invention introduces key modules such as vanishing point cropping, 3D query generation, and range positional embedding, which can effectively deal with problems such as sparse target pixels and unstable depth estimation in long-distance scenes, and achieve accurate detection and positioning of multiple categories of targets (vehicles, pedestrians, cyclists, etc.) on the road side at long distances.
[0030] The long-distance three-dimensional target detection method for roadside scenes of the present invention utilizes prior geometric information to guide the three-dimensional target detection so as to maintain a high detection accuracy at a long distance.
[0031] The present invention can be widely used in fields such as intelligent transportation infrastructure perception, vehicle-road collaboration and autonomous driving, and can improve the intelligence and reliability of road monitoring and safety protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 The present invention is a flow chart of a road-side long-distance three-dimensional target detection algorithm guided by prior geometric information.
[0033] Figure 2 This is a comparison table of the latest results of the embodiments of the present invention in the range of 0-100 meters with the DAIR-V2X-1 verification set.
[0034] Figure 3This is a visualization of the impact of the vanishing point cropping module according to an embodiment of the present invention.
[0035] Figure 4 This is a comparison table of the latest results of the embodiments of the present invention in the range of 0-200 meters with the DAIR-V2X-1 verification set.
[0036] Figure 5 This is a 3D query generator ablation study on DAIR-V2X-I according to an embodiment of the present invention.
[0037] Figure 6 This is the loss curve during model training without using a 3D query generator in an embodiment of the present invention.
[0038] Figure 7 This is a study on vanishing point cropping and ablation based on DAIR-V2X-I according to an embodiment of the present invention.
[0039] Figure 8 This is a study on range position embedding ablation of an embodiment of the present invention on DAIR-V2X-I. DETAILED DESCRIPTION
[0040] The present invention is further described below by way of embodiments in conjunction with the accompanying drawings.
[0041] The process of the roadside long-distance three-dimensional target detection method based on prior geometric information guidance of the present invention is shown in Figure 1 .
[0042] Step S1: Vanishing Point Cropping
[0043] In long-distance scenes, the target often occupies only a very small pixel area in the image, and the features are blurred and easy to be ignored. Therefore, this step analyzes the convergence direction of the road in the image plane, finds the vanishing point and crops and enlarges the area, including:
[0044] 1. Determine the vanishing point
[0045] Sample or detect nearly parallel lines such as road edges and lane lines, and use their intersection points on the image plane to estimate the vanishing point position (x v ,y v ). If there are multiple road lines with different directions, the precise vanishing point can be obtained by taking the average of the intersection of multiple straight lines or using the least squares method to iteratively fit.
[0046] 2. Cropping and affine transformation
[0047] The cropping area is defined around the vanishing point area (W, H). If the coordinates of the upper left corner of the cropping area are (x min ,y min), the coordinates of the lower right corner are (x max ,y max ), the camera’s intrinsic parameter matrix is represented by K v :
[0048]
[0049] For the region of interest (RoI), its equivalent camera intrinsic parameters can be expressed as:
[0050]
[0051] in This operation magnifies distant areas while keeping the overall structure continuous.
[0052] 3. Output “pseudo multi-view” images
[0053] The resulting cropped and enlarged image I crop With the original image I orig The subsequent networks are fed in parallel to enhance the visual resolution of distant objects while preserving the global context.
[0054] Step S2: 3D Query Generator
[0055] In traditional monocular 3D detection, if the entire three-dimensional space is uniformly sampled to generate queries, it is easy to cause the search range to be too large and it is difficult to effectively focus on the target. This step uses the two-dimensional detection box or prior key point as the ground center point and projects it to the three-dimensional world coordinate system to generate a high-quality 3D query. The main process is as follows:
[0056] 1. Ground center point extraction
[0057] Perform two-dimensional detection or prior box filtering on the input image to extract the lower edge center (u, v) of the target in the image coordinate system, which approximately represents the reference point where the object contacts the ground.
[0058] 2. Projection to a 3D coordinate system
[0059] Assuming that the internal and external parameters of the camera are the internal parameter matrix K and the external parameter matrix T respectively, the ground plane equation can be expressed as:
[0060] a i X+b i Y+c i Z+d i =0
[0061] The following formula can be used to complete the 2D to 3D projection:
[0062]
[0063] Where K is the camera internal parameter, T is the external parameter, d i is the constant term in the ground plane equation, λ is the depth scaling factor; (x′, y) is the normalized direction vector, (u, v) is the image coordinate, (x w ,y w , z w ) is the three-dimensional reference point in the world coordinate system.
[0064] 3. 3D Query Alignment
[0065] During network training, a randomly initialized 3D query position (Q x , Q y , Q z ) and the projection result (x w ,y w , z w ) to match and fine-tune the query point and the target’s true position to maintain a high degree of consistency, thereby reducing the “invalid search” range.
[0066] Step S3: Range Positional Embedding
[0067] Using the prior interval obtained by the above 3D query, this step uses the PGDepth depth prediction network based on this prior, iteratively scales or expands the possible depth interval, and performs position encoding fusion in the feature space. It can be specifically divided into:
[0068] 1. Depth Range Generator
[0069] Assume that the network is discretized to obtain N depth candidates And output the corresponding probability P n The discrete depth mean and variance can be made as follows:
[0070]
[0071] And U D Used as the depth interval for the next iteration to adaptively determine the upper and lower bounds of the near and far depths.
[0072] 2. Create Frustum
[0073] The ROI area is divided into uniformly sized grids (u, v) on the image plane. For each (u, v) with depth d n Combined into 3D points (x n ,y n , z n), it can be reversely projected to the world coordinate system to form a three-dimensional viewing cone.
[0074] 3. Depth Aware Embedding
[0075] The coordinates of the above 3D point (x n ,y n , z n ) is used for position encoding, fused with the ROI features of the corresponding grid, and converted into Key Pos and Key / Value features for use by Transformer through MLP, so that both image semantics and depth priors are available in the subsequent network reasoning stage.
[0076] Step S4: Spatial Decoder and Network Training
[0077] After completing cropping, magnification, 3D query and depth embedding, the features and query are handed over to the Transformer decoder for 3D object detection training. It mainly includes the following sub-steps:
[0078] 1. Decoding structure
[0079] A multi-layer Transformer decoder is used, the 3D Query in step S2 is used as the Query vector, and the features in step S3 (including Key Pos and Depth Aware features) are used as Key and Value. The category and 3D bounding box of the target are estimated through self-attention and cross-attention operations.
[0080] 2. Spatial Attention Mask
[0081] The attention mask M is constructed according to the Euclidean distance between the 3D Query in three-dimensional space and the ground reference point. During cross-attention, each query is matched with only several of its most adjacent positions with high weights to reduce the mismatch caused by sparse long-distance features.
[0082] 3. Loss Function
[0083] Network training includes two losses: 2D detection and 3D detection:
[0084] 2D loss: Cross entropy or Focal Loss can be used to constrain target classification;
[0085] 3D loss: The Hungarian matching algorithm is used to perform a one-to-one correspondence between the predicted box and the true value box. The successfully matched targets are subjected to L1 regression loss and Focal Loss classification loss to ensure that the 3D bounding box and category are accurately regressed.
[0086] Embodiment:
[0087] The experimental environment is as follows: the hardware platform uses a server equipped with 2RTX-4090GPUs. The experiment uses the Adamw optimizer for 60 training runs, with an initial learning rate of 2×10 -4 ,All experiments are performed on 2 RTX-4090 GPUs.
[0088] Comparative experiment: In order to comprehensively evaluate the actual performance of the present invention in roadside 3D target detection, the present invention is compared with a variety of existing algorithms in two different detection ranges of 0-100 meters and 0-200 meters on the DAIR-V2X-I dataset, including traditional lidar-based PointPillars, SECOND, MVXNet, Imvoxenet, and multi-view / monocular detection methods M3D-RPN, MonoDLE, MonoDETR, BEVFormer, BEVDepth, BEVHeight, etc. The detection results are shown in the figure below. Figure 2 As shown in the figure, compared with mainstream visual detection methods such as BEVHeight, BEVDepth, and MonoDETR in the range of 0-100m, the average precision (AP) of the present invention for three types of targets, vehicles, pedestrians, and cyclists, has achieved higher scores, among which the AP scores are improved by about 0.34, 0.45, and 0.43 percentage points respectively at the three difficulty levels of easy / medium / difficult (the values are based on vehicles as an example). In particular, in order to evaluate the effectiveness of the present invention in long-distance scenarios (within 200 meters), the BEVHeight method was retrained and tested; the results are shown in the figure below. Figure 4 Compared with BEVHeight, PGDetection has achieved significant improvement in vehicle categories, with AP increases of approximately 1.03, 1.82, and 1.87 points in the easy, medium, and hard difficulty levels, respectively.
[0089] Ablation experiment:
[0090] 3D Query Generator: Randomly generated 3D queries (i.e., without using 3D prior projection information) are compared with the solution using 3D Query Generator. Figure 5 and Figure 6 The results show that if the 3D query generation module is removed, the network will have serious missed detection and training non-convergence at long distances, and the detection accuracy will almost completely fail; after using 3DQueryGenerator, the average precision (AP) of vehicles, pedestrians, and cyclists at different levels of difficulty, such as easy, medium, and difficult, has been significantly improved. The analysis shows that 3D Query Generator can effectively map the center point of the lower edge of the two-dimensional detection box and the ground plane to the three-dimensional world coordinate system, making it easier to align the query with the real target position in the high-dimensional sparse space, significantly reducing the scope of "invalid search".
[0091] Range Positional Embedding: This module consists of three parts: Depth Range Generator, Create Frustum, and Depth Aware Embedding. Figure 8 As shown in the figure, removing any of the submodules will result in a significant drop in the average precision (AP). Specifically, if the ROI feature lacks depth information (Depth Aware Embedding is removed), the detection accuracy of vehicles, pedestrians, and cyclists at three levels of difficulty is greatly reduced; if a simple 3D position embedding is used instead of Create Frustum, the performance drops by about 7 AP points compared to the cone construction method. Experiments show that Range Positional Embedding can organically combine image semantics with depth estimation results, and better locate and distinguish targets at different distances in three-dimensional space, providing important support for long-distance detection.
[0092] Vanishing Point Cropping: In the long-distance scenario of 0 to 200 m, a comparative experiment was conducted on whether to use Vanishing Point Cropping. Figure 7 After enabling vanishing point cropping, the average accuracy of vehicles, pedestrians, and cyclists at the three difficulty levels of easy / medium / difficult is significantly improved; in the visualization results, such as Figure 3 It can also be observed that vanishing point cropping can detect distant objects that were missed by the original method. This shows that Vanishing Point Cropping significantly improves the network's attention to distant small objects by magnifying the sparse pixel areas at a distance, providing a more stable detection basis for long-distance 3D perception.
Claims
1. A roadside long-distance three-dimensional target detection method based on prior geometric information guidance, characterized in that: The specific steps are as follows: S1: Image cropping based on vanishing points After acquiring the roadside video or image frame, determining at least one vanishing point by detecting the intersection of the road extension direction lines; A cropping region is defined around the vanishing point and enlarged or affine transformed to obtain a cropped image I crop ; The original image I orig With cropped image I crop The subsequent networks are fed in parallel to enhance the feature recognition of distant target areas while ensuring the overall field of view; Step S2: 3D query generation Using a two-dimensional detector or existing prior information, obtain the lower edge center point (u, v) corresponding to the target on the image plane as the reference point in contact with the ground; combine the camera internal and external parameter matrix and the ground plane equation to project the reference point into the three-dimensional world coordinate system to obtain the initial three-dimensional coordinates (x w ,y w , z w ); During the training process, each query is aligned in three-dimensional space according to the distance from the reference point to form a high-quality three-dimensional query Q 3D , providing prior constraints for subsequent decoding; Step S3: Depth Range Position Embedding The depth prediction module discretizes the depth of the input image or the region of interest (ROI) to obtain a series of depth candidate values {d n } and its probability distribution {P n }, adaptively shrink or expand the most likely depth range according to the uncertainty; associate each pixel (u, v) with the depth d n Back-projecting to a three-dimensional coordinate system to form a viewing cone with depth prior; and superimposing the depth probability distribution and the three-dimensional position encoding in the feature map to obtain an embedded representation with both semantic and geometric information; Step S4: Spatial Transformer Decoder The original image I obtained in step S1 orig and cropped image I crop The multi-view image is extracted through the convolutional backbone network and combined with the deep embedding in step S3 as the key and value of the decoder; the three-dimensional query Q in step S2 is 3D It is input as a query into the decoder, and predicts the category and bounding box of the target in three-dimensional space through multiple layers of self-attention and cross-attention modules; A spatial attention mask is introduced to make the query focus only on the reference coordinate points closest to itself, so as to improve the focus of distant targets.
2. The method for detecting long-distance three-dimensional targets on the road side based on guidance of prior geometric information according to claim 1, characterized in that: In step S1, when determining the vanishing point, if multiple road lines with different directions are included, the intersection average of the multiple straight lines is taken or the least squares method is used for iterative fitting to obtain the precise vanishing point.
3. The method for detecting long-distance three-dimensional targets on the road side based on guidance of prior geometric information according to claim 1, characterized in that: In step S1, the cropped area is enlarged and translated using an affine transformation matrix, and the high resolution is preferentially retained for the area near the vanishing point to enhance the detectability of distant targets.
4. The method for detecting long-distance three-dimensional targets on the road side based on guidance of prior geometric information according to claim 1, characterized in that: In step S2, the projection process generated by the three-dimensional query satisfies the following relationship: Where K is the camera internal parameter, T is the external parameter, d i is the constant term in the ground plane equation, λ is the depth scaling factor; (x′, y′) is the normalized direction vector, (u, v) is the image coordinate, (x w ,y w z w ) is the three-dimensional reference point in the world coordinate system.
5. The method for detecting long-distance three-dimensional targets on the road side based on guidance of prior geometric information according to claim 1, characterized in that: In step S3, the depth range position embedding includes the following sub-steps: For depth candidate Perform iterative updates and statistically predict the mean value D p and variance U D , making U D As the depth range for the next update; The ROI features of the region of interest are divided into multiple scales, and each pixel (uu, v) is associated with a discrete depth d n After back-projection to three-dimensional space, use the three-dimensional coordinates (x n ,y n , z n ) to do position encoding and convert the depth probability P n Element-wise fusion with feature maps; The fused features are converted into Key and Value forms that can be used by the Transformer decoder using a multi-layer perceptron (MLP) or convolution method.
6. The method for detecting long-distance three-dimensional targets on the road side based on guidance of prior geometric information according to claim 1, characterized in that: In step S4, the Transformer decoder performs a three-dimensional query Q 3D When performing multi-head attention operations with Key and Value, an attention mask based on three-dimensional Euclidean distance is introduced, which only allows high-weight matching of the query with several of its closest three-dimensional reference points to reduce false matching caused by sparse long-distance features.
7. The method for detecting long-distance three-dimensional targets on the road side based on guidance of prior geometric information according to claim 1, characterized in that: In step S4, the network training includes two losses: 2D detection and 3D detection. The 2D loss uses cross entropy or FocalLoss to constrain target classification. The 3D loss uses the Hungarian matching algorithm to perform a one-to-one correspondence between the predicted box and the true value box. The successfully matched target is subjected to L1 regression loss and Focal Loss classification loss to ensure that the three-dimensional bounding box and category are accurately regressed.
Citation Information
Cited By
Roadside monocular 3D target detection method, device and system, and storage medium
CN120877219A
Roadside vision three-dimensional target sensing method, device, equipment and medium
CN120953945A