A roadside vision three-dimensional target perception method, device, equipment and medium

CN120953945BActive Publication Date: 2026-10-09SHANGHAI UNIV OF ENG SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511076139.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2026-10-09
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

[0004]然而,路侧视觉与车侧视觉的场景差异(如视角、目标密度等)导致现有的车侧视觉三维目标检测方法应用在路侧场景下性能显著下降,尤其在复杂交通场景下多目标检测精度不足

Benefits of technology

[0024] This invention overcomes the limitations of single-vehicle perception and improves the accuracy of multi-target detection. The SimAM-T attention mechanism enhances feature extraction capabilities, while the dynamically incremental discretization method and probability estimation improve the accuracy of depth and height information determination. Geometric perspective analysis further refines the target depth and height information. This invention plays a significant role in promoting the development and application of autonomous driving technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953945B_ABST
    Figure CN120953945B_ABST
Patent Text Reader

Abstract

The application discloses a roadside visual three-dimensional target perception method, device, equipment and medium, wherein the method comprises the following steps: performing feature extraction on a complex road scene to obtain target features in a roadside image; performing projection transformation on the target features to determine the positions of the targets in a three-dimensional space; determining depth information and height information of the targets based on the positions of the targets in the three-dimensional space; analyzing geometric perspective relationships between the targets to correct the depth information and the height information of the targets; and obtaining detection results of each target in the three-dimensional space based on the corrected depth information and the height information. The application can break through the single-vehicle perception bottleneck and improve multi-target detection accuracy. The SimAM-T attention mechanism enhances feature extraction capability, and the dynamic incremental discretization method and probability estimation improve the accuracy of depth and height information determination. The application has an important promoting effect on the development and application of automatic driving technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle-road-cloud integration and autonomous driving, specifically to a roadside visual three-dimensional target perception method, device, equipment, and medium. Background Technology

[0002] The core of the perception system for autonomous and intelligent connected vehicles lies in 3D target detection. Currently, the perception range of vehicle-side vision sensors is limited (usually less than 100 meters) and is easily affected by obstacles, inclement weather, lighting conditions, and the intensity of surface reflection and motion of objects. To overcome the bottleneck of single-vehicle perception, vehicle-road-cloud collaborative technology has become a key direction. Among them, roadside perception systems compensate for the shortcomings of single-vehicle local perception through regional collaborative perception, providing a new path for beyond-line-of-sight target detection.

[0003] While LiDAR (Light Detection and Ranging) provides high-precision point cloud data in roadside perception systems, its high cost severely restricts large-scale deployment. Multi-sensor fusion solutions, although improving detection accuracy through heterogeneous data complementarity, face challenges such as high computational complexity and fusion distortion caused by spatiotemporal registration errors. In contrast, roadside vision perception systems based on monocular cameras offer significant economic and industrialization advantages. Roadside cameras are typically deployed on streetlight poles over ten meters high, achieving a perception distance exceeding 200 meters, while a larger pitch angle effectively reduces blind spots. Furthermore, the perception results can be transmitted in real-time to surrounding connected vehicles via V2X communication protocols, overcoming the limitations of single-vehicle sensors and greatly improving the driving safety of autonomous vehicles.

[0004] However, the differences between roadside vision and vehicle-side vision (such as viewing angle and target density) lead to a significant performance drop in the performance of existing vehicle-side vision 3D target detection methods when applied to roadside scenarios, especially in complex traffic scenarios where the accuracy of multi-target detection is insufficient. Summary of the Invention

[0005] To address the technical problems mentioned above, this invention provides a roadside visual three-dimensional target perception method, comprising the following steps:

[0006] S1. Extract features from complex road scenes to obtain target features in roadside images;

[0007] S2. Perform a projection transformation on the target features to determine the target's position in three-dimensional space;

[0008] S3. Based on the target's position in three-dimensional space, determine the target's depth and height information;

[0009] S4. Analyze the geometric perspective relationships between targets and correct the depth and height information of the targets;

[0010] S5. Based on the corrected depth and height information, obtain the detection results of each target in three-dimensional space.

[0011] Preferably, step S1 includes: introducing the SimAM attention mechanism based on energy function theory into the visual 3D target detection task, and enhancing the feature extraction capability by constructing a 3D attention weight matrix; at the same time, improving the SimAM module to SimAM-T, introducing a temporal smoothing regularization term to reduce energy fluctuations caused by single-frame noise.

[0012] Preferably, step S2 includes: transforming two-dimensional image points to three-dimensional coordinate points using the camera's intrinsic and extrinsic parameter matrices; introducing a virtual coordinate system and a reference plane, and achieving complementarity of height and depth information through transformation matrices and the geometric relationship of similar triangles, thereby projecting the two-dimensional image points into three-dimensional space.

[0013] Preferably, step S3 includes: discretizing continuous depth and height values ​​into multiple bins, introducing a probability distribution map into the network detection head branch, and then combining it with the sigmoid function to determine the depth and height information of the target in three-dimensional space.

[0014] Preferably, step S4 includes: constructing a dense directed graph, determining the target influence through three scores: depth confidence score, two-dimensional distance score, and classification similarity score, and then estimating the depth based on the geometric perspective depth relationship between targets, thereby correcting the target depth and height information.

[0015] The present invention also provides a roadside visual three-dimensional target perception device, the device being used to implement the above method, comprising: an extraction module, a transformation module, a positioning module, a correction module, and a detection module;

[0016] The extraction module is used to extract features from complex road scenes and obtain target features in roadside images;

[0017] The transformation module is used to perform projection transformation on the target features to determine the position of the target in three-dimensional space;

[0018] The positioning module is used to determine the depth and height information of the target based on its position in three-dimensional space.

[0019] The correction module is used to analyze the geometric perspective relationship between targets and correct the depth and height information of the targets;

[0020] The detection module is used to obtain the detection results of each target in three-dimensional space based on the corrected depth and height information.

[0021] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described method.

[0022] The present invention also provides a computer-readable storage medium storing a computer program that, when executed, implements the above-described method.

[0023] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0024] This invention overcomes the limitations of single-vehicle perception and improves the accuracy of multi-target detection. The SimAM-T attention mechanism enhances feature extraction capabilities, while the dynamically incremental discretization method and probability estimation improve the accuracy of depth and height information determination. Geometric perspective analysis further refines the target depth and height information. This invention plays a significant role in promoting the development and application of autonomous driving technology. Attached Figure Description

[0025] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a schematic diagram of the SimAM-T attention module embedded in the network backbone according to an embodiment of the present invention;

[0027] Figure 2 This is a schematic diagram of a depth-based two-dimensional to three-dimensional projection transformation according to an embodiment of the present invention;

[0028] Figure 3 This is a schematic diagram showing the approximate relationship of the target height from the vehicle side view in an embodiment of the present invention;

[0029] Figure 4 This is a schematic diagram of a height-based two-dimensional to three-dimensional projection transformation according to an embodiment of the present invention;

[0030] Figure 5 This is a schematic diagram of monocular vision probability estimation according to an embodiment of the present invention;

[0031] Figure 6 This is a schematic diagram of the height distribution of the boxes under different α values ​​according to an embodiment of the present invention;

[0032] Figure 7 This is a schematic diagram of the depth bin distribution under different β values ​​according to an embodiment of the present invention;

[0033] Figure 8This is a multi-target geometric perspective depth relationship diagram according to an embodiment of the present invention;

[0034] Figure 9 This is a framework diagram of the roadside visual three-dimensional target perception algorithm model according to an embodiment of the present invention;

[0035] Figure 10 This is a schematic diagram of the computer device structure according to an embodiment of the present invention.

[0036] Explanation of reference numerals in the attached figures:

[0037] 10. Processor; 20. Memory; 30. Input device; 40. Output device. Detailed Implementation

[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0039] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0040] First, let's introduce the technical terminology used in this invention. The term "target" as used here primarily refers to objects that need to be detected in roadside scenarios, specifically including the following categories:

[0041] Vehicles, including sedans, SUVs, and trucks, are among the key targets of roadside visual 3D target perception. In traffic scenarios, accurately perceiving vehicle position, orientation, and speed is crucial for autonomous vehicles to anticipate road conditions and plan routes. For example, images captured by roadside cameras, using the 3D target perception method described in this paper, can determine the specific positions of various types of vehicles on the road in 3D space, including the depth distance between the vehicle and the roadside camera, as well as the vehicle's position in the lateral and longitudinal directions of the road. This provides autonomous vehicles with precise location information of surrounding vehicles, helping them make better driving decisions such as following, overtaking, and avoiding obstacles.

[0042] Pedestrians: As a vulnerable group in traffic scenarios, pedestrian detection is crucial for ensuring traffic safety. Roadside vision systems can perceive pedestrian activities on sidewalks, crosswalks, and roads, including their location and direction of movement. For example, timely and accurate detection of a pedestrian's position and movement trend when they cross the road allows autonomous vehicles to react in advance by slowing down or stopping, thus preventing traffic accidents.

[0043] Traffic signs and markings: Although they are not dynamic targets in themselves, they play a crucial role in enabling autonomous vehicles to correctly understand road conditions and obey traffic rules. Roadside vision systems can identify various traffic signs (such as speed limit signs, no-entry signs, etc.) and markings (such as lane lines, stop lines, etc.) on the road, and convert this information into position and orientation information in three-dimensional space, providing autonomous vehicles with accurate road rule guidance so that they can drive in accordance with regulations.

[0044] Obstacles: These include construction obstacles and fallen objects on the road. These obstacles may pose a threat to the normal operation of vehicles and therefore need to be detected and located in a timely manner. Roadside vision-based 3D target perception methods can determine the position, size, and shape of obstacles in 3D space, enabling autonomous vehicles to plan detours in advance, avoid obstacles, and ensure driving safety.

[0045] Example 1

[0046] This embodiment provides a roadside visual three-dimensional target perception method, the steps of which include:

[0047] S1. Extract features from complex road scenes to obtain target features in roadside images.

[0048] In 3D object detection tasks, improving feature extraction capabilities for objects in complex road scenes is crucial for enhancing detection accuracy. Attention mechanisms are modules that can be easily integrated into existing neural networks without requiring large-scale modifications to the original architecture. The aim is to improve feature extraction capabilities without significantly increasing model complexity, thereby improving detection performance. Common attention mechanisms include channel attention (SE) and spatial attention (CBAM), which process channel and spatial information respectively through cascading. Furthermore, the ECA (Efficient Channel Attention) mechanism avoids dimensionality reduction and uses one-dimensional convolution to achieve cross-channel interaction, improving the efficiency of channel attention while maintaining lightweight design. However, this decoupled processing strategy struggles to capture the global features and contextual relationships of the target. To address this issue, this embodiment introduces the SimAM (Simple, Parameter-Free Attention Module) attention mechanism, based on energy function theory, into the visual 3D object detection task. By constructing a 3D attention weight matrix, collaborative optimization across channel and spatial dimensions is achieved to enhance feature extraction capabilities.

[0049] SimAM has no additional parameters, thus it can be seamlessly embedded into other networks and is very lightweight. Its principle is derived from the neuroscience theory of spatial suppression, which posits that neurons exhibiting significant spatial suppression effects are usually more important in visual processing. SimAM assigns an importance weight to each neuron in the feature map, thereby helping the network to better focus on important features. An energy function is used to improve the discriminative power of the target feature from other surrounding features, thereby enhancing the feature extraction capability. The energy function is shown in Equation (1).

[0050]

[0051] in, and They are the target neuron t and other neurons x, respectively. i The linear representation of ω t and b t These are the weights and biases of the linear transformation, M is the number of neurons in the channel, and i represents the index of the neuron. t and y o Let represent the ideal values ​​of the target neuron and other neurons, respectively. Since the outputs of neural networks are normalized to [-1, 1], to reduce computational complexity, a binary label is introduced to simplify the energy function, i.e., y = ... t Let y be 1. o Set it to -1.

[0052] This binary partitioning can directly measure the linear separability of the target neuron and its surrounding neurons, thereby guiding the generation of attention weights. This design not only conforms to the inspiration of "spatial inhibition" in neuroscience, but also simplifies the solution process to obtain a closed-form solution. The simplified energy function is shown in equation (2).

[0053]

[0054] Here, λ is the regularization coefficient, which prevents the energy function from having excessive weight during the solution process and avoids overfitting.

[0055] The weight ω is obtained through analytical solution. t and bias b t Closed-form solution:

[0056]

[0057] Where, μ t and These are the mean and variance of other neurons within the channel, respectively. This method significantly reduces computational cost, avoiding the need to iteratively calculate μ at each location. t and The process is tedious, and all neurons within the channel participate in the computation. Minimum energy It can be calculated using equation (5):

[0058]

[0059] A smaller value indicates that the target neuron is more distinct from other neurons, meaning that the feature is more important. Therefore, the feature importance of each neuron can be determined by... This is determined by weighted updating of the feature map.

[0060]

[0061] Where E is the minimum energy of all neurons. The set, the formula is obtained by combining with the characteristics Figure X The feature map is updated by multiplying each element together.

[0062] In roadside visual 3D target perception scenarios, roadside cameras are fixed at a high position, and consecutive frames captured are highly similar in height. The background, including roads and buildings, remains almost unchanged, and the motion trajectories of targets such as vehicles and pedestrians are smooth. Based on these characteristics, this embodiment demonstrates an improved design of the SimAM module, named SimAM-T (SimAM-Temporal). This is because the feature importance of each neuron is... Directly related, therefore this embodiment is... A temporal smoothing regularization term is introduced to smooth energy changes, thereby focusing more on temporally stable characteristics, reducing energy fluctuations caused by single-frame noise, and improving robustness. The improved minimum energy function is shown in Equation (7).

[0063]

[0064] Where δ represents the temporal regularization coefficient, and T represents the total number of image frames. Let represent the minimum energy of target neuron t in frame τ, where τ∈[1,T]. Finally, the feature map is updated by weighting and multiplying it with each element in X.

[0065]

[0066] Among them, E T It is the minimum energy of all neurons after introducing temporal smoothing regularization. The collection, improved SimAM-T modules such as Figure 1 As shown.

[0067] S2. Perform projection transformation on the target features to determine the target's position in three-dimensional space.

[0068] Vehicle-mounted cameras are typically installed on the roof or front of the vehicle. The Z-axis of the camera's coordinate system is usually parallel to the ground and points forward of the vehicle. The depth of the detected target can be approximated by its distance to the camera. Roadside cameras, on the other hand, are usually installed at a higher height, resulting in a pitch angle between the Z-axis of their camera's coordinate system and the ground. Figure 2 As shown, the distance from the camera to the target is not equal to the target depth. To solve this problem, this invention designs a depth-based projection transformation from two-dimensional image points to three-dimensional spatial points.

[0069] The specific process of this projection transformation is as follows: First, the two-dimensional points on the image are represented by p... i =[u,v,1] T The expression represents the transformation of two-dimensional image points to three-dimensional coordinate points, as shown in equation (9), where u and v are pixel coordinates.

[0070]

[0071] Where K represents the camera's intrinsic parameter matrix, d represents the target's depth, and R and t are the camera's rotation and translation matrices, i.e., extrinsic parameter matrices. This represents the three-dimensional coordinates of a point in a two-dimensional image after transformation.

[0072] In the camera coordinate system, the Z-axis points forward, the X-axis points to the right, and the Y-axis points downward, with the origin O located at the center of the image. In the pixel coordinate system, the origin is typically located at the top left corner of the image, the u-axis points to the right, and the v-axis points downward. Therefore, the transformation relationship can be obtained: for a point x = [x, y, 1] in a two-dimensional image... T Its coordinates in the pixel coordinate system are x′=[x,y,1] T Satisfying: u=x+C u v = y + C v Combining the projection relationship d·x′=K·X, we can obtain:

[0073] v·d=f·y+C v (10)

[0074] From the vehicle side view, the approximate relationship between the depth and height of the two targets in the image is shown in Equation (11).

[0075]

[0076] Where y represents the height of the target in the camera coordinate system, considering that the camera's Z-axis is parallel to the ground and all targets are in contact with the ground, the distance y2-y1 between the center points of two targets in the image should be approximately equal to the half-height difference between the two targets, i.e., (h2-h1) / 2. Figure 3 As shown.

[0077] However, in roadside view scenarios, the approximate relationship of equation (11) does not hold. This is because the roadside camera has a certain pitch angle with the ground. This results in the y-axis distance difference y2-y1 between the two target center points in the image not being equivalent to the half-height difference (h2-h1) / 2. To solve this problem, this invention designs a height-based two-dimensional-three-dimensional transformation, such as... Figure 4 As shown.

[0078] This embodiment introduces a virtual coordinate system and a reference plane. The virtual coordinate system shares the origin with the camera coordinate system, but its Y-axis points downwards and is perpendicular to the ground, while its Z-axis points forwards and is parallel to the ground. The reference plane is defined as parallel to the image plane, with a depth value d for each image point. ref All are 1. Point p in the two-dimensional image... i When projected onto the reference plane, at this point,

[0079] p r =K -1 ·d ref [u,v,1] T =K -1 [u,v,1] T (12)

[0080] Where, p r It is image point p i Projection onto the reference plane.

[0081] Using a virtual coordinate system, similar triangles can be constructed from the diagram:

[0082]

[0083] At this point, the point on the reference plane and the point in the three-dimensional world... The relationship between them is:

[0084]

[0085] in, and These represent the transformation matrices from the reference plane to the virtual coordinate system and from the virtual coordinate system to the world coordinate system, respectively. H and h represent the heights of the camera and the object, respectively. Point p r Along the Y-axis to the origin O of the virtual coordinate system vir The distance.

[0086] The core idea of ​​this algorithm is to align the roadside camera coordinate system with the world coordinate system, thereby facilitating subsequent perception processing. The projection process is divided into four stages, each consisting of a specific geometric transformation. First, the input two-dimensional feature map F is processed. 2D ={p1,...,p i}, introduce a reference plane parallel to the camera plane with a fixed depth value of 1, and set each two-dimensional feature point p i =[u,v,1] T Projected onto reference plane point P r Next, an intermediate three-dimensional virtual coordinate system was introduced, and a transformation matrix was used to... Reference plane point P r Transformed to a three-dimensional virtual coordinate system to obtain Then, in the virtual coordinate system, based on the height H of the roadside camera and the height h of the target, the geometric relationship of similar triangles is used. This allows us to obtain the target's depth value, achieving complementarity between height and depth information. Finally, a transformation matrix is ​​used... Transform the target point in the virtual coordinate system to the target point in the world coordinate system. It realizes the projection transformation from two-dimensional image points to three-dimensional spatial points.

[0087] S3. Based on the target's position in three-dimensional space, determine the target's depth and height information.

[0088] To address the ambiguity issue in monocular vision depth estimation, this embodiment transforms the regression problem into a classification problem. That is, continuous depth values ​​are discretized into multiple "depth bins," each representing a specific range of depth values. By predicting which "depth bin" a target belongs to, the target's depth can be estimated. Figure 5 As shown.

[0089] In regression tasks, it is usually necessary to design new loss functions to adapt to the depth estimation problem, and only a single predicted value can be output. However, by transforming the regression task into a classification task, it is possible to calculate the probability distribution (such as the Softmax result) and output multiple possible depth solutions. The cross-entropy loss function can also be used, which has been widely validated and is well-suited for classification tasks. Considering the difference between roadside and vehicle-side views, and combining this invention's 2D-3D projection transformation module designed for height, this invention performs probabilistic estimation for both depth and height values. Considering that height h and depth d are continuous within a certain range, this invention employs a dynamically increasing discretization method, while simultaneously defining the height range [H... min H max ] and depth range [D min D max If the data is discretized into height and depth bins, and the number of discretized bins is N, then:

[0090]

[0091] Among them, h i and d iLet represent the indices of the discrete bins corresponding to height and depth, respectively. α and β represent the hyperparameters representing the range variations of the height and depth bins, respectively. The binning strategies corresponding to different hyperparameter values ​​are as follows: Figure 6 and Figure 7 As shown.

[0092] The discretized height and depth values ​​are represented as weight vectors ω. h ∈R N and ω d ∈R N Then, a network detection head branch is introduced to generate a probability distribution map H. PM and D PM , as in equation (16).

[0093] H P =ω h T softmax(H PM ), D P =ω T softmax(D PM (16)

[0094] Among them, H P and D P Let H represent the expected value of the probability distribution, denoted as probability height and probability depth. Then, the height prediction value H from the direct regression... R and depth prediction value H P The parameter σ is fused using a sigmoid function, as shown in equation (17):

[0095] H I =σH R +(1-σ)H P D I =σD R +(1-σ)D P (17)

[0096] To reduce the impact of direct regression height and depth values ​​on the final height and depth predictions, the parameter σ is set to 0.25. Here, H... I and D I These represent the final predicted height and depth values ​​for a single target, respectively.

[0097] S4. Analyze the geometric perspective relationships between targets and correct the depth and height information of the targets.

[0098] In roadside scenarios, cameras are deployed at a high position, detecting a much larger number of targets than from the vehicle's side view. However, targets at traffic intersections are usually neatly arranged between lanes, so the geometric positional relationships between targets can be utilized. According to equation (18), the geometric perspective depth relationship between target 1 and target 2 is denoted as...

[0099]

[0100] From the above formula, we can see that the depth of target 2 can be estimated from the depth of target 1. Therefore, the depth of a certain target can be estimated based on the depths of other targets. For ease of understanding, as follows... Figure 8 As shown, a dense directed graph is constructed, where any two targets are connected by two bidirectional edges representing a depth relationship. Assume there are N = {1, 2, ..., n} targets to predict, and given... In the case of estimating the depth of target i based on target j, for all targets i∈N, the geometric depth is defined as...

[0101] Considering the large number of targets and the potential for significant errors due to the varying depths of different target types (such as pedestrians), three scores are used to determine which targets have the greatest influence, including depth confidence scores. Two-dimensional distance score and classification similarity The depth confidence score can be derived from the probability depth distribution D of each target. P The scores are obtained from [the first score], and the other two scores are calculated as follows:

[0102]

[0103] in, It is the two-dimensional distance between the projection centers of targets i and j in the pixel coordinate system. f represents the pixel length of the image diagonal. i and f j This represents the output confidence vector of the two targets in the classification branch, and the total score is then expressed as S. j→i In summary, the geometric perspective depth relationship between multiple targets is given by equation (21).

[0104]

[0105] S4. Based on the corrected depth and height information, obtain the detection results of each target in three-dimensional space.

[0106] At this point, the probabilistic depth estimate for a single target and the geometric perspective depth prediction for multiple targets have been obtained. I and D G To integrate these two complex results, the total score S is... j→i The highest score S max D will be used as a parameter I and D G The fusion is performed to obtain the final fusion depth D for each target:

[0107] D=σ(S max )D I +(1-σ(S max ))D G . (twenty two)

[0108] Example 2

[0109] The following will describe in detail, with reference to this embodiment, how the present invention solves the technical problems in practical work.

[0110] Typically, 2D object detection aims to predict the 2D bounding box and class label for each object of interest, while 3D object detection aims to predict the position, orientation, and size of each object in the camera coordinate system. Position represents the object's 3D coordinates in the camera coordinate system. The object's orientation in the camera coordinate system is usually represented by the angle between the object and the XOZ plane of the camera coordinate system. The sign of the angle can be determined using a right-handed coordinate system: with your right hand gripping the downward-pointing Y-axis, the direction in which your thumb points in the same direction as the Y-axis, and your four fingers bend, is positive. Therefore, in the XOZ plane, clockwise is positive, and counterclockwise is negative. Target size refers to the actual length, width, and height information of the target; monocular 3D object detection extracts semantic feature information from a single 2D image to determine the target's 3D bounding box, which can be represented as: B = {B1, B2, ..., B...} n}. Each bounding box B i It consists of seven degrees of freedom: B i =(x i ,y i ,z i ,l i ,w i ,h i ,θ i Here, x, y, and z represent the three-dimensional position coordinates (m) of the center point of each three-dimensional bounding box. l, w, and h represent the length, width, and height (m) of the three-dimensional bounding box, respectively. θ represents the global azimuth angle of each target in space, that is, the angle between the target and the XOZ plane of the camera coordinate system, ranging from [-π, π].

[0111] However, the model in this embodiment does not directly predict the target's seven degrees of freedom (x, y, z, w, l, h, θ), but rather predicts the target's nine degrees of freedom (Δx, Δy, d, w, l, h, sin(θ, C). θ(c) In 2D object detection tasks, the model needs to regress the distances from the target's center point to the four sides of the 2D detection box (top, bottom, left, and right), denoted as t, b, l, and r, respectively. However, in 3D object detection tasks, it is difficult to regress the distances from the target's center to the six faces of the 3D bounding box. Furthermore, directly regressing the target's absolute 3D coordinates (x, y, z) is easily affected by uneven distribution of local features. Therefore, regressing the target's "2.5D" center coordinates is more appropriate, which can be summarized as regressing the offsets Δx and Δy of the target's predicted 2D center point from the true center point, along with its corresponding depth d. Since the camera's Z-axis direction is consistent with the shooting direction, the target's Z-axis coordinate z is the target's depth d from the camera. The target's 2D image center point can be transformed into a 3D center point in the camera coordinate system using the camera's intrinsic parameter matrix, thus achieving the regression of the center point coordinates. Directly regressing the azimuth angle θ easily encounters problems of angular periodicity and numerical discontinuity. Using sin(θ) as the azimuth angle representation maps the direction angle to the interval [-1, 1], ensuring that the output matches the range of the activation function and avoiding numerical overflow or scaling issues. However, the azimuth angle θ represents the target's orientation in the camera coordinate system, not its facing direction. Therefore, a direction category C is introduced. θ The target orientation is determined by distinguishing whether the heading angle is acute or obtuse. Centrality (c) determines which predicted points are closer to the true center of the target, thus suppressing low-quality predicted boxes that are far from the target center.

[0112] The overall process is as follows Figure 9 As shown: First, the roadside image is input into a neural network, which extracts a feature map rich in semantic information. This feature map is then passed to a 3D weight-based attention module, which generates 3D weights based on the extracted features. The 2D-3D coordinate transformation is achieved by combining camera parameters, the ground plane equation, and a depth and height-based projection module. Since the depth and height information of the roadside scene are complementary, the task of depth and height regression of a single isolated object is reformulated as a discrete probability classification problem. Furthermore, considering the geometric relationships between multiple targets, the depth estimates for single and multiple targets are weighted to obtain a fused depth. These features are then passed to the detection head to output the 3D object detection results.

[0113] This embodiment designs an experiment on the open-source roadside dataset Rope3D, which uses AP|R40 as the evaluation index, as shown in Equation (23).

[0114]

[0115] in, This represents the precision of a certain recall threshold r∈{1 / 40,2 / 40,...,1}.

[0116] The Rope3D dataset further divides the comprehensive evaluation index into several sub-indicators: ACS (Average Center Similarity), which is used to evaluate the accuracy of the predicted target center point. ACS is derived by calculating the distance and similarity between the predicted target center point and the actual target center point, as shown in Equation (24).

[0117]

[0118] Where D is the positive sample set, C s It is the norm of the true value of the target center. It is the Euclidean distance between the predicted center value and the true center value of sample s. AOS (Average Orientation Similarity): measures the degree of orientation estimation; AOS is used to evaluate the accuracy of target localization.

[0119]

[0120] in, It is the angle difference of sample s. This indicates that the evaluation does not distinguish whether the head or tail of an object is facing the camera. ASS (Average Area Similarity): This measures the area occupied by the predicted target relative to the real target, where... It is the difference in absolute area, A s Represents the area of ​​the actual target.

[0121]

[0122] AGD (Average Distance Similarity): Since the position, orientation, length and width of the four vertices of the ground plane of a 3D bounding box are considered together, the similarity is measured by calculating the average distance between these four vertices.

[0123]

[0124] in, and s g Let G and G represent the g-th predicted point and the true point of sample s, respectively, and K = 4 represent the total number of ground points in the 3D bounding box.

[0125] To maintain consistency with other similarity metrics, AGS (3D bounding box corner similarity) is defined as:

[0126]

[0127] Let S = (ACS + AOS + AAS + AGS) / 4, and then combine them into Rope by reweighting AP and S (ω1 = 8 and ω2 = 2). score (Rope 3D dataset score) serves as another evaluation metric for the model's detection performance on the Rope 3D dataset;

[0128] Rope score =(ω1*AP+ω2*S) / (ω1+ω2). (29)

[0129] To evaluate the performance of the roadside 3D object detection algorithm of this invention, experiments were conducted on an NVIDIA RTX 4070 Ti Super 16G GPU. The batch size was set to 1, the learning rate to 0.001, and the SGD optimizer was used, with a training batch size of 48. Model training and performance evaluation were performed on the Rope3D dataset, and several existing pure vision-based roadside 3D object detection models were selected for comparison. The experimental results on the Rope3D dataset are shown in Table 1. All algorithms used single-frame image data.

[0130] Table 1

[0131]

[0132] Example 3

[0133] This embodiment also provides a roadside visual three-dimensional target perception device, including: an extraction module, a transformation module, a localization module, a correction module, and a detection module; the extraction module is used to extract features from complex road scenes and obtain target features in roadside images; the transformation module is used to perform projection transformation on the target features to determine the position of the target in three-dimensional space; the localization module is used to determine the depth and height information of the target based on the position of the target in three-dimensional space; the correction module is used to analyze the geometric perspective relationship between targets and correct the depth and height information of the targets; the detection module is used to obtain the detection result of each target in three-dimensional space based on the corrected depth and height information.

[0134] Example 4

[0135] The present invention also provides a computer device having the above-described roadside visual three-dimensional target perception device.

[0136] See Figure 10 , Figure 10 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 10As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 10 Take a processor 10 as an example.

[0137] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0138] The memory 20 stores instructions executable by at least one processor 10 to cause the processor 10 to perform the methods shown in the above embodiments. The memory 20 may include a stored program area and a stored data area, wherein the stored program area may store applications required for at least one function of the operating system; and the stored data area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0139] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk, or solid-state drive; the memory 20 may also include combinations of the above types of memory. The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 may be connected via a bus or other means. Figure 10Taking a bus connection as an example, input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some optional embodiments, the display device may be a touch screen.

[0140] Example 5

[0141] This embodiment also provides a computer-readable storage medium. The methods described above according to embodiments of the present invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code that, when executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0142] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A roadside visual three-dimensional target perception method, characterized in that the steps include: include: S1. Extract features from complex road scenes to obtain target features in roadside images; S2. Perform a projection transformation on the target features to determine the target's position in three-dimensional space. The steps include: transforming the two-dimensional image points corresponding to the target features to three-dimensional coordinate points using the camera's intrinsic and extrinsic parameter matrices; introducing a virtual coordinate system and a reference plane, and projecting the two-dimensional image points into three-dimensional space through the transformation matrix and the geometric relationships of similar triangles; wherein the virtual coordinate system shares the origin with the camera coordinate system, but the virtual coordinate system... Y The axis is perpendicular to the ground. Z The axis faces forward and is parallel to the ground. The reference plane is defined as parallel to the image plane, and the depth value of each image point is 1; the two-dimensional image points When projected onto the reference plane, at this point... in, Image points Projection onto the reference plane; The projection process consists of four stages; the first stage involves processing the input two-dimensional feature map. Introduce a reference plane parallel to the camera plane with a fixed depth value of 1, and set each two-dimensional feature point... Projected onto reference plane point Next, an intermediate three-dimensional virtual coordinate system was introduced, and a transformation matrix was used to... Reference plane point Transformed to a three-dimensional virtual coordinate system to obtain Then, in the virtual coordinate system, based on the height of the roadside camera... H and the height of the target h Using the geometric relationship of similar triangles This is used to obtain the target's depth value, achieving complementarity between height and depth information; finally, a transformation matrix is ​​used... Transform the target point in the virtual coordinate system to the target point in the world coordinate system. It realizes the projection transformation from two-dimensional image points to three-dimensional spatial points; S3. Based on the target's position in three-dimensional space, determine the target's depth and height information; S4. Analyze the geometric perspective relationships between targets and correct the depth and height information of the targets; S5. Based on the corrected depth and height information, obtain the detection results of each target in three-dimensional space.

2. The roadside visual three-dimensional target perception method according to claim 1, characterized in that, Step S1 includes: introducing the SimAM attention mechanism based on energy function theory into the visual 3D object detection task, and constructing a 3D attention weight matrix; the steps include: in, and Target neurons t Other neurons xi The linear representation of, ωt and bt These are the weights and biases of the linear transformation. M It is the number of neurons in the channel. i Index representing a neuron; yt and yo These represent the ideal values ​​for the target neuron and other neurons, respectively. Binary partitioning directly measures the linear separability of the target neuron from its surrounding neurons, guiding the generation of attention weights. The simplified energy function is shown below: in, λ These are regularization coefficients, used to prevent excessive weighting of the energy function during the solution process and avoid overfitting; the weights are obtained through analytical solutions. ωt and bias bt Closed-form solution: in, μ t and These are the mean and variance of the other neurons within the channel, respectively; Minimum Energy The calculation is as follows: ; The importance of each neuron's features is determined by... To determine this, the feature map is ultimately updated using a weighted algorithm: in, E It is the minimum energy of all neurons The set, through the feature map X Each element is multiplied to weight and update the feature map; The SimAM module is improved to SimAM-T, and a timing smoothing regularization term is introduced. The steps include: in, δ Represents the time series regularization coefficient. T Indicates the total number of frames in the image. Indicates the first τ Target neuron in frame t The minimum energy, ; Ultimately with X The feature map is updated by multiplying each element in the matrix. in, ET It is the minimum energy of all neurons after introducing temporal smoothing regularization. A set of.

3. The roadside visual three-dimensional target perception method according to claim 1, characterized in that, The steps in S3 include: discretizing continuous depth and height values ​​into several bins, introducing a probability distribution map into the network detection head branch, and then combining the sigmoid function to determine the depth and height information of the target in three-dimensional space.

4. The roadside visual three-dimensional target perception method according to claim 1, characterized in that, Step S4 includes: constructing a dense directed graph; determining the target influence through three scores: depth confidence score, two-dimensional distance score, and classification similarity score; then estimating the depth based on the geometric perspective depth relationship between targets; and correcting the target's depth and height information. In the dense directed graph, any two targets are connected by two bidirectional edges representing the depth relationship, and a set... A prediction target, given In the case of target j To estimate the target i The depth, for all targets Defined as geometric depth ; Three scores are used to determine which targets are most influential, including the deep confidence score. Two-dimensional distance score and classification similarity The depth confidence score is derived from the probability depth distribution of each target. DP The scores are obtained from [the first score], and the other two scores are calculated as follows: in, The goal i and j The two-dimensional distance between the projection centers in the pixel coordinate system. Indicates the pixel length of the image diagonal. fi and fi This represents the output confidence vector of the two targets in the classification branch, and then the total score is expressed as... In summary, the geometric perspective depth relationship between multiple targets is shown in the following equation: 。 5. A roadside visual three-dimensional target perception device, the device being used to implement the method according to any one of claims 1-4, characterized in that, include: Extraction module, transformation module, positioning module, correction module, and detection module; The extraction module is used to extract features from complex road scenes and obtain target features in roadside images; The transformation module is used to perform projection transformation on the target features to determine the target's position in three-dimensional space. The process includes: transforming the target features from two-dimensional image points to three-dimensional coordinate points using the camera's intrinsic and extrinsic parameter matrices; introducing a virtual coordinate system and a reference plane, and projecting the two-dimensional image points into three-dimensional space through the transformation matrix and the geometric relationship of similar triangles; wherein, the virtual coordinate system shares the origin with the camera coordinate system, but the virtual coordinate system... The axis is perpendicular to the ground. The axis faces forward and is parallel to the ground; the reference plane is defined as parallel to the image plane, and the depth value of each image point is 1; the two-dimensional image points... p i When projected onto the reference plane, at this point... in, p r Image points p i Projection onto the reference plane; The projection process consists of four stages; the first stage involves processing the input two-dimensional feature map. Introduce a reference plane parallel to the camera plane with a fixed depth value of 1, and set each two-dimensional feature point... Projected onto reference plane point Next, an intermediate three-dimensional virtual coordinate system was introduced, and a transformation matrix was used to... Reference plane point Transformed to a three-dimensional virtual coordinate system to obtain Then, in the virtual coordinate system, based on the height of the roadside camera... H and the height of the target h Using the geometric relationship of similar triangles This is used to obtain the target's depth value, achieving complementarity between height and depth information; finally, a transformation matrix is ​​used... Transform the target point in the virtual coordinate system to the target point in the world coordinate system. It realizes the projection transformation from two-dimensional image points to three-dimensional spatial points; The positioning module is used to determine the depth and height information of the target based on its position in three-dimensional space. The correction module is used to analyze the geometric perspective relationship between targets and correct the depth and height information of the targets; The detection module is used to obtain the detection results of each target in three-dimensional space based on the corrected depth and height information.

6. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Vehicle-mounted visual real-time multi-target multi-task joint sensing method and device

    CN111310574A

  • Roadside long-distance three-dimensional target detection method based on prior geometric information guidance

    CN120107902A