Intelligent traffic cone barrel anti-collision early warning method based on binocular vision
By using a binocular vision-based intelligent traffic cone collision avoidance warning method, and combining high-precision binocular camera calibration with an improved RT-DETR algorithm and multimodal trajectory prediction technology, the method achieves vehicle behavior prediction and real-time collision warning in complex road environments. This addresses the shortcomings of existing technologies and improves the accuracy and reliability of the warning.
Patent Information
- Application Number
- CN202510640617.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-10-31
AI Technical Summary
Existing intelligent traffic cone collision warning technologies are inadequate in terms of warning timeliness, detection range, adaptability to complex environments, and vehicle behavior prediction, making it difficult to meet the needs of modern complex and ever-changing traffic environments.
A binocular vision-based intelligent traffic cone collision avoidance warning method is adopted. Through high-precision binocular camera calibration, improved RT-DETR algorithm, binocular vision depth perception and StrongSORT multi-target tracking algorithm, combined with multimodal trajectory prediction technology of generative adversarial network, the method realizes real-time vehicle speed and distance measurement, and provides multi-level warning through collision risk index.
It significantly improves the accuracy and reliability of early warnings, reduces the false alarm rate, and can provide timely and accurate collision warnings in complex road environments, thereby enhancing the safety management level of road construction areas.
Smart Images

Figure CN120877552A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of road traffic technology, and in particular to an intelligent traffic cone collision avoidance and early warning method based on binocular vision. Background Technology
[0002] With the continued acceleration of global urbanization, the scale of transportation road networks is expanding at an unprecedented rate, leading to increasingly frequent road construction, maintenance, and temporary traffic control scenarios. In these scenarios, traffic cones, as crucial traffic management tools, undertake key tasks such as separating traffic flows, guiding traffic direction, and protecting the safety of construction sites, playing an irreplaceable role in reducing traffic accidents and alleviating traffic congestion. However, traditional traffic cones have limited functionality, relying mainly on their physical shape and reflective markings for warnings. This results in limited warning effectiveness and a lack of proactive perception and early warning capabilities, making them ill-suited to the complex and ever-changing modern traffic environment. Therefore, their intelligent upgrading, particularly the improvement of collision avoidance warning capabilities, has become a critical issue that urgently needs to be addressed in the field of traffic management.
[0003] Currently, intelligent collision avoidance and early warning technologies for traffic cones have mainly developed along three major technical routes: attitude sensors, radar systems, and visual perception.
[0004] 1. Collision Avoidance Warning Scheme Based on Attitude Sensors: This scheme integrates devices such as accelerometers and gyroscopes inside the cone to monitor the cone's attitude changes in real time, thereby determining whether a collision or tipping event has occurred and triggering a warning. Although this scheme is effective in detecting direct collisions, its technical limitations are significant. The main issues are a noticeable delay in warning response, making it impossible to predict potential hazards in advance; furthermore, the attitude sensors are susceptible to interference from environmental factors (such as wind and ground vibration), resulting in a high false alarm rate and reducing the overall reliability of the warning system.
[0005] 2. Radar-based collision avoidance warning solution: This solution utilizes millimeter-wave radar or lidar technology to measure the distance and relative speed between the vehicle and the cones by emitting and receiving electromagnetic waves, thereby assessing the risk of collision. While this solution can achieve a certain degree of distance detection, its ability to identify the specific type, size, and driving intention of the vehicle is limited, making it difficult to provide a comprehensive threat assessment. Furthermore, radar systems are susceptible to environmental factors such as reflections from metallic objects and adverse weather conditions, leading to decreased detection accuracy; and radar equipment is expensive and complex to maintain, increasing the deployment and maintenance costs of the system.
[0006] 3. Vision-based collision avoidance warning solution: This solution uses cameras to capture real-time images of the construction area, combines computer vision and object detection algorithms to identify vehicles approaching traffic cones, and estimates their relative distance to the cones. However, existing vision-based solutions have significant limitations. For example, the performance of object detection algorithms degrades in complex road environments, inclement weather, or under changing lighting conditions, leading to missed detections or false alarms; monocular cameras lack stereo depth information, resulting in low distance estimation accuracy and susceptibility to differences in vehicle size; and existing vision systems have limited ability to predict future vehicle trajectories, making it difficult to accurately predict lane-changing and turning behaviors in complex road conditions.
[0007] Therefore, existing intelligent traffic cone collision avoidance warning technologies have shortcomings in terms of warning timeliness, detection range, adaptability to complex environments, and vehicle behavior prediction. In view of this, this invention proposes an intelligent traffic cone collision avoidance warning method based on binocular vision. This method aims to achieve comprehensive, accurate, and timely warnings of potential collision risks by utilizing high-precision binocular camera calibration, real-time high-precision vehicle detection, real-time ranging and speed measurement technology combining depth perception and multi-target tracking, and multimodal trajectory prediction technology based on generative adversarial networks (GANs). This will effectively improve the safety management level of road construction areas. Summary of the Invention
[0008] The purpose of this invention is to provide an intelligent traffic cone collision avoidance and early warning method based on binocular vision, comprising the following steps:
[0009] Step S1: Perform high-precision calibration of the binocular camera based on circle detection;
[0010] Step S2: Collect video data of the road construction area, preprocess, label, enhance and segment the data to construct a vehicle detection dataset;
[0011] Step S3: Improve the RT-DETR (Real-Time Detection Transformer) algorithm and train it;
[0012] Step S4: Based on the improved RT-DETR (Real-Time Detection Transformer) algorithm, combined with binocular vision depth perception and StrongSORT multi-target tracking algorithm, real-time vehicle speed measurement and distance measurement are achieved.
[0013] Step S5: Based on the vehicle position and speed information obtained in step S4, predictive analysis of the vehicle trajectory is realized. By calculating the probability of the intersection between the predicted vehicle trajectory and the traffic cone protection area, the collision risk index (CRI) is calculated, and multi-level early warning is realized based on the collision risk index.
[0014] Furthermore, in step S1, the method for high-precision calibration of a binocular camera based on circle detection includes the following steps:
[0015] S11, Calibration board selection and image acquisition:
[0016] A 7×7 circular array calibration board with a solid circle diameter of 20mm and a center-to-center distance of 40mm between adjacent circles is used, with a printing accuracy error of ≤0.01mm. A binocular camera is mounted on an intelligent cone to synchronously acquire images of the calibration board in 20 different poses (horizontal tilt ±45°, vertical tilt ±30°, evenly distributed at a distance of 0.5-3m).
[0017] S12, Binocular camera calibration:
[0018] ① Accurate extraction of the centroid of the ellipse: First, for the calibration board image captured by the stereo camera, an improved image processing technique is used to extract the centroid of the solid circle ellipse; bilateral filtering is used instead of traditional Gaussian filtering for image preprocessing, which can maintain clear edges while effectively removing noise; the bilateral filtering expression is as follows:
[0019]
[0020] After filtering, improved Canny edge detection is performed, calculating the image gradient magnitude G and gradient direction θ:
[0021]
[0022] θ = arctan(G) y / G x (1.3) Finally, the sub-pixel precise coordinates of the centroid of the ellipse are obtained by calculating the geometric moments:
[0023]
[0024] Wherein, geometric moment m pq Defined as:
[0025]
[0026] ② Parameter optimization combining Zhang's calibration and bundled adjustment: Based on the extracted centroid coordinates of the ellipse, the initial values of the camera parameters are determined using Zhang's calibration method; Zhang's method is based on the homography relationship between the planar calibration plate and the image.
[0027]
[0028] Where H is the homography matrix:
[0029] H=[h1 h2 h3]=λA[r1 r2 t] (1.8) A is the camera intrinsic parameter matrix, which includes parameters such as focal length and optical center position;
[0030] Solving for camera parameters by minimizing the reprojection error:
[0031]
[0032] Then, a bundled adjustment optimization algorithm is introduced to construct the normal equation:
[0033]
[0034] By solving, both camera parameters and 3D point coordinates are optimized:
[0035]
[0036] X = N 22 -1 (W2-N 21 t) (1.12)
[0037] ③ Extrinsic parameter optimization based on diagonal constraints: To further improve the system's measurement accuracy, the diagonal length of the calibration plate is used as a known quantity for extrinsic parameter optimization; the diagonal length is calculated as follows:
[0038]
[0039] Introduce the proportionality coefficient α = 1 / t x get:
[0040]
[0041] The formula for calculating the proportionality coefficient α is:
[0042]
[0043] The calibration plate has two diagonals, denoted as L. l and L r The corresponding translation vector is:
[0044]
[0045] Taking into account the constraints of the two diagonals, the extrinsic parameter matrix is finally optimized as follows:
[0046] T = (T l +T r ) / 2 (1.17)
[0047] Furthermore, in step S2, video data of the road construction area is collected, and the data is preprocessed, labeled, enhanced, and segmented to construct a vehicle detection dataset. The method is as follows:
[0048] S21, Data Acquisition:
[0049] By deploying high-definition binocular cameras in the road construction area, video data with a resolution of 1080P and a frame rate of 30fps is collected to ensure coverage of scenes under different time periods, weather, and lighting conditions. At the same time, construction areas, road sections with cones and warning signs, and complex vehicle interaction scenarios covered by different weather and lighting conditions are selected from BDD100K, Cityscapes, and KITTI datasets and integrated into the vehicle detection dataset to enrich the diversity of the data.
[0050] S22, Data Preprocessing:
[0051] Spatiotemporal sampling was performed on the collected video footage. For typical scenes with stable traffic flow, normal vehicle speed, and good lighting conditions, 3 frames per second were extracted. For special event scenes such as sudden vehicle deceleration or lane change, vehicles approaching the boundary of a construction area, severe weather conditions, or insufficient lighting, 10 frames per second were extracted as training data. Subsequently, the extracted images underwent enhancement processing, including exposure correction, color balance, noise reduction, and sharpening, to improve image quality. Finally, all training images were uniformly scaled to a resolution of 1280×720 pixels and normalized to ensure data standard consistency.
[0052] S23, Data Labeling:
[0053] The processed images are labeled to mark the precise location and category information of various vehicles in the images;
[0054] S24, Data Augmentation:
[0055] By using traditional image enhancement techniques such as flipping, rotating, scaling, cropping, and brightness adjustment, the dataset size is expanded and the sample diversity is increased. At the same time, DCGAN (Deep Convolutional Generative Adversarial Network) technology is applied to synthesize rare scene samples in reality, such as extreme weather and special lighting, to improve the model's adaptability in complex environments.
[0056] S25, Dataset Partitioning
[0057] The constructed labeled images are divided according to the ratio of training set:validation set:test set = 7:2:1.
[0058] The method for improving and training the RT-DETR (Real-Time Detection Transformer) algorithm in step S3 is as follows:
[0059] S31, Improved Backbone Network Design:
[0060] ① Combining Faster-Block and RepConv from FasterNet, a new feature extraction module is designed: the RepFaster module. This RepFaster module employs a multi-branch structure during training, including a standard 3×3 convolutional branch and a pointwise 1×1 convolutional branch, as shown below:
[0061] F out =W 3×3 *X+W 1×1 *X+b (3.1)
[0062] Among them, W 3×3 W represents the parameters of the 3×3 convolution kernel. 1×1 Let X represent the parameters of a 1×1 convolution kernel, X be the input feature map, and b be the bias term. During the inference phase, multi-branch structures are merged into a single convolution operation through parameter equivalence.
[0063] W equiv =W 3×3 +expand(W 1×1 (3.2)
[0064] b equiv =b (3.3)
[0065] F out =W equiv *X+b equiv (3.4)
[0066] Among them, expand(W) 1×1 This indicates that the parameters of a 1×1 convolution kernel are expanded to make it compatible with a 3×3 convolution kernel;
[0067] In partial convolution, the input feature X is divided into n div Groups are formed by processing each group of features separately and then merging them, as shown below:
[0068]
[0069]
[0070] ② Integrate the RepFaster module into the basic modules of the ResNet backbone network to form a new backbone network; the traditional ResNet basic block is represented as:
[0071] F out =F(x)+x (3.7)
[0072] F(x)=W2*σ(W1*x) (3.8)
[0073] Where σ is the activation function, and W1 and W2 are convolution parameters; the improved structure replaces W2 with the RepFaster module:
[0074] F(x)=RepFaster(σ(W1*x)) (3.9)
[0075] S32, Improved intra-scale feature interaction design:
[0076] The HiLo attention mechanism is introduced into the Transformer encoder layer to improve the AIFI module in the original network. The processing details of the HiLo attention mechanism are as follows:
[0077] ① By capturing local details such as vehicle outlines and logos through high-frequency branching, the problem of missed detection under occlusion and motion blur in traditional methods is solved. The specific implementation details are as follows:
[0078] First, the feature map is segmented into multiple local regions using window partitioning:
[0079] X windows =reshape(X,[B,h) group ,window_size,w group ,window_size,C]) (3.10)
[0080] Among them, h group =H / / window_size,w group =W / / window_size, B is the batch size, H and W are the height and width of the feature map, C is the number of channels, and the window size is s×s;
[0081] Perform linear projection on the divided local regions:
[0082] Q h ,K h V h =linear projection (X windows (3.11)
[0083] in, d is the dimension of each attention head;
[0084] Calculate attention score:
[0085]
[0086] Weighted summation:
[0087] HiFiout = Attention h ·V h (3.13)
[0088] The window features are recombined into a complete feature map by inverse window partitioning:
[0089] X HiFi =WindowReverse(HiFiout) (3.14)
[0090] ② By modeling global context information such as road scenes and relative vehicle positions through the low-frequency branch (LoFi), the problem of missed detections and false detections caused by insufficient modeling of global context information in traditional methods is solved. The specific implementation details are as follows:
[0091] First, downsampling is performed using average pooling to reduce the resolution:
[0092] X down =AvgPool2d(X,kernel_size=window_size) (3.15)
[0093] Where AvgPool2d represents the average pooling operation;
[0094] Perform linear projection:
[0095] Q l =linear_projection q (X) (3.16)
[0096] K l V l =linear_projection kv (X down (3.17)
[0097] in, H down =H / / window size W down =W / / window size ;
[0098] Calculate attention score:
[0099]
[0100] Weighted summation:
[0101] LoFi out =Attention l ·V l (3.19)
[0102] ③ Feature fusion: High-frequency and low-frequency features are concatenated along the channel dimension to form a complete output feature.
[0103] Output = concat([X HiFi ,X LoFi ],dim=C) (3.20)
[0104] S33, Improved Feature Fusion Module:
[0105] Design a "Multi-Scale Dilated Convolutional Attention Pyramid Network (MDC-APN)" to replace the original feature fusion module CCFM. This MDC-APN has the following key design features:
[0106] ① Multi-scale dilated convolution module (MDC): Processes features in parallel through dilated convolutions with different dilation rates, effectively expanding the receptive field while preserving detailed information.
[0107]
[0108] Where d = r1, r2, and r3 represent different dilation rates, with the default value set to [1, 2, 3]. This design allows the model to perceive contextual information at different scales simultaneously, making it particularly effective for detecting vehicles at different distances. Compared to standard convolution, dilated convolution significantly expands the receptive field without increasing the number of parameters or computational cost.
[0109] Receptive field = k + (k-1) × (d-1) (3.22)
[0110] Where k is the kernel size and d is the dilation rate. When k = 3, the receptive fields of the three parallel branches are 3, 5 and 7, respectively, covering a variety of scales commonly used in vehicle detection.
[0111] ②CSP-MDC structure: Combines the design principles of CSP networks with MDC modules to achieve more efficient feature extraction.
[0112]
[0113] ③ Dual-path feature pyramid structure: MDC-APN adopts a dual-path design, including two parallel feature extraction paths: upsampling and downsampling.
[0114]
[0115] Among them, X high and X low These represent high-level and low-level features, respectively. Gating represents the gating mechanism, which is used to adaptively select effective features.
[0116] ④ Gated Feature Selection Mechanism: To suppress redundant features and enhance the representation of effective features, MDC-APN introduces a feature gating mechanism:
[0117] G=σ(Conv1×1 (Concat[F sampled ,F target (3.25)
[0118] F gated =G⊙F sampled +(1-G)⊙F target (3.26)
[0119] Among them, F sampled F represents the sampled features. target ⊙ represents the target layer features;
[0120] S34, Training the improved RT-DETR model:
[0121] The improved RT-DETR algorithm is trained using a two-stage strategy. The first stage involves pre-training on the MS COCO dataset with a batch size of 4 and an initial learning rate of 1×10⁻⁶. -4 The first stage employs a cosine annealing decay strategy, using the AdamW optimizer, and trains for 300 epochs at a 640×640 pixel resolution to provide good initialization parameters for the model. The second stage uses the vehicle detection dataset constructed in step S2 for domain fine-tuning, with a batch size of 4 and an initial learning rate of 5×10⁻⁶. -5 The model was warmed up for 10 rounds and trained for 200 rounds using the AdamW optimizer at a resolution of 1024×1024 pixels to adapt it to the characteristics of traffic scenarios.
[0122] Furthermore, in step S4, based on the improved RT-DETR (Real-Time Detection Transformer) algorithm, combined with binocular vision depth perception and the StrongSORT multi-target tracking algorithm, the method for achieving real-time vehicle speed and distance measurement is as follows:
[0123] A binocular camera captures the parallax of the same scene through its left and right lenses, and calculates the target distance based on triangulation. The formula is as follows:
[0124]
[0125] Where Z is the target distance (meters), f is the camera focal length (pixels), B is the binocular baseline distance (meters), and d is the disparity (pixels) of the matching points in the left and right images.
[0126] The disparity map is calculated using the SGBM (semi-global block matching) algorithm. After generating the depth map, the depth value of the center point of the vehicle bounding box is extracted. Combined with calibration parameters, dynamic calibration is performed to ensure ranging accuracy.
[0127] In terms of velocity measurement, the detection box input of RT-DETR will be improved into the StrongSORT algorithm. Kalman filtering will be used to predict the target's motion state, and cross-frame target matching will be performed using appearance features (ReID model) and motion features (Mahathano distance). Instantaneous velocity will be calculated based on the positional changes (ΔX, ΔY) and time intervals (Δt) of the same target in consecutive frames.
[0128]
[0129] Instantaneous jitter is eliminated by using a sliding window averaging method (window size = 5 frames), and the lateral speed error when the vehicle turns is corrected by combining a heading angle compensation model, and finally a smooth speed value is output.
[0130] Furthermore, in step S5, based on the vehicle position and speed information obtained in step S4, the vehicle trajectory is predicted and analyzed. By calculating the probability of the predicted vehicle trajectory intersecting with the traffic cone protection area, the collision risk index (CRI) is calculated. The method for implementing multi-level early warning based on the collision risk index is as follows:
[0131] S51, Define the vehicle trajectory and trajectory prediction problem:
[0132] A vehicle trajectory is defined as a sequence of position coordinates of a vehicle over a continuous period of time; for a trajectory containing n points, its formal representation is:
[0133] T = {Xi, Yi: i ∈ [1...n]} (5.1)
[0134] The mapping relationship between trajectory point index i and timestamp satisfies:
[0135] time(i)-time(j)=K×(ij) (5.2)
[0136] In the formula, K is a fixed time step;
[0137] The trajectory prediction problem is defined as follows: given a sequence of trajectories of m observed points (Xi, Yi, i∈[1...m]), predict the value of the trajectory at a future time step (Xi, Yi, i∈[m+1...n]) such that the error between the predicted trajectory and the actual trajectory t∈T is minimized.
[0138] S52, a multimodal trajectory generation model based on GAN:
[0139] To address the unique challenges of vehicle trajectory prediction in road construction areas, where multiple possible travel paths exist, an improved Generative Adversarial Network (GAN) structure is employed to achieve multimodal trajectory prediction. This structure comprises a generator network and a discriminator network, as detailed below:
[0140] ① The generator network consists of two sub-networks, and its architecture is as follows:
[0141] First sub-network: used to process the observed trajectory, containing an LSTM layer with 32 neurons, a second LSTM layer with 16 neurons, and a dense layer with 16 neurons;
[0142] The second subnetwork receives the output of the first subnetwork and a 2-dimensional latent vector, and contains a dense layer of 16 neurons and an output layer.
[0143] The mathematical expression for a generator is:
[0144] G(x,z)=f2(concat[f1(x),z]) (5.3)
[0145] Where x is the observed trajectory, z is the potential vector, and f1 and f2 represent the mapping functions of the two sub-networks, respectively;
[0146] ② The discriminator network consists of three sub-networks, and its architecture is as follows:
[0147] First subnetwork: Used to evaluate the rationality of the generated trajectory, containing an LSTM layer with 64 neurons and a dense layer with 32 neurons;
[0148] The second sub-network is used to evaluate the matching degree between the generated trajectory and the observed trajectory, and contains an LSTM layer with 64 neurons and a dense layer with 32 neurons.
[0149] The third subnetwork is used to integrate the outputs of the first two subnetworks and consists of three dense layers with 32, 16, and 1 neurons respectively.
[0150] The mathematical expression of the discriminator is:
[0151] D(x,y)=f3(concat[f1(y),f2(concat[x,y])]) (5.4)
[0152] Where x is the observed trajectory, y is the generated trajectory or the true trajectory, and f1, f2 and f3 represent the mapping functions of the three sub-networks, respectively;
[0153] S53, Multimodal Behavior Modeling and Differentiation Guarantee:
[0154] To ensure the diversity and effectiveness of predicted trajectories, a multimodal behavior differentiation constraint mechanism is introduced:
[0155] ① Behavioral Differentiation Criteria: Ensure that the generated K trajectories (K=3 in this system) represent different behavioral patterns, satisfying:
[0156]
[0157] Where Ei and Ej are the normalized average displacement errors (N-ADE) of trajectories i and j, and T is the threshold parameter (set to 0.5);
[0158] ② Differentiated trajectory generation: Differentiated trajectories are generated using the following algorithm:
[0159] K predicted trajectories are generated sequentially. For each newly generated trajectory, the N-ADE ratio with the already generated trajectory is calculated. If the ratio is higher than the threshold (1-T), the trajectory is accepted; otherwise, it is regenerated. A maximum of 100 generation attempts are made to ensure algorithm efficiency.
[0160] S54, Trajectory Prediction Evaluation Metrics:
[0161] The following metrics are used to evaluate the quality of trajectory predictions:
[0162] Average displacement error (ADE): The root mean square error of all corresponding points between the predicted trajectory and the true trajectory;
[0163]
[0164] Final Displacement Error (FDE): The error between the final point of the predicted trajectory and the actual trajectory;
[0165]
[0166] Normalized average displacement error (N-ADE): A standardized evaluation metric that takes into account the effects of trajectory length and variance;
[0167]
[0168] Normalized final displacement error (N-FDE):
[0169]
[0170] Where l(t) is the trajectory length (meters), v(t) is the trajectory variance, and Kl and Kv are constants (set to 0.04 and 0.003 respectively);
[0171] S55, Early Warning Decision-Making:
[0172] ① Input the vehicle detection results obtained in step S4 into the improved GAN model and generate three possible future trajectories. Analyze the probability of each predicted trajectory intersecting with the traffic cone area and calculate the Collision Risk Index (CRI):
[0173] CRI = max(P) i ×R i ×V i (5.10)
[0174] in:
[0175] P i Let be the probability of trajectory i, output by the GAN multimodal trajectory prediction model, representing the probability of this behavior pattern, with a normalized range of [0,1].
[0176] R i As a risk factor, considering both spatial overlap and time urgency: R i =IoU×exp(-t / τ); where IoU is the intersection-exchange ratio of the vehicle trajectory and the protected area, t is the predicted collision time (seconds), and τ is the time constant;
[0177] Vi is the vehicle speed factor, which considers the impact of vehicle speed on collision risk: V i =min(v / v) ref ,1), where v is the current speed of the vehicle, v ref For reference speed;
[0178] The cone protection zone adopts a dynamic design, and its radius is calculated using the following formula:
[0179] R b =R b +K×V (5.11)
[0180] Among them, R p R is the radius (in meters) of the protected area. b The base protection radius is 1 meter (default), K is the speed coefficient (default 0.05), and V is the vehicle speed (km / h).
[0181] ② Determine the warning level based on the calculated CRI value:
[0182] CRI < 0.3: Safe, no warning required.
[0183] 0.3≤CRI<0.6: Potential risk, issue Level 1 warning.
[0184] 0.6≤CRI<0.8: Moderate risk, issue a level 2 warning.
[0185] CRI≥0.8: High risk, emergency warning issued.
[0186] The system recalculates the CRI value every 100ms to ensure the timeliness and accuracy of the warning.
[0187] Through this multi-factor-based dynamic risk assessment mechanism, the system can adapt to complex and ever-changing road construction environments, provide more accurate and reliable collision warnings, and effectively protect the safety of construction workers.
[0188] Compared with the prior art, the present invention has the following beneficial effects:
[0189] This invention proposes an intelligent traffic cone collision avoidance and early warning method based on binocular vision, aiming to provide timely and accurate collision warning information for road construction workers and improve the safety of road construction areas. This invention employs binocular camera calibration technology based on circle detection to achieve high-precision calibration of the binocular camera, laying the foundation for subsequent detection and analysis and ensuring the accuracy and reliability of the data. Then, through data acquisition and data augmentation techniques, high-quality and diverse training and testing data are provided for the improved RT-DETR algorithm, enhancing the model's generalization ability and robustness. This invention integrates a RepFaster feature extraction module, a HiLo attention mechanism, and a multi-scale dilated convolutional attention pyramid network, significantly improving the accuracy and efficiency of vehicle recognition and enabling precise detection of various vehicle targets. Based on the detection results of the improved RT-DETR algorithm, combined with binocular vision depth perception and the StrongSORT multi-target tracking algorithm, real-time speed and distance measurement of vehicles are achieved, providing accurate input data for trajectory prediction and early warning algorithms. Finally, GAN multimodal trajectory prediction technology is used to generate multiple future trajectories representing different behavioral patterns, and multi-level early warning decisions are achieved through the comprehensive collision risk index (CRI). This technology not only significantly reduces the false alarm rate but also dramatically improves early warning accuracy, achieving a technological leap from simple perception to intelligent prediction. This invention effectively addresses the limitations of existing prediction methods in complex road scenarios by using multimodal trajectory prediction and analysis of vehicle behavior intentions to ensure reliable early warning information under various complex road conditions. Attached Figure Description
[0190] Figure 1 This is an overall flowchart of the intelligent traffic cone collision avoidance and early warning method based on binocular vision of the present invention;
[0191] Figure 2 Flowchart for binocular camera calibration;
[0192] Figure 3 Create a flowchart for the vehicle inspection dataset;
[0193] Figure 4 The structure diagram for improving the RT-DETR (Real-Time Detection Transformer) algorithm;
[0194] Figure 5 Here is a diagram of the RepFaster module structure;
[0195] Figure 6 This is a block diagram of the Multi-Scale Dilated Convolutional Attention Pyramid Network (MDC-APN);
[0196] Figure 7 This is a structural diagram of the distance and speed measurement module;
[0197] Figure 8 This is a logic diagram for trajectory prediction and early warning. Detailed Implementation
[0198] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings, but it should be understood that the scope of protection of the present invention is not limited to the specific embodiments.
[0199] Example 1
[0200] As attached Figure 1 As shown, the intelligent traffic cone collision avoidance warning method based on binocular vision includes the following steps:
[0201] Step S1: Perform high-precision calibration of the binocular camera based on circle detection. (See appendix) Figure 2 The calibration process includes the following steps:
[0202] S11, Calibration board selection and image acquisition:
[0203] A 7×7 circular array calibration board with a solid circle diameter of 20mm and a center-to-center distance of 40mm between adjacent circles is used, with a printing accuracy error of ≤0.01mm. A binocular camera is mounted on an intelligent cone to synchronously acquire images of the calibration board in 20 different poses (horizontal tilt ±45°, vertical tilt ±30°, evenly distributed at a distance of 0.5-3m).
[0204] S12, Binocular camera calibration:
[0205] ① Accurate extraction of the centroid of the ellipse: First, for the calibration board image captured by the stereo camera, an improved image processing technique is used to extract the centroid of the solid circle ellipse; bilateral filtering is used instead of traditional Gaussian filtering for image preprocessing, which can maintain clear edges while effectively removing noise; the bilateral filtering expression is as follows:
[0206]
[0207] After filtering, improved Canny edge detection is performed, calculating the image gradient magnitude G and gradient direction θ:
[0208]
[0209] θ = arctan(G) y / G x (1.3)
[0210] Finally, the sub-pixel precise coordinates of the ellipse's centroid are obtained by calculating the geometric moments:
[0211]
[0212] Wherein, geometric moment m pqDefined as:
[0213]
[0214] ② Parameter optimization combining Zhang's calibration and bundled adjustment: Based on the extracted centroid coordinates of the ellipse, the initial values of the camera parameters are determined using Zhang's calibration method; Zhang's method is based on the homography relationship between the planar calibration plate and the image.
[0215]
[0216] Where H is the homography matrix:
[0217] H=[h1 h2 h3]=λA[r1 r2t] (1.8) A is the camera intrinsic parameter matrix, which includes parameters such as focal length and optical center position;
[0218] Solving for camera parameters by minimizing the reprojection error:
[0219]
[0220] Then, a bundled adjustment optimization algorithm is introduced to construct the normal equation:
[0221]
[0222] By solving, both camera parameters and 3D point coordinates are optimized:
[0223]
[0224] X = N 22 -1 (W2-N 21 t) (1.12)
[0225] ③ Extrinsic parameter optimization based on diagonal constraints: To further improve the system's measurement accuracy, the diagonal length of the calibration plate is used as a known quantity for extrinsic parameter optimization; the diagonal length is calculated as follows:
[0226]
[0227] Introduce the proportionality coefficient α = 1 / t x get:
[0228]
[0229] The formula for calculating the proportionality coefficient α is:
[0230]
[0231] The calibration plate has two diagonals, denoted as L. l and L r The corresponding translation vector is:
[0232]
[0233] Taking into account the constraints of the two diagonals, the extrinsic parameter matrix is finally optimized as follows:
[0234] T = (T l +T r ) / 2 (1.17)
[0235] Step S2: Collect video data from the road construction area, preprocess, label, enhance, and segment the data to construct a vehicle detection dataset. (See appendix) Figure 3 The specific method is as follows:
[0236] S21, Data Acquisition:
[0237] By deploying high-definition binocular cameras in the road construction area, video data with a resolution of 1080P and a frame rate of 30fps is collected to ensure coverage of scenes under different time periods, weather, and lighting conditions. At the same time, construction areas, road sections with cones and warning signs, and complex vehicle interaction scenarios covered by different weather and lighting conditions are selected from BDD100K, Cityscapes, and KITTI datasets and integrated into the vehicle detection dataset to enrich the diversity of the data.
[0238] S22, Data Preprocessing:
[0239] Spatiotemporal sampling was performed on the collected video footage. For typical scenes with stable traffic flow, normal vehicle speed, and good lighting conditions, 3 frames per second were extracted. For special event scenes such as sudden vehicle deceleration or lane change, vehicles approaching the boundary of a construction area, severe weather conditions, or insufficient lighting, 10 frames per second were extracted as training data. Subsequently, the extracted images underwent enhancement processing, including exposure correction, color balance, noise reduction, and sharpening, to improve image quality. Finally, all training images were uniformly scaled to a resolution of 1280×720 pixels and normalized to ensure data standard consistency.
[0240] S23, Data Labeling:
[0241] The processed images are labeled to mark the precise location and category information of various vehicles in the images;
[0242] S24, Data Augmentation:
[0243] By using traditional image enhancement techniques such as flipping, rotating, scaling, cropping, and brightness adjustment, the dataset size is expanded and the sample diversity is increased. At the same time, DCGAN (Deep Convolutional Generative Adversarial Network) technology is applied to synthesize rare scene samples in reality, such as extreme weather and special lighting, to improve the model's adaptability in complex environments.
[0244] S25, Dataset Partitioning
[0245] The constructed labeled images are divided into training set:validation set:test set in a ratio of 7:2:1. This ensures that the model can fully learn the sample features during training, effectively verify the training effect, and objectively evaluate the final performance.
[0246] Step S3: In actual vehicle detection scenarios, there are many complex and highly uncertain factors such as lighting differences, motion blur, occlusion, and weather. To reduce the impact of these uncertainties and improve the accuracy and real-time performance of vehicle detection, an improved RT-DETR (Real-Time Detection Transformer) algorithm is adopted and trained. The structure diagram of the improved RT-DETR (Real-Time Detection Transformer) algorithm is attached. Figure 4 As shown, the method for improving the algorithm and training it is as follows:
[0247] S31, Improved Backbone Network Design:
[0248] ① To improve the feature extraction capability, model accuracy, and robustness of the vehicle detection model, while optimizing computational efficiency, maintaining low latency, and high throughput, a new feature extraction module, the RepFaster module, is designed by combining Faster-Block and RepConv from FasterNet. Its structure is shown in the attached figure. Figure 5 As shown, this module reduces computational redundancy and memory access, improving computational efficiency by combining partial convolution and pointwise convolution; it also enhances the model's ability to capture complex features by reparameterizing multiple convolution kernels. Specifically, the RepFaster module employs a multi-branch structure during training, including a standard 3×3 convolution branch and a pointwise 1×1 convolution branch, as shown below:
[0249] F out =W 3×3 *+W 1×1 *X+b (3.1)
[0250] Among them, W 3×3 W represents the parameters of the 3×3 convolution kernel. 1×1 Let X represent the parameters of a 1×1 convolution kernel, X be the input feature map, and b be the bias term. During the inference phase, multi-branch structures are merged into a single convolution operation through parameter equivalence.
[0251] W equiv =W 3×3 +expand(W 1×1 (3.2)
[0252] bequiv =b (3.3)
[0253] F out =W equiv *X+b equiv (3.4)
[0254] Among them, expand(W) 1×1 The ) indicates that the parameters of the 1×1 convolution kernel are expanded to make it compatible with the 3×3 convolution kernel; this design allows the model to have different structures during the training and inference phases, effectively balancing the model's expressive power and computational efficiency;
[0255] In partial convolution, the input feature X is divided into n div Groups are formed by processing each group of features separately and then merging them, as shown below:
[0256]
[0257]
[0258] This design significantly reduces computational cost while maintaining feature extraction capabilities;
[0259] ② Integrate the RepFaster module into the basic modules of the ResNet backbone network to form a new backbone network; the traditional ResNet basic block is represented as:
[0260] F out =F(x)+x (3.7)
[0261] F(x)=W2*σ(W1*x) (3.8)
[0262] Where σ is the activation function, and W1 and W2 are convolution parameters; the improved structure replaces W2 with the RepFaster module:
[0263] F(x)=RepFaster(σ(W1*x)) (3.9)
[0264] This combination leverages ResNet's residual learning capabilities and RepFaster's efficient feature extraction capabilities, making it more suitable for vehicle detection tasks.
[0265] S32, Improved intra-scale feature interaction design:
[0266] To adapt to the complex road environment and real-time requirements of traffic cone collision avoidance scenarios, and to further enhance the model's ability to capture vehicle features, especially the simultaneous perception of vehicles at different distances, a HiLo attention mechanism (High-Low Frequency Attention) is introduced into the Transformer encoder layer to improve the original network's AIFI (Attention-based Intra-scale Feature Interaction) module. The HiLo attention mechanism, by separating high-frequency and low-frequency feature processing, more efficiently captures local details and global contextual information in vehicle detection. The processing details of the HiLo attention mechanism are as follows:
[0267] ① By capturing local details such as vehicle outlines and logos through high-frequency branching (HiFi), the problem of missed detection under occlusion and motion blur in traditional methods is solved. The specific implementation details are as follows:
[0268] First, the feature map is segmented into multiple local regions using window partitioning:
[0269] X windows =reshape(X,[B,h) group ,window_size,w group ,window_size,C]) (3.10)
[0270] Among them, h group =H / / window_size,w group =W / / window_size, B is the batch size, H and W are the height and width of the feature map, C is the number of channels, and the window size is s×s;
[0271] Perform linear projection on the divided local regions:
[0272] Q h ,K h V h =linear projection (X windows (3.11)
[0273] in, d is the dimension of each attention head;
[0274] Calculate attention score:
[0275]
[0276] Weighted summation:
[0277] HiFiout = Attention h ·V h(3.13)
[0278] The window features are recombined into a complete feature map by inverse window partitioning:
[0279] X HiFi =WindowReverse(HiFiout) (3.14)
[0280] ② By modeling global context information such as road scenes and relative vehicle positions through the low-frequency branch (LoFi), the problem of missed detections and false detections caused by insufficient modeling of global context information in traditional methods is solved. The specific implementation details are as follows:
[0281] First, downsampling is performed using average pooling to reduce the resolution:
[0282] X down =AvgPool2d(X,kernel_size=window_size) (3.15)
[0283] Where AvgPool2d represents the average pooling operation;
[0284] Perform linear projection:
[0285] Q l =linear_projection q (X) (3.16)
[0286] K l V l =linear_projection kv (X down (3.17)
[0287] in, H down =H / / window size W down =W / / window size ;
[0288] Calculate attention score:
[0289]
[0290] Weighted summation:
[0291] LoFi out =Attention l ·V l (3.19)
[0292] ③ Feature fusion: High-frequency and low-frequency features are concatenated along the channel dimension to form a complete output feature.
[0293] Output = concat([X HiFi ,X LoFi ],dim=C) (3.20)
[0294] By using the HiLo attention mechanism, the model can maintain computational efficiency while paying attention to local details and global context, significantly improving vehicle detection performance in complex traffic scenarios, especially in scenarios where vehicles at multiple scales coexist.
[0295] S33, Improved Feature Fusion Module:
[0296] To further enhance the RT-DETR algorithm's ability to detect vehicles at different scales, especially its accuracy in recognizing distant, small vehicles and partially occluded vehicles, this invention proposes a novel feature fusion module, "Multi-Scale Dilated Convolutional Attention Pyramid Network (MDC-APN)," to replace the original feature fusion module CCFM. The structure of this MDC-APN is shown in the attached figure. Figure 6 As shown, it has the following key design features:
[0297] ① Multi-scale dilated convolution module (MDC): Processes features in parallel through dilated convolutions with different dilation rates, effectively expanding the receptive field while preserving detailed information.
[0298]
[0299] Where d = r1, r2, and r3 represent different dilation rates, with the default value set to [1, 2, 3]. This design allows the model to perceive contextual information at different scales simultaneously, making it particularly effective for detecting vehicles at different distances. Compared to standard convolution, dilated convolution significantly expands the receptive field without increasing the number of parameters or computational cost.
[0300] Receptive field = k + (k-1) × (d-1) (3.22)
[0301] Where k is the kernel size and d is the dilation rate. When k = 3, the receptive fields of the three parallel branches are 3, 5 and 7, respectively, covering a variety of scales commonly used in vehicle detection.
[0302] ②CSP-MDC structure: Combines the design concept of CSP (Cross Stage Partial) network with MDC module to achieve more efficient feature extraction:
[0303]
[0304] This method of extracting features from only some channels reduces computation by about 50% while maintaining model performance.
[0305] ③ Dual-path feature pyramid structure: MDC-APN adopts a dual-path design, including two parallel feature extraction paths: upsampling and downsampling.
[0306]
[0307] Among them, X high and X low These represent high-level and low-level features, respectively. Gating represents the gating mechanism, which is used to adaptively select effective features. This dual-path design enables the network to utilize both high-level semantic information and low-level detailed information simultaneously, enhancing its ability to detect small and large vehicles.
[0308] ④ Gated Feature Selection Mechanism: To suppress redundant features and enhance the representation of effective features, MDC-APN introduces a feature gating mechanism:
[0309] G=σ(Conv 1×1 (Concat[F sampled ,F target (3.25)
[0310] F gated =G⊙F sampled +(1-G)⊙F target (3.26)
[0311] Among them, F sampled F represents the sampled features. target ⊙ represents the target layer features, and ⊙ represents element-wise multiplication. This gating mechanism can adaptively adjust the weights of different features according to their importance, improving the effectiveness of feature representation. It is particularly suitable for handling complex scenarios in vehicle detection.
[0312] S34, Training the improved RT-DETR model:
[0313] The improved RT-DETR algorithm is trained using a two-stage strategy. The first stage involves pre-training on the MS COCO dataset with a batch size of 4 and an initial learning rate of 1×10⁻⁶. -4 The first stage employs a cosine annealing decay strategy, using the AdamW optimizer, and trains for 300 epochs at a 640×640 pixel resolution to provide good initialization parameters for the model. The second stage uses the vehicle detection dataset constructed in step S2 for domain fine-tuning, with a batch size of 4 and an initial learning rate of 5×10⁻⁶. -5 The model underwent 10 warm-up rounds and 200 training rounds using the AdamW optimizer at a resolution of 1024×1024 pixels to adapt it to the characteristics of traffic scenarios. Throughout the training process, optimization techniques such as gradient accumulation, exponential moving average, mixed precision training, and progressive learning were combined to improve training efficiency and model stability.
[0314] Step S4: Based on the improved RT-DETR (Real-Time Detection Transformer) algorithm, combined with binocular vision depth perception and StrongSORT multi-target tracking algorithm, real-time vehicle speed and distance measurement are achieved (the module structure diagram for distance and speed measurement is attached). Figure 7 As shown in the figure, this provides accurate input data for trajectory prediction and early warning. The specific steps are as follows:
[0315] A binocular camera captures the parallax of the same scene through its left and right lenses, and calculates the target distance based on triangulation. The formula is as follows:
[0316]
[0317] Where Z is the target distance (meters), f is the camera focal length (pixels), B is the binocular baseline distance (meters), and d is the disparity (pixels) of the matching points in the left and right images.
[0318] The disparity map is calculated using the SGBM (semi-global block matching) algorithm. After generating the depth map, the depth value of the center point of the vehicle bounding box is extracted. Combined with calibration parameters, dynamic calibration is performed to ensure ranging accuracy.
[0319] In terms of velocity measurement, the detection box input of RT-DETR will be improved into the StrongSORT algorithm. Kalman filtering will be used to predict the target's motion state, and cross-frame target matching will be performed using appearance features (ReID model) and motion features (Mahathano distance). Instantaneous velocity will be calculated based on the positional changes (ΔX, ΔY) and time intervals (Δt) of the same target in consecutive frames.
[0320]
[0321] Instantaneous jitter is eliminated by using a sliding window averaging method (window size = 5 frames), and the lateral speed error when the vehicle turns is corrected by combining a heading angle compensation model, and finally a smooth speed value is output.
[0322] Step S5: Based on the vehicle position and speed information obtained in step S4, predictive analysis of the vehicle trajectory is performed. By calculating the probability of the predicted vehicle trajectory intersecting with the traffic cone protection area, the collision risk index (CRI) is calculated. Multi-level early warning is implemented based on the collision risk index. The specific method is as follows:
[0323] S51, Define the vehicle trajectory and trajectory prediction problem:
[0324] A vehicle trajectory is defined as a sequence of position coordinates of a vehicle over a continuous period of time; for a trajectory containing n points, its formal representation is:
[0325] T = {Xi, Yi: i ∈ [1...n]} (5.1)
[0326] The mapping relationship between trajectory point index i and timestamp satisfies:
[0327] time(i)-time(j)=K×(ij) (5.2)
[0328] In the formula, K is a fixed time step;
[0329] The trajectory prediction problem is defined as follows: given a sequence of trajectories of m observed points (Xi, Yi, i∈[1...m]), predict the value of the trajectory at a future time step (Xi, Yi, i∈[m+1...n]) such that the error between the predicted trajectory and the actual trajectory t∈T is minimized.
[0330] S52, a multimodal trajectory generation model based on GAN:
[0331] To address the unique challenges of vehicle trajectory prediction in road construction areas, where multiple possible travel paths exist, an improved Generative Adversarial Network (GAN) structure is employed to achieve multimodal trajectory prediction. This structure comprises a generator network and a discriminator network, as detailed below:
[0332] ① The generator network consists of two sub-networks, and its architecture is as follows:
[0333] First sub-network: used to process the observed trajectory, containing an LSTM layer with 32 neurons, a second LSTM layer with 16 neurons, and a dense layer with 16 neurons;
[0334] The second subnetwork receives the output of the first subnetwork and a 2-dimensional latent vector, and contains a dense layer of 16 neurons and an output layer.
[0335] The mathematical expression for a generator is:
[0336] G(x,z)=f2(concat[f1(x),z]) (5.3)
[0337] Where x is the observed trajectory, z is the potential vector, and f1 and f2 represent the mapping functions of the two sub-networks, respectively;
[0338] ② The discriminator network consists of three sub-networks, and its architecture is as follows:
[0339] First subnetwork: Used to evaluate the rationality of the generated trajectory, containing an LSTM layer with 64 neurons and a dense layer with 32 neurons;
[0340] The second sub-network is used to evaluate the matching degree between the generated trajectory and the observed trajectory, and contains an LSTM layer with 64 neurons and a dense layer with 32 neurons.
[0341] The third subnetwork is used to integrate the outputs of the first two subnetworks and consists of three dense layers with 32, 16, and 1 neurons respectively.
[0342] The mathematical expression of the discriminator is:
[0343] D(x,y)=f3(concat[f1(y),f2(concat[x,y])]) (5.4)
[0344] Where x is the observed trajectory, y is the generated trajectory or the true trajectory, and f1, f2 and f3 represent the mapping functions of the three sub-networks, respectively;
[0345] S53, Multimodal Behavior Modeling and Differentiation Guarantee:
[0346] To ensure the diversity and effectiveness of predicted trajectories, a multimodal behavior differentiation constraint mechanism is introduced:
[0347] ① Behavioral Differentiation Criteria: Ensure that the generated K trajectories (K=3 in this system) represent different behavioral patterns, satisfying:
[0348]
[0349] Where Ei and Ej are the normalized average displacement errors (N-ADE) of trajectories i and j, and T is the threshold parameter (set to 0.5);
[0350] ② Differentiated trajectory generation: Differentiated trajectories are generated using the following algorithm:
[0351] K predicted trajectories are generated sequentially. For each newly generated trajectory, the N-ADE ratio with the already generated trajectory is calculated. If the ratio is higher than the threshold (1-T), the trajectory is accepted; otherwise, it is regenerated. A maximum of 100 generation attempts are made to ensure algorithm efficiency.
[0352] S54, Trajectory Prediction Evaluation Metrics:
[0353] The following metrics are used to evaluate the quality of trajectory predictions:
[0354] Average displacement error (ADE): The root mean square error of all corresponding points between the predicted trajectory and the true trajectory;
[0355]
[0356] Final Displacement Error (FDE): The error between the final point of the predicted trajectory and the actual trajectory;
[0357]
[0358] Normalized average displacement error (N-ADE): A standardized evaluation metric that takes into account the effects of trajectory length and variance;
[0359]
[0360] Normalized final displacement error (N-FDE):
[0361]
[0362] Where l(t) is the trajectory length (meters), v(t) is the trajectory variance, and Kl and Kv are constants (set to 0.04 and 0.003 respectively);
[0363] S55, Early Warning Decision-Making:
[0364] ① Input the vehicle detection results obtained in step S4 into the improved GAN model and generate three possible future trajectories. Analyze the probability of each predicted trajectory intersecting with the traffic cone area and calculate the Collision Risk Index (CRI):
[0365] GRI = max(P) i ×R i ×V i (5.10)
[0366] in:
[0367] P i Let be the probability of trajectory i, output by the GAN multimodal trajectory prediction model, representing the probability of this behavior pattern, with a normalized range of [0,1].
[0368] R i As a risk factor, considering both spatial overlap and time urgency: R i =IoU×Exp(-t / τ); where IoU is the intersection-exchange ratio of the vehicle trajectory and the protected area, t is the predicted collision time (seconds), and τ is the time constant;
[0369] Vi is the vehicle speed factor, which considers the impact of vehicle speed on collision risk: V i =min(v / v) ref ,1), where v is the current speed of the vehicle, v ref For reference speed;
[0370] The cone protection zone adopts a dynamic design, and its radius is calculated using the following formula:
[0371] R p =R b +K×V (5.11)
[0372] Among them, R p R is the radius (in meters) of the protected area.b The base protection radius is 1 meter (default), K is the speed coefficient (default 0.05), and V is the vehicle speed (km / h).
[0373] ② Determine the warning level based on the calculated CRI value:
[0374] CRI < 0.3: Safe, no warning required.
[0375] 0.3≤CRI<0.6: Potential risk, issue Level 1 warning.
[0376] 0.6≤CRI<0.8: Moderate risk, issue a level 2 warning.
[0377] CRI≥0.8: High risk, emergency warning issued.
[0378] The system recalculates the CRI value every 100ms to ensure the timeliness and accuracy of warnings. The logic for trajectory prediction and warning is as follows: Figure 8 As shown.
[0379] Through this multi-factor-based dynamic risk assessment mechanism, the system can adapt to complex and ever-changing road construction environments, provide more accurate and reliable collision warnings, and effectively protect the safety of construction workers.
Claims
1. A method for intelligent traffic cone collision avoidance and early warning based on binocular vision, characterized in that, Includes the following steps: Step S1: Perform high-precision calibration of the binocular camera based on circle detection; Step S2: Collect video data of the road construction area, preprocess, label, enhance and segment the data to construct a vehicle detection dataset; Step S3: Improve the RT-DETR algorithm and train it; Step S4: Based on the improved RT-DETR algorithm, combined with binocular vision depth perception and StrongSORT multi-target tracking algorithm, real-time vehicle speed measurement and distance measurement are achieved. Step S5: Based on the vehicle position and speed information obtained in step S4, predictive analysis of the vehicle trajectory is performed. By calculating the probability of the vehicle's predicted trajectory intersecting with the traffic cone protection area, a collision risk index is calculated, and multi-level early warning is implemented based on the collision risk index.
2. The intelligent traffic cone collision avoidance and early warning method according to claim 1, characterized in that, In step S1, the method for high-precision calibration of a stereo camera based on circle detection includes the following steps: S11, Calibration board selection and image acquisition: A 7×7 circular array calibration board with a solid circle diameter of 20mm and a center-to-center distance of 40mm between adjacent circles is used, with a printing accuracy error ≤0.01mm; a binocular camera is mounted on a smart cone to synchronously acquire images of the calibration board in 20 different poses; S12, Binocular camera calibration: ① Accurate extraction of the centroid of the ellipse: First, for the calibration board image captured by the stereo camera, an improved image processing technique is used to extract the centroid of the solid circle ellipse; bilateral filtering is used instead of traditional Gaussian filtering for image preprocessing, which can maintain clear edges while effectively removing noise; the bilateral filtering expression is as follows: After filtering, improved Canny edge detection is performed, calculating the image gradient magnitude G and gradient direction θ: θ=arctane(G y / G x ) (1.3) Finally, the sub-pixel precise coordinates of the ellipse's centroid are obtained by calculating the geometric moments: Wherein, geometric moment m pq Defined as: ② Parameter optimization combining Zhang's calibration and bundled adjustment: Based on the extracted centroid coordinates of the ellipse, the initial values of the camera parameters are determined using Zhang's calibration method; Zhang's method is based on the homography relationship between the planar calibration plate and the image. Where H is the homography matrix: H=[h1 h2 h3]=λA[r1 r2 t] (1.8) A is the camera intrinsic parameter matrix, which includes parameters such as focal length and optical center position; Solving for camera parameters by minimizing the reprojection error: Then, a bundled adjustment optimization algorithm is introduced to construct the normal equation: By solving, both camera parameters and 3D point coordinates are optimized: X=N 22 -1 (W2-N 21 t) (1.12) ③ Extrinsic parameter optimization based on diagonal constraints: To further improve the system's measurement accuracy, the diagonal length of the calibration plate is used as a known quantity for extrinsic parameter optimization; the diagonal length is calculated as follows: Introduce the proportionality coefficient α = 1 / t x get: The formula for calculating the proportionality coefficient α is: The calibration plate has two diagonals, denoted as L. l and L r The corresponding translation vector is: Taking into account the constraints of the two diagonals, the extrinsic parameter matrix is finally optimized as follows: T=(T l +T r ) / 2 (1.17) 3. The intelligent traffic cone collision avoidance and early warning method according to claim 1, characterized in that, In step S2, video data of the road construction area is collected, and the data is preprocessed, labeled, enhanced, and segmented to construct a vehicle detection dataset. S21, Data Acquisition: By deploying high-definition binocular cameras in the road construction area, video data with a resolution of 1080P and a frame rate of 30fps is collected to ensure coverage of scenes under different time periods, weather, and lighting conditions. At the same time, construction areas, road sections with cones and warning signs, and complex vehicle interaction scenarios covered by different weather and lighting conditions are selected from BDD100K, Cityscapes, and KITTI datasets and integrated into the vehicle detection dataset to enrich the diversity of the data. S22, Data Preprocessing: Spatiotemporal sampling was performed on the collected video footage. For typical scenes with stable traffic flow, normal vehicle speed, and good lighting conditions, 3 frames per second were extracted. For special event scenes such as sudden vehicle deceleration or lane change, vehicles approaching the boundary of a construction area, severe weather conditions, or insufficient lighting, 10 frames per second were extracted as training data. Subsequently, the extracted images underwent enhancement processing, including exposure correction, color balance, noise reduction, and sharpening, to improve image quality. Finally, all training images were uniformly scaled to a resolution of 1280×720 pixels and normalized to ensure data standard consistency. S23, Data Labeling: The processed images are labeled to mark the precise location and category information of various vehicles in the images; S24, Data Augmentation: The dataset size is expanded and sample diversity is increased by using traditional image enhancement techniques such as flipping, rotating, scaling, cropping and brightness adjustment. At the same time, deep convolutional generative adversarial networks are applied to synthesize rare scene samples in reality, such as extreme weather and special lighting, to improve the model's adaptability in complex environments. S25, Dataset Partitioning The constructed labeled images are divided according to the ratio of training set:validation set:test set = 7:2:
1.
4. The intelligent traffic cone collision avoidance and early warning method according to claim 1, characterized in that, The method for improving and training the RT-DETR algorithm in step S3 is as follows: S31, Improved Backbone Network Design: ① Combining Faster-Block and RepConv from FasterNet, a new feature extraction module is designed: the RepFaster module. This RepFaster module employs a multi-branch structure during training, including a standard 3×3 convolutional branch and a pointwise 1×1 convolutional branch, as shown below: F out =W 3×3 *X+W 1×1 *X+b (3.1) Among them, W 3×3 W represents the parameters of the 3×3 convolution kernel. 1×1 Let X represent the parameters of a 1×1 convolution kernel, X be the input feature map, and b be the bias term. During the inference phase, multi-branch structures are merged into a single convolution operation through parameter equivalence. W equiv =W 3×3 +expand(W 1×1 ) (3.2) b equiv =b (3.3) F out =W equiv *X+b equiv (3.4) Among them, expand(W) 1×1 This indicates that the parameters of a 1×1 convolution kernel are expanded to make it compatible with a 3×3 convolution kernel; In partial convolution, the input feature X is divided into n div Groups are formed by processing each group of features separately and then merging them, as shown below: ② Integrate the RepFaster module into the basic modules of the ResNet backbone network to form a new backbone network; the traditional ResNet basic block is represented as: F out =F(x)+x (3.7) F(x)=W2*σ(W1*x) (3.8) Where σ is the activation function, and W1 and W2 are convolution parameters; the improved structure replaces W2 with the RepFaster module: F(x)=RepFaster(σ(W1*x)) (3.9) S32, Improved intra-scale feature interaction design: The HiLo attention mechanism is introduced into the Transformer encoder layer to improve the AIFI module in the original network. The processing details of the HiLo attention mechanism are as follows: ① By capturing local details such as vehicle outlines and logos through high-frequency branching, the problem of missed detection under occlusion and motion blur in traditional methods is solved. The specific implementation details are as follows: First, the feature map is segmented into multiple local regions using window partitioning: X windows =reshape(X,[B,h group ,window_size,w group ,window_size,C]) (3.10) Among them, h group =H / / window_size,w group =W / / window_size, B is the batch size, H and W are the height and width of the feature map, C is the number of channels, and the window size is s×s; Perform linear projection on the divided local regions: Q h ,K h ,V h =linear projection (X windows ) (3.11) in, d is the dimension of each attention head; Calculate attention score: Weighted summation: HiFiout=Attention h ·V h (3.13) The window features are recombined into a complete feature map by inverse window partitioning: X HiFi =WindowReverse(HiFiout) (3.14) ② By modeling global context information such as road scenes and relative vehicle positions using low-frequency branches, the problem of missed detections and false detections caused by insufficient modeling of global context information in traditional methods is solved. The specific implementation details are as follows: First, downsampling is performed using average pooling to reduce the resolution: X down =AvgPool2d(X,kernel_size=window_size) (3.15) Where AvgPool2d represents the average pooling operation; Perform linear projection: Q l =linear_projection q (X) (3.16) K l ,V l =linear_projection kv (X down ) (3.17) Among them, H down = H / / window size W down = W / / window size ; Calculate attention score: Weighted summation: LoFi out =Attention l ·V l (3.19) ③ Feature fusion: High-frequency and low-frequency features are concatenated along the channel dimension to form a complete output feature. Output=concat([X HiFi ,X LoFi ],dim=C) (3.20) S33, Improved Feature Fusion Module: We designed a "Multi-Scale Dilated Convolutional Attention Pyramid Network (MDC-APN)" to replace the original feature fusion module CCFM. This MDC-APN has the following key design features: ① Multi-scale dilated convolution module (MDC): Processes features in parallel through dilated convolutions with different dilation rates, effectively expanding the receptive field while preserving detailed information. Where d = r1, r2, and r3 represent different dilation rates, with the default value set to [1, 2, 3]. This design allows the model to perceive contextual information at different scales simultaneously, making it particularly effective for detecting vehicles at different distances. Compared to standard convolution, dilated convolution significantly expands the receptive field without increasing the number of parameters or computational cost. Receptive field = k + (k-1) × (d-1) (3.22) Where k is the kernel size and d is the dilation rate. When k = 3, the receptive fields of the three parallel branches are 3, 5 and 7, respectively, covering a variety of scales commonly used in vehicle detection. ②CSP-MDC structure: Combines the design principles of CSP networks with MDC modules to achieve more efficient feature extraction. ③ Dual-path feature pyramid structure: MDC-APN adopts a dual-path design, including two parallel feature extraction paths: upsampling and downsampling. Among them, X high and X low These represent high-level and low-level features, respectively. Gating represents the gating mechanism, which is used to adaptively select effective features. ④ Gated Feature Selection Mechanism: To suppress redundant features and enhance the representation of effective features, MDC-APN introduces a feature gating mechanism: G=σ(Conv 1×1 (Concat[F sampled ,F target ])) (3.25) F gated =G⊙F sampled +(1-G)⊙F target (3.26) Among them, F sampled F represents the sampled features. target ⊙ represents the target layer features; S34, Training the improved RT-DETR model: The improved RT-DETR algorithm training adopts a two-stage strategy. In the first stage, pre-training is performed on the MS COCO dataset with a batch size of 4, an initial learning rate of 1×10⁻⁴, and a cosine annealing decay strategy. The AdamW optimizer is used, and the model is trained for 300 epochs at a resolution of 640×640 pixels to provide good initialization parameters for the model. In the second stage, the vehicle detection dataset constructed in step S2 is used for domain fine-tuning. A batch size of 4 and an initial learning rate of 5×10⁻⁵ are used, and the model is warmed up for 10 epochs. The AdamW optimizer is used, and the model is trained for 200 epochs at a resolution of 1024×1024 pixels to adapt the model to the characteristics of traffic scenarios.
5. The intelligent traffic cone collision avoidance and early warning method according to claim 1, characterized in that, In step S4, based on the improved RT-DETR algorithm, combined with binocular vision depth perception and StrongSORT multi-target tracking algorithm, the method for real-time vehicle speed and distance measurement is as follows: A binocular camera captures the parallax of the same scene through its left and right lenses, and calculates the target distance based on triangulation. The formula is as follows: Where Z is the target distance, f is the camera focal length, B is the binocular baseline distance, and d is the disparity of the matching points in the left and right images; The disparity map is calculated using a semi-global block matching algorithm. After generating the depth map, the depth value of the center point of the vehicle bounding box is extracted. Combined with calibration parameters, the system is dynamically calibrated to ensure ranging accuracy. In terms of velocity measurement, the detection box of RT-DETR will be improved and input into the StrongSORT algorithm. Kalman filtering will be used to predict the target's motion state, and cross-frame target matching will be performed using appearance features and motion features. Appearance features are extracted using the ReID model, while motion features are calculated based on Mahalanobis distance. Instantaneous velocity is calculated based on the positional change (ΔX, ΔY) and time interval (Δt) of the same target in consecutive frames. Instantaneous jitter is eliminated by using the sliding window averaging method, and the lateral speed error during vehicle turning is corrected by combining the heading angle compensation model, and finally a smooth speed value is output.
6. The intelligent traffic cone collision avoidance and early warning method according to claim 1, characterized in that, In step S5, based on the vehicle position and speed information obtained in step S4, the vehicle trajectory is predicted and analyzed. By calculating the probability of the predicted vehicle trajectory intersecting with the traffic cone protection area, a collision risk index is calculated. The method for implementing multi-level early warning based on the collision risk index is as follows: S51, Define the vehicle trajectory and trajectory prediction problem: A vehicle trajectory is defined as a sequence of position coordinates of a vehicle over a continuous period of time; for a trajectory containing n points, its formal representation is: T = {Xi, Yi: i ∈ [1...n]} (5.1) The mapping relationship between trajectory point index i and timestamp satisfies: time(i)-time(j)=K×(ij) (5.2) In the formula, K is a fixed time step; The trajectory prediction problem is defined as follows: given a sequence of trajectories of m observed points (Xi, Yi, i∈[1...m]), predict the value of the trajectory at a future time step (Xi, Yi, i∈[m+1...n]) such that the error between the predicted trajectory and the actual trajectory t∈T is minimized. S52, a multimodal trajectory generation model based on generative adversarial networks: To address the unique challenges of vehicle trajectory prediction in road construction areas, where multiple possible travel paths exist, an improved generative adversarial network (GAN) structure is employed to achieve multimodal trajectory prediction. This structure comprises a generator network and a discriminator network, as detailed below: ① The generator network consists of two sub-networks, and its architecture is as follows: First sub-network: used to process the observed trajectory, containing an LSTM layer with 32 neurons, a second LSTM layer with 16 neurons, and a dense layer with 16 neurons; The second subnetwork receives the output of the first subnetwork and a 2-dimensional latent vector, and contains a dense layer of 16 neurons and an output layer. The mathematical expression for a generator is: G(x,z)=f2(concat[f1(x),z])(5.3) where x is the observed trajectory, z is the potential vector, and f1 and f2 represent the mapping functions of the two sub-networks respectively; ② The discriminator network consists of three sub-networks, and its architecture is as follows: First subnetwork: Used to evaluate the rationality of the generated trajectory, containing an LSTM layer with 64 neurons and a dense layer with 32 neurons; The second sub-network is used to evaluate the matching degree between the generated trajectory and the observed trajectory, and contains an LSTM layer with 64 neurons and a dense layer with 32 neurons. The third subnetwork is used to integrate the outputs of the first two subnetworks and consists of three dense layers with 32, 16, and 1 neurons respectively. The mathematical expression of the discriminator is: D(x,y)=f3(concat[f1(y),f2(concat[x,y])])(5.4) Where x is the observed trajectory, y is the generated trajectory or the true trajectory, and f1, f2 and f3 represent the mapping functions of the three sub-networks, respectively; S53, Multimodal Behavior Modeling and Differentiation Guarantee: To ensure the diversity and effectiveness of predicted trajectories, a multimodal behavior differentiation constraint mechanism is introduced: ① Behavioral Differentiation Criteria: Ensure that the generated K trajectories represent different behavioral patterns and satisfy the following: Where Ei and Ej are the normalized average displacement errors N-ADE of trajectories i and j, and T is the threshold parameter, set to 0.5; ② Differentiated trajectory generation: Differentiated trajectories are generated using the following algorithm: K predicted trajectories are generated sequentially. For each newly generated trajectory, the N-ADE ratio with the already generated trajectories is calculated. If the ratio is higher than the threshold (1-T), the trajectory is accepted; otherwise, it is regenerated. A maximum of 100 generation attempts are made to ensure algorithm efficiency. S54, Trajectory Prediction Evaluation Metrics: The following metrics are used to evaluate the quality of trajectory predictions: Average displacement error (ADE): The root mean square error of all corresponding points between the predicted trajectory and the actual trajectory; Final displacement error (FDE): The error between the final point of the predicted trajectory and the actual trajectory; Normalized average displacement error N-ADE: A standardized evaluation index that takes into account the effects of trajectory length and variance; Normalized final displacement error N-FDE: Where l(t) is the trajectory length, v(t) is the trajectory variance, Kl and Kv are constants, Kl is set to 0.04 and Kv is set to 0.003; S55, Early Warning Decision-Making: ① Input the vehicle detection results obtained in step S4 into the improved generative adversarial network model to generate three possible future trajectories. Analyze the probability of each predicted trajectory intersecting with the traffic cone area and calculate the collision risk index: CRI=max(P i ×R i ×V i ) (5.10) in: P i Let be the probability of trajectory i, output by the generative adversarial network multimodal trajectory prediction model, representing the probability of this behavior pattern, with a normalized range of [0,1]. R i As a risk factor, considering both spatial overlap and time urgency: R i =IoU×exp(-t / τ); where IoU is the intersection-exchange ratio of the vehicle trajectory and the protected area, t is the predicted collision time, and τ is the time constant; Vi is the vehicle speed factor, which considers the impact of vehicle speed on collision risk: V i =min(v / v) ref ,1), where v is the current speed of the vehicle, v ref For reference speed; The cone protection zone adopts a dynamic design, and its radius is calculated using the following formula: R p =R b +K×V (5.11) Among them, R p R is the radius of the protected area. b Where K is the basic protection radius, V is the speed coefficient, and V is the vehicle speed. ② Determine the warning level based on the calculated collision risk index value: A collision risk index value of <0.3 indicates safety and no warning is required. If the calculated collision risk index value is 0.3 or less and < 0.6, it indicates a potential risk, and a Level 1 warning is issued. If the calculated collision risk index value is 0.6 or less and < 0.8, it indicates a moderate risk, and a level-two warning is issued. A collision risk index value of ≥0.8 indicates high risk, triggering an emergency warning. The system recalculates the collision risk index every 100ms to ensure the timeliness and accuracy of the warning.
Citation Information
Cited By
Intelligent safety protection method and system based on trajectory prediction and collision situation identification
CN121598212A
Traffic cone real-time target detection method, system and equipment based on YOLOv8 identification model
CN121861595A