Intersection non-motor vehicle conflict early warning system and method based on deep learning
By adopting deep learning technology and multi-dimensional analysis methods in the intersection conflict warning system, the problem of low detection accuracy of existing systems in complex traffic environments is solved, and high-precision non-motor vehicle detection and trajectory prediction is achieved, which significantly improves the accuracy of conflict judgment and the real-time system.
Patent Information
- Application Number
- CN202510098552.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The existing intersection conflict warning system has low detection accuracy in complex traffic environments, especially in the case of light changes, severe weather or traffic congestion, making it difficult to effectively detect and track non-motor vehicles, resulting in missed or mis-checked inspections.
The intersection non-motor vehicle conflict warning system based on deep learning is adopted, including video preprocessing module, detection module, tracking module, conflict judgment module and acousto-optical early warning module. The improved YOLO-V3 algorithm was used for object detection, combined with the GRU-FST-FCN model for trajectory prediction, and risk judgment was made through a multi-dimensional conflict assessment mechanism.
It significantly improves the accuracy and real-time nature of non-motor vehicle detection, improves the accuracy of trajectory prediction, enhances the comprehensiveness and accuracy of conflict judgments, reduces the probability of missed and false alarms, and ensures the real-time nature of the system.
Smart Images

Figure CN119992841A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent transportation technology, and in particular to a non-motor vehicle conflict warning system at an intersection based on deep learning and a method thereof. Background Art
[0002] With the acceleration of urbanization and the rapid growth of motor vehicle ownership, traffic safety issues are becoming increasingly prominent, especially in complex intersections. As an important traffic participant, the conflict between non-motor vehicles and motor vehicles has become one of the main causes of traffic accidents. Traditional traffic management methods are difficult to effectively cope with this challenge. Therefore, the development of advanced warning systems has become an important research direction in the current field of traffic safety.
[0003] Existing intersection conflict warning systems mainly rely on fixed cameras and simple image processing algorithms. These systems usually use traditional computer vision techniques such as background subtraction or optical flow to detect and track vehicles. However, these methods often perform poorly in complex traffic environments, especially in conditions of changing lighting, bad weather, or heavy traffic, where the detection accuracy drops significantly. In addition, existing systems have a particularly difficult time dealing with small targets such as non-motor vehicles, and often miss or misdetect them.
[0004] On the other hand, some studies have tried to introduce machine learning methods to improve the performance of warning systems. However, most of these methods remain at the level of shallow feature extraction and cannot fully capture the complex patterns of non-motor vehicle movement. At the same time, they often ignore the interactive relationship between traffic participants, resulting in inaccurate and in-time warning results.
[0005] In addition, existing warning systems usually use a single judgment standard, such as a simple distance threshold or speed threshold, which is difficult to adapt to complex and changing traffic scenarios. In terms of warning triggering, most systems use a single alarm method, which cannot effectively distinguish different degrees of dangerous situations and easily cause driver alarm fatigue or misjudgment. Summary of the invention
[0006] In view of the above problems, there is an urgent need for a system that can accurately detect and track non-motor vehicles, accurately predict their movement trajectories, comprehensively evaluate conflict risks, and provide effective warnings. The present invention is aimed at this demand and proposes a non-motor vehicle conflict warning system and method at intersections based on deep learning.
[0007] The present invention proposes a non-motor vehicle conflict warning system at an intersection based on deep learning, comprising:
[0008] Video preprocessing module for:
[0009] Crop the surveillance video frames of the intersection;
[0010] Perform frame rate conversion and image quality processing on video data;
[0011] The detection module is connected to the video preprocessing module for:
[0012] Receiving processed video data sent by the video preprocessing module;
[0013] Based on the processed video data, using the YOLO-V3 algorithm to detect the position of the non-motor vehicle target object in the video;
[0014] Through deep learning and feature extraction, data including the position and size of the non-motor vehicle outline is obtained;
[0015] A tracking module, in communication with the detection module, is used to:
[0016] Receiving the non-motor vehicle contour position and size data sent by the detection module;
[0017] Based on the non-motor vehicle profile position and size data, tracking and obtaining speed data of the non-motor vehicle, the speed data including acceleration and speed;
[0018] The conflict determination module is in communication with the tracking module and is used to:
[0019] receiving the non-motor vehicle speed data sent by the tracking module;
[0020] Calculating potential conflict points based on the real-time position and speed data of the non-motor vehicle in combination with road boundary data;
[0021] Rank potential conflict points according to their conflict level;
[0022] The sound and light warning module is connected to the conflict judgment module for:
[0023] Receiving the potential conflict point analysis result sent by the conflict judgment module;
[0024] Based on the analysis results of the potential conflict points, an alarm reminder is issued through an audible and visual early warning device.
[0025] Preferably, the detection module comprises:
[0026] Object detection unit, used to:
[0027] The target in each frame of video is divided into three parts: detection layer, classification layer, and regression layer;
[0028] Use 3*3 convolution for downsampling to extract scale-invariant features;
[0029] Use 1*1 convolution to extract color information;
[0030] A classification unit, connected to the target detection unit, is used to:
[0031] Use softmax as a classifier;
[0032] Predict the category to which each object belongs in the classification layer;
[0033] A regression unit, connected to the target detection unit and the classification unit, is used to:
[0034] Use IOU loss as detection constraint;
[0035] Regress the center point coordinates of multiple targets on the entire image;
[0036] Predict the width and height of the target box.
[0037] Preferably, the tracking module comprises:
[0038] Speed calculation unit for:
[0039] Calculate the velocity in the x direction:
[0040] Calculate the velocity in the y direction:
[0041] Among them, v i .X,v i+1 .X is the angle between the non-motor vehicle and the x-axis in the i-th and i+1-th frames, respectively, and v i .Y,v i+1 .Y is the y-axis coordinate of the non-motor vehicle in the i-th and i+1-th frames;
[0042] A motion direction calculation unit is connected to the speed calculation unit and is used for:
[0043] Calculate the motion direction vector Q of the non-motor vehicle,
[0044] A trajectory prediction unit, connected to the motion direction calculation unit, is used to:
[0045] The moving direction vector Q of the non-motor vehicle and the road boundary are calculated to obtain the possible driving path of the non-motor vehicle;
[0046] The possible driving path of non-motor vehicles is the extension line of the road boundary segment.
[0047] Preferably, the conflict judgment module includes:
[0048] Angle calculation unit for:
[0049] Calculate the angle between the possible driving path and the potential conflict point;
[0050] The angle value is calculated as follows:
[0051] A distance calculation unit, connected to the angle calculation unit, is used to:
[0052] Calculate the distance between two objects: d = Hypotenuse (cosθ, sinθ) × v × frame rate;
[0053] Among them, θ is the direction angle between the two objects, Hypotenuse is the inverse tangent function, and v is the average speed of the two objects;
[0054] A sorting unit, connected to the angle calculation unit and the distance calculation unit, is used to:
[0055] Sort non-motor vehicles according to their travel direction;
[0056] Combine the vehicle's driving direction, acceleration and distance information for comprehensive ranking;
[0057] The conflict level is calculated by calculating the angle between the directions.
[0058] Preferably, the sound and light warning module comprises:
[0059] Sound and light equipment for:
[0060] Send out sound and light signals for early warning;
[0061] A controller, connected to the sound and light device, is used to:
[0062] Control the start and stop of sound and light equipment;
[0063] A communicator, connected to the controller, is used to:
[0064] The sound and light warning information is transmitted back to the central platform via Ethernet or WIFI.
[0065] Preferably, the target detection unit is further used for:
[0066] Divide the entire image into S*S grids;
[0067] Each grid predicts B anchor boxes;
[0068] Reduce the width and height of the entire image to 1 / 4 of the original to get more focus.
[0069] Preferably, the regression unit is further used for:
[0070] Generate 4 new anchor boxes with the predicted anchor box position as the center point;
[0071] The final detection frame is obtained by scaling as the contour position of the non-motor vehicle.
[0072] Preferably, the trajectory prediction unit adopts a GRU-FST-FCN model, including:
[0073] GRU network, used for:
[0074] Receive input frames and output corresponding distances;
[0075] Calculate the difference between the current frame and the previous frame based on the previous frame data;
[0076] The FST (Fast Spatial Transform) module, connected to the GRU network, is used to:
[0077] Perform spatial transformation on the features output by the GRU network;
[0078] FCN (Fully Convolutional Network), connected to the FST module, is used to:
[0079] Calculate the output displacement of non-motor vehicles;
[0080] The position coordinates of the non-motor vehicle are obtained and the motion trajectory of the non-motor vehicle is predicted.
[0081] Preferably, the sorting unit is further used for:
[0082] First, sort by distance, and when the distances are equal, sort by angle;
[0083] Set the minimum distance alarm threshold and maximum speed alarm threshold;
[0084] When the alarm conditions are met, two sound alarms are issued to the driver, one for a red light signal and one for a yellow light signal.
[0085] The deep learning-based intersection non-motor vehicle conflict warning method based on the system comprises the following steps:
[0086] S1: pre-processing the surveillance video frames of the intersection;
[0087] S2: Use the YOLO-V3 algorithm to detect the position of the non-motor vehicle target object in the video, and obtain data including the position and size of the non-motor vehicle outline through deep learning and feature extraction;
[0088] S3: Using the data obtained in step S2, tracking and obtaining speed data of the non-motor vehicle, the speed data including acceleration and speed;
[0089] S4: Calculate potential conflict points based on the real-time position and speed data of the non-motor vehicle and the road boundary data, and sort the potential conflict points by conflict level;
[0090] S5: Analyze potential conflict points, and use sound and light warning devices to warn you of the analysis results;
[0091] Wherein, step S2 further includes:
[0092] The target in each frame of video is divided into three parts: detection layer, classification layer, and regression layer;
[0093] Use 3*3 convolution for downsampling to extract scale-invariant features;
[0094] Use 1*1 convolution to extract color information;
[0095] Use softmax as classifier and IOU loss as detection constraint;
[0096] Step S3 also includes:
[0097] The GRU-FST-FCN model is used to learn and predict the motion trajectory of non-motor vehicles;
[0098] Step S4 also includes:
[0099] Calculate the angle between the possible driving path and the potential conflict point;
[0100] Calculate the distance between two objects;
[0101] Combine the vehicle's driving direction, acceleration and distance information for comprehensive ranking;
[0102] Step S5 also includes:
[0103] Set the minimum distance alarm threshold and maximum speed alarm threshold;
[0104] When the alarm conditions are met, an audible alarm with both red and yellow light signals is issued to the driver.
[0105] The beneficial effects of the present invention are mainly reflected in the following aspects:
[0106] The system of the present invention effectively solves the problems existing in the prior art by innovatively combining deep learning technology with multi-dimensional analysis methods. First, the present invention adopts an improved YOLO-V3 algorithm for target detection, which greatly improves the accuracy and real-time performance of non-motor vehicle detection. Secondly, the innovative GRU-FST-FCN model can accurately capture the motion characteristics of non-motor vehicles and achieve high-precision trajectory prediction. In addition, the present invention introduces a multi-dimensional conflict assessment mechanism, which comprehensively considers factors such as distance, angle, and speed, making conflict judgment more comprehensive and accurate.
[0107] In terms of system architecture, the present invention realizes close cooperation of various functional units through modular design. The video preprocessing module provides high-quality input data for subsequent analysis, and the detection module and tracking module cooperate with each other to ensure the continuity and accuracy of target recognition and trajectory prediction. The conflict judgment module integrates the output of the previous module to conduct a comprehensive risk assessment. Finally, the sound and light warning module adopts a hierarchical warning mechanism to effectively balance the sensitivity and reliability of the warning.
[0108] From the effect point of view, the present invention has achieved significant improvements in many aspects. In terms of detection accuracy, it has been improved by about 30% compared with traditional methods, especially under complex lighting and weather conditions. In terms of trajectory prediction accuracy, the average error is reduced to within 0.5 meters, laying the foundation for accurate conflict prediction. The accuracy of conflict judgment reaches more than 95%, greatly reducing the probability of missed reports and false alarms. At the same time, the real-time performance of this system is also guaranteed, and the average processing delay is controlled within 50 milliseconds to meet the needs of actual traffic scenarios.
[0109] In general, this invention proposes a comprehensive, efficient and reliable solution for non-motor vehicle conflict warning at intersections through the deep integration of deep learning technology and traffic engineering knowledge. It can not only effectively improve the level of traffic safety, but also provide new ideas and methods for the development of intelligent transportation systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0110] Figure 1 It is a logic block diagram of the whole system of the present invention. DETAILED DESCRIPTION
[0111] See also Figure 1 The present invention provides a non-motor vehicle conflict warning system and method at an intersection based on deep learning. The system is mainly used to detect and warn potential conflicts between non-motor vehicles and motor vehicles at intersections to improve traffic safety. The present invention will be described in detail below in conjunction with specific implementation methods.
[0112] like Figure 1As shown, the system of the present invention includes a video preprocessing module 1, a detection module 2, a tracking module 3, a conflict judgment module 4 and an acoustic and visual warning module 5. These modules are connected via a data bus or a network and work together to complete the task of warning the conflict of non-motor vehicles at the intersection.
[0113] The video preprocessing module 1 is used to perform preliminary processing on the intersection surveillance video. Specifically, the module first crops the original video frame to focus on the key area of the intersection. Preferably, the resolution of the cropped image can be set to 1920×1080 pixels, which can ensure the image quality without causing a huge computational burden for subsequent processing. Subsequently, the video preprocessing module 1 will perform frame rate conversion on the cropped video. In one embodiment of the present invention, the original 30fps video can be converted to 15fps, which can ensure image continuity and reduce the amount of calculation. Finally, the module will also perform image quality enhancement processing on the video, such as contrast adjustment, sharpening, etc., to facilitate subsequent target detection.
[0114] The detection module 2 is connected to the video preprocessing module 1 for receiving the preprocessed video data and detecting the non-motor vehicle target based on this. The present invention adopts the improved YOLO-V3 algorithm for target detection. The core idea of the YOLO-V3 algorithm is to divide the whole picture into S×S grids, each grid predicts B bounding boxes, and each bounding box contains 5 prediction values: x, y, w, h and confidence. Among them, x and y represent the offset of the center of the bounding box relative to the upper left corner of the grid unit, w and h represent the width and height of the bounding box relative to the whole picture, and the confidence reflects the possibility that the bounding box contains the target and the accuracy of the predicted box.
[0115] Preferably, the present invention sets S to 13 and B to 3. This setting shows a good balance between detection effect and computational efficiency in multiple experiments of the present invention. Detection module 2 also uses 1×1 convolution to extract color information and 3×3 convolution to perform downsampling to extract scale-invariant features. This multi-scale feature extraction method can effectively improve the detection accuracy of non-motor vehicle targets of different sizes.
[0116] In the classification process, the detection module 2 uses softmax as a classifier. The mathematical expression of the softmax function is:
[0117]
[0118] Among them, z i represents the raw score of the ith category, and K is the total number of categories. In the present invention, we mainly focus on the non-motor vehicle category, so K is usually set to 2 (non-motor vehicle and background).
[0119] Detection module 2 also uses IOU (Intersection over Union) loss as a detection constraint. The calculation formula of IOU is:
[0120]
[0121] Among them, Area ofOverlap represents the overlapping area of the predicted box and the true box, and Area ofUnion represents the union area of the predicted box and the true box. In the present invention, we set the IOU threshold to 0.5, that is, when the IOU is greater than 0.5, the detection result is considered valid.
[0122] The tracking module 3 is connected to the detection module 2 for receiving the detected non-motor vehicle position information and performing target tracking. The module first calculates the speed of the non-motor vehicle in the x and y directions using the following calculation formula:
[0123] Calculate the velocity in the x direction:
[0124] Calculate the velocity in the y direction:
[0125] Among them, v i .X,v i+1 .X is the angle between the non-motor vehicle and the x-axis in the i-th and i+1-th frames, respectively, and v i .Y,v i+1 .Y is the y-axis coordinate of the non-motor vehicle in the i-th and i+1-th frames; in one embodiment of the present invention, the frame rate is set to 15fps.
[0126] Based on the calculated speed information, the tracking module 3 calculates the motion direction vector Q of the non-motor vehicle:
[0127]
[0128] Among them, v i Speed represents the speed of the i-th frame. This method can effectively smooth out speed fluctuations in a short period of time and provide more stable direction prediction.
[0129] Subsequently, the tracking module 3 calculates the motion direction vector Q of the non-motor vehicle and the road boundary to obtain the possible driving path of the non-motor vehicle. It is worth noting that the present invention defines the possible driving path of the non-motor vehicle as the extension line of the road boundary line segment, which can effectively handle the possible turning behavior of the non-motor vehicle at the intersection.
[0130] Through the above steps, the system of the present invention can accurately detect and track non-motor vehicles in the intersection area, providing basic data support for subsequent conflict warnings. This deep learning-based detection and tracking method has higher accuracy and robustness than traditional methods, and can better cope with complex traffic scenarios. The conflict judgment module 4 is communicatively connected to the tracking module 3, and is used to receive speed data of non-motor vehicles and make judgments on potential conflicts based on these data. The conflict judgment module 4 of the present invention adopts an innovative multi-dimensional analysis method, which comprehensively considers factors such as angle, distance and speed to achieve more accurate conflict prediction.
[0131] First, the angle calculation unit 41 in the conflict judgment module 4 calculates the angle value between the possible driving path of the non-motor vehicle and the potential conflict point. The calculation formula of the angle value is as follows:
[0132]
[0133] Among them, (v i .X,v i .Y) represents the position coordinates of the non-motor vehicle in the i-th frame, (Q X ,Q Y ) represents the motion direction vector of the non-motor vehicle, and H is a preset constant. In a preferred embodiment of the present invention, H is set to 1000, which is verified by a large number of experiments to ensure the calculation accuracy while avoiding the problem of numerical overflow.
[0134] Next, the distance calculation unit 42 calculates the distance between the non-motor vehicle and the potential conflict point. The calculation formula is as follows:
[0135] d = Hypotenuse (cosΦ, sinΦ) · v · frame rate,
[0136] Among them, Φ is the direction angle between the two objects, and Hypotenuse is the inverse tangent function (Note: Hypotenuse here should be understood as the length of the hypotenuse of a right triangle, that is, v is the average speed of the two objects. Preferably, the present invention adopts a frame rate of 15fps, which can reduce the computational burden while ensuring the tracking accuracy.
[0137] Based on the above calculation results, the sorting unit 43 will sort the potential conflict points by conflict level. The sorting process first considers the driving direction of the non-motor vehicle, and then combines the acceleration and distance information of the vehicle for comprehensive ranking. In one embodiment of the present invention, a weighted scoring method is used for sorting, and the specific formula is as follows:
[0138]
[0139] Among them, w 1、w 2 、w 3 are weight coefficients for distance, acceleration and angle, respectively. Preferably, w 1 、w 2 、w 3 They can be set to 0.5, 0.3 and 0.2 respectively. This weight distribution fully considers the impact of various factors on conflict risk and can effectively identify high-risk potential conflict points.
[0140] The sound and light warning module 5 is connected to the conflict judgment module 4 for issuing a warning signal according to the judgment result. The module includes a sound and light device 51, a controller 52 and a communicator 53. The sound and light device 51 is used to issue sound and light signals for warning, the controller 52 is responsible for controlling the start and stop of the sound and light device, and the communicator 53 transmits the warning information back to the central platform via Ethernet or WIFI.
[0141] In a preferred embodiment of the present invention, the sound and light warning adopts a dual-level warning mechanism. When the risk of conflict is low, the system will emit a yellow flashing signal and a low-frequency alarm sound; when the risk of conflict is high, the system will emit a red fast flashing signal and a high-frequency alarm sound. This hierarchical warning mechanism can better attract the attention of drivers and pedestrians and effectively reduce the probability of traffic accidents.
[0142] In order to further improve the accuracy of target detection, the target detection unit 21 in the detection module 2 of the present invention also adopts a grid division strategy. Specifically, the unit divides the entire image into S×S grids, and each grid predicts B anchor boxes. In one embodiment of the present invention, S is set to 13 and B is set to 3. This setting shows a good balance between detection effect and computational efficiency in multiple experiments.
[0143] In addition, in order to obtain more focus points, the target detection unit 21 will reduce the width and height of the entire image to 1 / 4 of the original. This downsampling operation can effectively reduce the amount of calculation while retaining sufficient image detail information, which is suitable for real-time traffic scene analysis.
[0144] In terms of target positioning, the regression unit 22 of the present invention adopts an innovative multi-anchor frame strategy. The unit generates four new anchor frames with the predicted anchor frame position as the center point, and then obtains the final detection frame as the non-motor vehicle lane contour position through scaling. This method can more accurately locate non-motor vehicle targets of different sizes and shapes, and improves the adaptability of the system.
[0145] Through the above design, the system of the present invention can comprehensively and accurately analyze the traffic conditions at the intersection, identify potential conflict risks in a timely manner, and reduce the probability of traffic accidents through an effective early warning mechanism. The system is not only suitable for current traffic management needs, but also has good scalability and adaptability, and can cope with more complex traffic scenarios in the future. The tracking module 3 of the present invention adopts an innovative GRU-FST-FCN model for non-motor vehicle trajectory prediction. The model consists of three parts: a GRU network 31, an FST (fast spatial transform) module 32, and an FCN (fully convolutional network) 33, which can effectively capture the temporal characteristics and spatial characteristics of non-motor vehicle movement.
[0146] The GRU network 31 is the core component of the model, which is used to process time series data. It receives input frames and outputs corresponding distance information. The update rule of the GRU network is as follows:
[0147] r t =σ(W r ·[h t-1 ,x t ]),
[0148] z t =σ(W z ·[h t-1 ,x t ]),
[0149]
[0150] Among them, r t is the reset gate, z t is the update gate, h t is the hidden state, x t is the input, W r , W z and W are weight matrices, and σ is the sigmoid function. In a preferred embodiment of the present invention, the hidden layer size is set to 128, which has shown good performance in multiple experiments.
[0151] The FST module 32 is connected to the GRU network 31 and is used to perform spatial transformation on the features output by the GRU network. The core idea of FST is to learn a transformation matrix θ and then perform an affine transformation on the input feature map. The learning process of the transformation matrix θ is as follows:
[0152] θ=f loc (x),
[0153] Among them, f loc is a small convolutional neural network, and x is the input feature map. In the present invention, f locIt adopts a structure of two layers of convolution plus one layer of full connection. This lightweight design can ensure the flexibility of transformation without significantly increasing the computational burden.
[0154] FCN (Fully Convolutional Network) 33 is the last component of the model and is used to accurately calculate the output displacement of non-motor vehicles. The main feature of FCN is that the fully connected layer in the traditional CNN is replaced by a convolutional layer, which allows the network to accept inputs of any size. The forward propagation process of FCN can be expressed as
[0155] y=f(W*x+b),
[0156] Wherein, * represents the convolution operation, W is the convolution kernel, b is the bias term, and f is the activation function. The present invention uses ReLU as the activation function, and its mathematical expression is:
[0157] f(x)=max(0,x)
[0158] This choice can effectively alleviate the gradient vanishing problem and accelerate the convergence process of the network.
[0159] Through the GRU-FST-FCN model, the present invention can accurately predict the movement trajectory of non-motor vehicles and provide reliable data support for subsequent conflict warnings.
[0160] In terms of conflict judgment, the sorting unit 43 of the present invention adopts a multi-level sorting strategy. First, sorting is performed by distance, and when the distances are equal, sorting is performed by angle. This method can more comprehensively evaluate the danger level of potential conflicts.
[0161] In addition, the present invention also sets a minimum distance alarm threshold and a maximum speed alarm threshold. In a preferred embodiment, the minimum distance alarm threshold is set to 5 meters, and the maximum speed alarm threshold is set to 40km / h. When the distance between the non-motor vehicle and other vehicles is less than 5 meters, or the relative speed is greater than 40km / h, the system will trigger an alarm. These thresholds are derived based on a large amount of actual traffic data analysis, and can effectively balance the sensitivity and accuracy of early warning.
[0162] When the alarm conditions are met, the system of the present invention will issue two different sound alarms to the driver. One is a red light signal, indicating a high risk; the other is a yellow light signal, indicating a moderate risk. This graded warning mechanism can help the driver better judge the degree of danger and make appropriate responses.
[0163] Finally, the present invention also proposes a non-motor vehicle conflict warning method at an intersection based on deep learning. The method comprises the following steps:
[0164] S1: pre-processing the surveillance video frames of the intersection;
[0165] S2: Use the YOLO-V3 algorithm to detect the position of the non-motor vehicle target object in the video, and obtain data including the position and size of the non-motor vehicle outline through deep learning and feature extraction;
[0166] S3: Using the data obtained in step S2, tracking and obtaining speed data of the non-motor vehicle, the speed data including acceleration and speed;
[0167] S4: Calculate potential conflict points based on the real-time position and speed data of the non-motor vehicle and the road boundary data, and sort the potential conflict points by conflict level;
[0168] S5: Analyze potential conflict points, and use sound and light warning devices to warn you of the analysis results.
[0169] In step S2, this method divides the target in each frame of video into three parts: detection layer, classification layer, and regression layer. 3*3 convolution is used for downsampling to extract scale-invariant features, and 1*1 convolution is used to extract color information. At the same time, softmax is used as a classifier and IOU loss is used as a detection constraint.
[0170] In step S3, the method uses the GRU-FST-FCN model to learn and predict the motion trajectory of non-motor vehicles. This deep learning model can effectively capture the complex patterns of non-motor vehicle motion and improve the accuracy of prediction.
[0171] In step S4, the method calculates the angle between the possible driving path and the potential conflict point, calculates the distance between the two objects, and performs a comprehensive ranking based on the vehicle's driving direction, acceleration, and distance information.
[0172] In step S5, the method sets a minimum distance warning threshold and a maximum speed warning threshold. When the warning conditions are met, an audible warning of both a red light and a yellow light is issued to the driver.
[0173] Through this method, the present invention can comprehensively and accurately analyze the traffic conditions at intersections, timely identify potential conflict risks, and reduce the probability of traffic accidents through an effective early warning mechanism. This method is not only suitable for current traffic management needs, but also has good scalability and adaptability, and can cope with more complex traffic scenarios in the future.
[0174] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. The non-motor vehicle conflict warning system at intersection based on deep learning is characterized by: include: Video preprocessing module for: Crop the surveillance video frames of the intersection; Perform frame rate conversion and image quality processing on video data; The detection module is connected to the video preprocessing module for: Receiving processed video data sent by the video preprocessing module; Based on the processed video data, using the YOLO-V3 algorithm to detect the position of the non-motor vehicle target object in the video; Through deep learning and feature extraction, data including the position and size of the non-motor vehicle outline is obtained; A tracking module, in communication with the detection module, is used to: Receiving the non-motor vehicle contour position and size data sent by the detection module; Based on the non-motor vehicle profile position and size data, tracking and obtaining speed data of the non-motor vehicle, the speed data including acceleration and speed; The conflict determination module is in communication with the tracking module and is used to: receiving the non-motor vehicle speed data sent by the tracking module; Calculating potential conflict points based on the real-time position and speed data of the non-motor vehicle in combination with road boundary data; Rank potential conflict points according to their conflict level; The sound and light warning module is connected to the conflict judgment module for: Receiving the potential conflict point analysis result sent by the conflict judgment module; Based on the analysis results of the potential conflict points, an alarm reminder is issued through an audible and visual early warning device.
2. The system according to claim 1, characterized in that The detection module comprises: Object detection unit, used to: The target in each frame of video is divided into three parts: detection layer, classification layer, and regression layer; Use 3*3 convolution for downsampling to extract scale-invariant features; Use 1*1 convolution to extract color information; A classification unit, connected to the target detection unit, is used to: Use softmax as a classifier; Predict the category to which each object belongs in the classification layer; A regression unit, connected to the target detection unit and the classification unit, is used to: Use IOU loss as detection constraint; Regress the center point coordinates of multiple targets on the entire image; Predict the width and height of the target box.
3. The system according to claim 1, characterized in that The tracking module includes: Speed calculation unit for: Calculate the velocity in the x direction: Calculate the velocity in the y direction: Among them, v i .X,v i+1 .X is the angle between the non-motor vehicle and the x-axis in the i-th and i+1-th frames, respectively, and v i .Y,v i+1 .Y is the y-axis coordinate of the non-motor vehicle in the i-th and i+1-th frames; A motion direction calculation unit is connected to the speed calculation unit and is used for: Calculate the motion direction vector Q of the non-motor vehicle, A trajectory prediction unit, connected to the motion direction calculation unit, is used to: The moving direction vector Q of the non-motor vehicle and the road boundary are calculated to obtain the possible driving path of the non-motor vehicle; The possible driving path of non-motor vehicles is the extension line of the road boundary segment.
4. The system according to claim 1, characterized in that The conflict judgment module includes: Angle calculation unit for: Calculate the angle between the possible driving path and the potential conflict point; The angle value is calculated as follows: A distance calculation unit, connected to the angle calculation unit, is used to: Calculate the distance between two objects: d = Hypotenuse (cosθ, sinθ) × v × frame rate; Among them, θ is the direction angle between the two objects, Hypotenuse is the inverse tangent function, and v is the average speed of the two objects; A sorting unit, connected to the angle calculation unit and the distance calculation unit, is used to: Sort non-motor vehicles according to their travel direction; Combine the vehicle's driving direction, acceleration and distance information for comprehensive ranking; The conflict level is calculated by calculating the angle between the directions.
5. The system according to claim 1, characterized in that The sound and light warning module comprises: Sound and light equipment for: Send out sound and light signals for early warning; A controller is connected to the sound and light device and is used to: Control the start and stop of sound and light equipment; A communicator, connected to the controller, is used to: The sound and light warning information is transmitted back to the central platform via Ethernet or WIFI.
6. The system according to claim 2, characterized in that The target detection unit is also used for: Divide the entire image into S*S grids; Each grid predicts B anchor boxes; Reduce the width and height of the entire image to 1 / 4 of the original to get more focus.
7. The system according to claim 2, characterized in that The regression unit is also used to: Generate 4 new anchor boxes with the predicted anchor box position as the center point; The final detection frame is obtained by scaling as the contour position of the non-motor vehicle.
8. The system according to claim 3, characterized in that The trajectory prediction unit adopts the GRU-FST-FCN model, including: GRU network, used for: Receive input frames and output corresponding distances; Calculate the difference between the current frame and the previous frame based on the previous frame data; The FST fast spatial transformation module is connected to the GRU network and is used to: Perform spatial transformation on the features output by the GRU network; The FCN fully convolutional network is connected to the FST module and is used to: Calculate the output displacement of non-motor vehicles; The position coordinates of the non-motor vehicle are obtained and the motion trajectory of the non-motor vehicle is predicted.
9. The system according to claim 4, characterized in that The sorting unit is also used for: First, sort by distance, and when the distances are equal, sort by angle; Set the minimum distance alarm threshold and maximum speed alarm threshold; When the alarm conditions are met, two sound alarms are issued to the driver, one for a red light signal and one for a yellow light signal.
10. A non-motor vehicle conflict warning method at an intersection based on deep learning based on the system according to any one of claims 1 to 9, characterized in that: The following steps are involved: S1: Preprocessing the surveillance video frames of the intersection; S2: Use the YOLO-V3 algorithm to detect the position of the non-motor vehicle target object in the video, and obtain data including the position and size of the non-motor vehicle outline through deep learning and feature extraction; S3: Using the data obtained in step S2, tracking and obtaining speed data of the non-motor vehicle, the speed data including acceleration and speed; S4: Calculate potential conflict points based on the real-time position and speed data of the non-motor vehicle and the road boundary data, and sort the potential conflict points by conflict level; S5: Analyze potential conflict points, and use sound and light warning devices to warn you of the analysis results; Wherein, step S2 further includes: The target in each frame of video is divided into three parts: detection layer, classification layer, and regression layer; Use 3*3 convolution for downsampling to extract scale-invariant features; Use 1*1 convolution to extract color information; Use softmax as classifier and IOU loss as detection constraint; Step S3 also includes: The GRU-FST-FCN model is used to learn and predict the motion trajectory of non-motor vehicles; Step S4 also includes: Calculate the angle between the possible driving path and the potential conflict point; Calculate the distance between two objects; Combine the vehicle's driving direction, acceleration and distance information for comprehensive ranking; Step S5 also includes: Set the minimum distance alarm threshold and maximum speed alarm threshold; When the alarm conditions are met, an audible alarm with both red and yellow light signals is issued to the driver.
Citation Information
Patent Citations
Navigation method and device, electronic equipment and storage medium
CN113008261A
Intersection traffic conflict index calculation method considering vehicle contour
CN115547060A
Non-motor vehicle driving behavior modeling method under intersection conflict interaction scene
CN118445990A
Signal-free intersection traffic conflict identification and risk assessment system and method
CN118898903A
Methods and apparatuses for closed-loop evaluation for autonomous vehicles
US20240140486A1