A typical traffic event detection system based on deep learning
By utilizing a deep learning-based traffic incident detection system with an improved YOLOv8 network and a multi-dimensional visual feature model, the problem of unsatisfactory tracking performance in congested scenarios is solved, achieving more efficient and accurate traffic incident detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHONGQING UNIV OF POSTS & TELECOMM
- Filing Date
- 2023-06-12
- Publication Date
- 2026-05-01
AI Technical Summary
Existing traffic incident detection methods have poor tracking performance in congested traffic scenarios, resulting in long detection times, low accuracy, and difficulty in accurately locating the accident site.
A deep learning-based traffic incident detection system is adopted, including a target detection model module, a time-flow motion feature model module, and a traffic incident detection model module. It utilizes an improved YOLOv8 network, a multi-branch hole fusion module, an overlapping detection box retention mechanism, a time-flow model, and a multi-dimensional visual feature model, combined with Gaussian linear interpolation and motion information, to improve detection accuracy and robustness.
It enables more accurate location of traffic events in congested traffic scenarios, reduces false detection rate, improves detection speed and accuracy, and enhances the robustness of the model.
Smart Images

Figure CN116645563B_ABST
Abstract
Description
A typical traffic incident detection system based on deep learning Technical Field
[0001] This invention belongs to the field of vehicle-road cooperation and intelligent transportation, specifically relating to a typical traffic incident detection system based on deep learning. Background Technology
[0002] Currently, the automotive industry is gradually transforming towards intelligence, connectivity, and greening. Deep learning and vision technologies are increasingly being used in the transportation sector, and the development of vehicle-city collaboration poses new requirements for traffic efficiency and safety. Current research on traffic incident detection can be mainly divided into two methods: those based on macroscopic traffic flow data and those based on vehicle activity analysis. Methods based on macroscopic traffic flow data identify traffic accidents by comparing macroscopic traffic variables before and after the accident. This method is robust to the on-site environment; however, it has a long detection time, low accuracy, and cannot accurately pinpoint the accident location. Methods based on vehicle activity analysis and interactive modeling acquire vehicle positions from video captured by roadside cameras or extract vehicle speed, acceleration, angular velocity, or other variables from optical flow. They analyze the positional relationships or optical flow changes of related vehicles during the accident to perform accident detection. However, in congested traffic scenarios, their less-than-ideal tracking performance becomes a bottleneck, limiting their application. Summary of the Invention
[0003] In view of this, the purpose of this invention is to provide a typical traffic incident detection system based on deep learning to solve the technical problem of low efficiency in traffic incident detection.
[0004] To achieve the above objectives, the present invention provides the following technical solution:
[0005] A typical traffic incident detection system based on deep learning is disclosed. This system comprises a target detection model module, a temporal motion feature model module, and a traffic incident detection model module. The target detection model module is a traffic participant partial event detection model module constructed using an improved YOLOv8 network. The temporal motion feature model module is a temporal motion feature model module obtained based on a traffic participant trajectory estimation algorithm module. The traffic incident detection model module is a multi-dimensional visual feature event detection model module. The target detection model module includes a multi-branch hole fusion module with added coordinate attention and an overlapping detection box retention mechanism module. The temporal motion feature model module is a Gaussian linear interpolation temporal model module.
[0006] Optionally, the multi-branch hole fusion module with added coordinate attention gathers features along the horizontal and vertical directions by improving the YOLOv8 detection model, downsamples the feature map, and fuses traffic information from branches of different scales.
[0007] Optionally, the overlapping detection box retention mechanism module replaces the non-maximum suppression NMS calculation method in the YOLOv8 object detection model with soft non-maximum suppression, and uses repulsion loss as the detection box loss boxLoss calculation method.
[0008] Optionally, the Gaussian linear interpolation time-flow model module fills in the information lost in the tracking trajectory due to missed detection by Gaussian smoothing interpolation, and constructs a Gaussian process regressor with the center point of the target to predict the temporal position of the coordinates by simulating nonlinear motion through Gaussian process regression, thereby correcting the trajectory.
[0009] Optionally, the event detection model module of the multidimensional visual features supplements and interacts with the multidimensional visual features from the perspective of the target's dynamic and static features. By utilizing the relationship between motion information and traffic parameters and traffic events, several features describing the state of road events are obtained, feature coefficients for different events are designed, and a classification model is established.
[0010] Optionally, the traffic participant trajectory estimation algorithm module, by considering the influence of near-large and far-small on motion analysis, sets different speed coefficients for each type of vehicle and detects different types of trajectory conflicts that may lead to accidents.
[0011] The beneficial effects of this invention are as follows:
[0012] First, the multi-branch hole fusion module with increased coordinate attention proposed in this invention can maintain good performance when detecting targets that are far away from the camera, and can more accurately locate and identify targets of interest.
[0013] Secondly, the overlapping detection box retention mechanism module proposed in this invention can effectively prevent two prediction boxes from being filtered out by NMS because they are too close, thereby reducing missed detections and achieving better detection results when the detection boxes are highly overlapping.
[0014] Third, the time-flow algorithm module proposed in this invention can effectively distinguish between occlusion and collision of traffic objects by judging collision events through multiple conditions, thereby reducing the bit error rate of the traffic event detection model module.
[0015] Fourth, the Gaussian linear interpolation time flow model module proposed in this invention more realistically fills in the disconnected trajectory, obtains more accurate target motion parameters, and reduces the impact on the traffic incident detection model module.
[0016] Fifth, the multi-dimensional visual feature event detection model module proposed in this invention improves the robustness of the detection model, balances detection speed and accuracy, and detects multiple traffic event categories.
[0017] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0018] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:
[0019] Figure 1 is a system framework diagram of the present invention;
[0020] Figure 2 is a structural diagram of the improved YOLOv8 of the present invention;
[0021] Figure 3 is a structural diagram of the coordinate attention mechanism of the present invention;
[0022] Figure 4 is a structural diagram of the StrongSORT model of the present invention. Detailed Implementation
[0023] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0024] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0025] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0026] Please refer to Figures 1 to 4, which show a typical traffic incident detection system based on deep learning.
[0027] Figure 1 shows a typical traffic incident detection system based on deep learning, comprising a target detection model module, a temporal motion feature model module, and a traffic incident detection model module. The target detection model module utilizes an improved YOLOv8 deep learning algorithm to detect traffic participants and directly outputs rollover, damage, and fire incidents as target categories. The temporal motion feature model module uses an improved temporal algorithm to track detected traffic participants, generating driving trajectories and traffic feature parameters, and filling in missing trajectories. The traffic incident detection model module proposes a multi-dimensional feature classification model to detect typical traffic incidents. The multi-dimensional visual feature incident detection model module supplements and interacts with multi-dimensional visual features from the perspectives of target dynamic and static features. Utilizing the relationship between motion information, traffic parameters, and traffic incidents, it obtains several features describing the state of road incidents, designs feature coefficients for different incidents, and establishes a classification model. Specifically, it first detects the pixel positions where a collision is likely to occur based on the distance between vehicles and the overlap of vehicle detection boxes. The distance between vehicles is transformed into a problem of determining the distance between vehicle bounding boxes, i.e., the minimum distance between two quadrilaterals. Assuming the minimum bounding rectangles of two vehicle detection boxes are A1B1C1D1 and A2B2C2D2, the minimum distance between them is:
[0028]
[0029] Where, d pq d represents the Euclidean distance between the rectangle and the point; qq P1 represents the Euclidean distance between rectangles; P2 represents a point in A1B1C1D1; P2 represents a point in A2B2C2D2.
[0030] X represents the coordinates (x, y) of a point on the rectangle, d peThis represents the Euclidean distance between a point and an edge; Formula (1) simply means that the minimum distance between two rectangles is equal to the minimum of the minimum distance between a point on the first rectangle and the second rectangle and the minimum distance between a point on the second rectangle and the first rectangle. The minimum distance between a point on a rectangle and the rectangle itself is equal to the minimum distance from a point on the rectangle to the four sides of the rectangle.
[0031] When the minimum distance is less than a threshold, the target is judged to be either occluded or in a collision. Based on kinematic characteristics, the current dynamic features of the object are calculated using the target's historical trajectory information. Assuming the target's driving direction and speed remain unchanged, the vehicle's position in the next few frames is predicted, and it is determined whether the vehicle's future trajectory will intersect. The target is continued to be tracked, and it is determined whether the target's acceleration and direction change abruptly at the predicted intersection point. If pedestrians suddenly gather, feature weight coefficients are set for the above conditions, and a multi-dimensional feature classification model is constructed to determine whether a vehicle collision has occurred, reducing the false detection rate of the event detection system and making the detection model more accurate and faster. When the vehicle's position does not change significantly over a period of time, the average speed of most vehicles is below a threshold, and static parameters such as road traffic density and traffic flow are compared with normal levels, a congestion event is judged. The vehicle's driving direction vector is obtained by connecting the vehicle's 5-frame trajectory. When the angle between the driving direction vector and the lane direction vector is greater than a threshold, a reverse driving event is judged. When the vehicle speed is 0 and the vehicle is located in a designated no-stopping zone, a parking violation event is judged.
[0032] Roadside cameras are suspended at different heights, and different vehicle models have different dimensions. If the vehicle is far from the camera and installed at a high distance, the video will show a very small area around the vehicle; conversely, it will show a larger area. Because the vehicle's area varies in the video, when estimating speed in pixels, a smaller area appears to move slower than a larger area. To estimate the trajectory of traffic objects, the impact of near-to-far dimensions on motion analysis is considered. A different speed coefficient is set for each type of vehicle, and different types of trajectory conflicts that could lead to accidents are detected. Specifically, two speed estimation parameters are set: an area parameter and a distance parameter. The area parameter is inversely proportional to the area, and the distance parameter is directly proportional to the distance to the camera. When a vehicle has a larger area, the smaller the area parameter, and the lower its speed. Therefore, a different speed coefficient parameter needs to be set for each type of vehicle; for example, a smaller area parameter for trains and a larger area parameter for cars. The farther the vehicle is from the camera, the larger the distance parameter. The video area needs to be divided from top to bottom, and the parameter decreases linearly. Multiplying the vehicle's average pixel speed by the speed estimation parameter reduces the false detection rate of vehicle occlusion.
[0033] The specific process of the object detection model includes the following stages:
[0034] Dataset creation phase: An open-source dataset containing traffic events was collected, including rollovers, fires, collisions, congestion, and wrong-way driving. Traffic objects and events in the dataset were re-labeled manually. Traffic participants included cars, trucks, buses, motorcycles, bicycles, and pedestrians, forming a typical traffic event dataset. To address the issues of limited sample size and imbalanced distribution, a method was designed to randomly select mask regions to cover certain pixel blocks in images with fewer samples, increasing training data and improving the model's generalization ability.
[0035] Training Phase: The dataset is input into the improved event detection network. A random dropout strategy is employed, first randomly removing half of the hidden neurons while keeping the input and output neurons unchanged. Then, the input is forward-propagated through the modified network, and the resulting loss is back-propagated through the same network. After this process is completed with a small batch of training samples, the parameters of the remaining neurons are updated using stochastic gradient descent to improve the model's generalization ability and prevent overfitting. Fine-tuning is performed on some parameters, such as the learning rate and epochs, to finally obtain the converged model.
[0036] Testing Phase: Test set images and videos are input into the trained network model to evaluate its detection accuracy. Based on the test results, the model is fine-tuned to obtain the optimal model. Specifically, the network model's performance is tested using test set data. For vehicle detection, precision and recall are used to evaluate the model's accuracy. For event detection, accuracy and detection speed are used to evaluate the model's detection accuracy. If all the above metrics meet the requirements, the hyperparameter results from the network model training are used as the network model for actual video testing. If the requirements are not met, the sample set and network initialization parameters are updated, and the network model is retrained.
[0037] Usage Phase: Video data is extracted from actual road surveillance videos, preprocessed, and traffic flow thresholds are set based on historical traffic flow data. Direction vectors for no-parking lanes and driving lanes are calibrated. The preprocessed surveillance video is then input into an optimized detection model to obtain detection results for traffic participants. These results are then fed into a traffic anomaly detection algorithm to promptly identify abnormal traffic events on the road. Due to limited storage capacity in practical applications, if no target is detected within a certain period, it is determined that the target has left the monitoring field of view, and the stored target trajectory data is deleted.
[0038] Regarding the problem of small target detection and localization, a multi-branch dilated fusion module with coordinate attention is proposed. By improving the YOLOv8 detection model to aggregate features along the horizontal and vertical directions, downsample the feature map, and fuse traffic information of different scale branches. Specifically: The overall structure of the improved YOLOv8 is shown in Figure 2. The Backbone part is a feature extraction network used to extract information from images. Using the CSP idea and adopting the C2F module to further lightweight the model. The Backbone part aims at the problem that the use of hierarchical convolution and pooling layers in the CNN architecture leads to the loss of fine-grained feature information, especially the performance degradation on low-resolution images and small objects. By using the SPD-Conv module with spatial-to-depth convolution and non-hierarchical convolution, the hierarchical and pooling operations are completely eliminated, and while preserving the discriminative feature information, the feature map is downsampled. The SPD-Conv consists of an SPD layer and a non-strided convolution layer. The SPD layer uses an image transformation technique to downsample the feature maps inside and throughout the CNN. The SPD layer transforms the feature map (S, S, C1) into an intermediate feature map (S / scale, S / scale, scale2C1), and then concatenates these sub-feature maps along the channel dimension to obtain a feature map whose spatial dimension is reduced by a scale factor and the channel dimension is increased by a scale factor. After the SPD feature transformation layer, a non-strided convolution layer with C2 filters is added, where C2 < scale2C1, and it is further transformed into (S / scale, S / scale, C2) to preserve all discriminative feature information as much as possible.
[0039] The structure of the coordinate attention mechanism algorithm in this invention is shown in Figure 3. The coordinate attention mechanism is divided into two steps: coordinate information embedding and coordinate attention generation. By embedding position information into channel attention, the network can perform attention on a larger area while avoiding a large amount of computational overhead. That is, given the input X, pooling kernels with two spatial ranges (H, 1) or (1, W) are used to encode each channel along the horizontal and vertical coordinates respectively. Therefore, the output calculation formula for the c-th channel with height h is as follows:
[0040]
[0041] The output calculation formula for the c-th channel with width w is as follows:
[0042]
[0043] To mitigate the loss of positional information caused by 2D global pooling, channel attention is decomposed into two parallel 1D feature encoding processes. This effectively integrates spatial coordinate information into the generated attention map. Specifically, two 1D global pooling operations are used to aggregate the input features in the vertical and horizontal directions into two independent orientation-aware feature maps. The outputs of the width and height channels are then sequentially passed through a concat layer, a 1×1 convolutional layer, and a non-linear activation layer to obtain intermediate feature maps that encode spatial information in the horizontal and vertical directions.
[0044] The intermediate feature map is split into two independent tensors f along the spatial dimension. h and f w These two feature maps, each embedding specific directional information, are encoded into two attention maps. Each attention map captures the long-range dependency of the input feature map along a spatial direction, thus preserving the location information within the generated attention map. A 1×1 convolutional layer and a sigmoid activation layer transform the input feature map into a tensor with the same number of channels. These two tensors are then multiplied into the input feature map to enhance its representational power.
[0045] YOLOv8's neck part mimics the organization of an FPN network. It first downsamples, then upsamples, with two cross-layer fusion connections between the downsampling and upsampling branches. An SPD-Conv module is added before the C2F connection between these two cross-layer fusion connections. In the downsampling part, dilated convolutions adaptively learn different receptive fields in each feature map according to the different scales of the detected targets, thereby improving the accuracy of multi-scale object detection and recognition. This can be divided into two parts: multi-branch convolutional layers and multi-branch pooling layers. The multi-branch convolutional layers provide receptive fields of different sizes to the input feature map through dilated convolutions. The average pooling layers fuse traffic information from the receptive fields of the three branches to improve the accuracy of multi-scale prediction.
[0046] The overlapping bounding box preservation mechanism module replaces the NMS calculation method in the YOLOv8 object detection model with soft-NMS, and uses Repulsion Loss as the box loss calculation method. Specifically, the head part uses a decoupled head structure to separate the classification and detection heads, and calculates the loss separately for each. One head performs classification, using the binary cross-entropy loss function. The other head performs object recognition, using Bbox Loss, and the loss function uses Repulsion Loss instead of DFL. The Repulsion Loss expression is as follows:
[0047] L = L Attr +α*L RepGT +β*L RepBox (4)
[0048] Where L Attr To make the predicted bounding box closer to the ground truth bounding box, L Rep This is to keep the predicted bounding box away from the surrounding ground truth bounding boxes, L Rep It is L RepGT and L RepBox The collective term, with parameters α and β used to balance the weights of both. L Attr The expression is as follows:
[0049]
[0050] Where P + This represents all positive samples, which are the sets of detection boxes P divided according to a set IoU threshold. This means that for each detection box P, a ground truth bounding box with the maximum IoU value is matched, B P This represents the predicted bounding box obtained after regression offset from the detection box P.
[0051] Smooth L1 The expression is as follows:
[0052]
[0053] In other words, B P and Smooth the top left corner coordinates and width and height (x, y, w, h) respectively. L1 The calculations are then performed, and the results are summed. The goal of optimization is to shorten the distance between the two, making the predicted bounding box and its target bounding box as close as possible.
[0054] L RepGT The expression is as follows:
[0055]
[0056] Smooth ln The expression is:
[0057]
[0058] L RepGT middle This represents the bounding box with the largest IoU (Intersection over Union) excluding the bounding boxes that already match the predicted bounding box P. The expression is:
[0059]
[0060] L RepBox The expression is as follows:
[0061]
[0062] Where Ι is the identity function and ε is a constant. The larger the IoU between the predicted bounding box Pi and the surrounding predicted bounding boxes Pj, the greater the resulting loss will be. The RepBox loss can reduce the probability that the predicted bounding boxes of different regression targets will merge into one after NMS, making the detector more robust to scenes with dense targets.
[0063] The post-processing section addresses the issue of false deletions when predicted bounding boxes are too close together. It replaces the NMS calculation method in the YOLOv8 object detection model with soft-NMS, applying a Gaussian weight to the confidence reset function. The soft-NMS calculation formula is as follows:
[0064]
[0065] Where iou(M,b) i () represents the intersection-union ratio of the two boxes, M is the box with the highest score, and b i For the box to be processed, s i The score is given for the predicted bounding box. i The larger the intersection-union ratio of M, the better. i The more severe the drop, the better. By reducing the confidence of currently overlapping detection boxes, the impact of duplicate boxes is reduced, and more detection results are retained.
[0066] The StrongSORT architecture in this invention is shown in Figure 4. StrongSORT consists of an appearance branch and a motion branch. In the appearance branch, a more powerful appearance feature extractor, BOT, is used instead of the simple CNN in DeepSORT. StrongSORT uses ResNeSt50 as its backbone and is pre-trained on the DukeMTMC-Reid re-identification dataset, enabling it to extract more discriminative features. A feature update strategy replaces the feature library, updating the appearance state of the i-th trajectory at frame t using an exponential moving average (EMA).
[0067]
[0068] Where f i t It represents the appearance features of the current matching detection, where α is a momentum term. This represents the trajectory appearance state of frame t-1. The EMA update strategy utilizes information about feature changes between each frame to suppress detection noise. In the motion branch, ECC is used for camera motion compensation to estimate global rotation and translation between adjacent frames. Furthermore, ordinary Kalman filtering is easily attacked by low-quality detections and ignores information about the detection noise scale. An improved NSA Kalman algorithm is used, employing an adaptive noise covariance calculation method with pre-defined measurement noise covariance parameters. When the detection has less noise, the confidence score is higher, meaning the detection has a higher weight in state updates, which helps improve the accuracy of state updates.
[0069] During the matching process, both appearance and motion information are considered simultaneously, and the cost matrix is a weighted sum of appearance and motion costs. As trackers become more powerful, they become more robust to confusing associations, but additional prior constraints limit matching accuracy. StrongSORT uses Vanola global linear assignment instead of matching cascades. The Gaussian linear interpolation time-flow model module fills in the missing tracking trajectory information caused by missed detections through Gaussian smooth interpolation. It simulates nonlinear motion using Gaussian process regression, constructs a Gaussian process regressor based on the target's center point, predicts the temporal position of the coordinates, and corrects the trajectory. Specifically, to address the problem of trajectory disconnection due to missing target detection in tracking tasks, Gaussian linear interpolation is added to the time-flow algorithm, using Gaussian process regression to simulate nonlinear motion. The Gaussian smooth interpolation model of the i-th trajectory is expressed as:
[0070] p t =f (i) (t)+ε (13)
[0071] Where t∈F is the frame number, p t ∈P is the position coordinate (x,y,w,h) of the t-th frame, where (x,y) is the coordinate of the lower left corner of the tracked target, w and h represent the width and height of the detection box, and ε~N(0,σ) 2 () is Gaussian noise. Nonlinear motion modeling is achieved by fitting a function f. (i) The solution, assuming it follows a Gaussian process, is as follows:
[0072] f (i) ∈GP(0,k(·,·)) (14)
[0073] in It is the radial basis function and the hyperparameter λ controls the smoothness of the trajectory. It is a function that adapts to the trajectory length. K(·,·) is a suffix that can be either k(x,x') or k(y,y'). (x,x') represents the x-coordinates of two trajectory points.
[0074] Based on the properties of Gaussian processes, given a new frame set F * Its smooth position P * The prediction method is expressed as follows:
[0075] P * =K(F * ,F)(K(F,F)+σ 2 I) -1 P (15)
[0076] Where K(·,·) is the covariance function based on k(·,·), F represents the known frame set, and F* represents the frame set to be tested.
[0077] When using Gaussian process regression, the input observation coordinates include not only the position when the trajectory is not interrupted, but also the position information estimated by linear frame interpolation after the trajectory is interrupted. After integrating all this information, the position at a certain interruption time is estimated to fill the trajectory gap.
[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A typical traffic incident detection system based on deep learning, characterized in that: The system sequentially comprises a target detection model module, a time-flow motion feature model module, and a traffic event detection model module. The target detection model module performs target detection on input roadside surveillance video images and outputs the detection results of traffic participants. The time-flow motion feature model module includes a Strong SORT module and a GSI-time-flow model module. The Strong SORT module takes the detection results from the target detection model module as input and outputs a preliminary trajectory. The GSI-time-flow model module takes the Strong SORT results as input. The SORT module outputs preliminary trajectory information, which is then converted to corrected trajectory information. The traffic event detection model module is a classification model module based on multi-dimensional visual features. It receives the trajectory information output by the temporal motion feature model module, performs traffic event classification based on multi-dimensional visual features, and outputs traffic event detection results. The target detection model module is a traffic participant partial event detection model module built with an improved YOLOv8 network. This module has a parallel multi-branch hole fusion module with added coordinate attention and an overlapping detection box retention mechanism module. The overlapping detection box retention mechanism module replaces the non-maximum suppression NMS calculation method in the YOLOv8 target detection model with soft non-maximum suppression, and uses repulsion loss as the detection box loss box loss calculation method. The temporal motion feature model module is based on the traffic participant motion trajectory estimation algorithm module. It fills in the information lost in the tracking trajectory due to missed detections by Gaussian smoothing interpolation, and constructs a Gaussian process regressor based on the center point of the target to predict the temporal position of the coordinates and correct the trajectory.
2. The typical traffic incident detection system based on deep learning according to claim 1, characterized in that: The multi-branch hole fusion module with added coordinate attention gathers features along the horizontal and vertical directions by improving the YOLOv8 detection model, downsamples the feature map, and fuses traffic information from branches of different scales.
3. The typical traffic incident detection system based on deep learning according to claim 1, characterized in that: The event detection model module of the multidimensional visual features supplements and interacts with the multidimensional visual features from the perspective of target dynamic and static features. By utilizing the relationship between motion information and traffic parameters and traffic events, it obtains several features describing the state of road events, designs feature coefficients for different events, and establishes a classification model.
4. A typical traffic incident detection system based on deep learning according to claim 1, characterized in that: The traffic participant trajectory estimation algorithm module, by considering the impact of near-large and far-small on motion analysis, sets different speed coefficients for each type of vehicle and detects different types of trajectory conflicts that may lead to accidents.
Citation Information
Patent Citations
Highway traffic incident detection method based on deep learning
CN114724063A
Novel target detection system and method under roadside view angle
CN115346177A