A traffic anomaly event detection method, device and computer storage medium

By obtaining the monitoring video stream in the traffic monitoring system and estimating the keyframe reference value using spatiotemporal feature correlation, key video frames are determined to detect abnormal events, and the problem of inefficient detection in the prior art is solved, and more efficient and accurate traffic abnormal event detection is achieved.

CN113673311BActive Publication Date: 2025-05-30ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110758806.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-05
Publication Date
2025-05-30
Estimated Expiration
2041-07-05

AI Technical Summary

Technical Problem

The existing traffic monitoring system relies on human eye observation, making it difficult to monitor a large number of camera scenes at the same time and accurately identify traffic abnormal events, resulting in ineffective detection.

Method used

By acquiring the monitoring video stream, the spatial and temporal characteristics correlation between the video frame and the associated video frame is used to estimate the key frame reference value, thereby determining the key video frame, and then detecting whether there are abnormal events in the monitoring video stream.

Benefits of technology

It improves the detection efficiency of traffic anomalies, avoids the impact of non-critical video frames on detection, and enhances the accuracy and real-timeness of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113673311B_ABST
    Figure CN113673311B_ABST
Patent Text Reader

Abstract

The present application provides a traffic anomaly event detection method, device, and computer storage medium. The traffic anomaly event detection method includes: obtaining a monitoring video stream; estimating a key frame reference value for each video frame based on the spatio-temporal feature correlation between each video frame and its associated video frame in the monitoring video stream; and determining key video frames from the monitoring video stream based on the key frame reference values of each video frame; and determining whether there is an anomaly event in the monitoring video stream based on the key video frames. Through the above method, the traffic anomaly event detection method of the present application improves the detection efficiency of traffic anomaly events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video processing, and particularly to a method and device for detecting traffic abnormal events and a computer storage medium. Background Art

[0002] To ensure the stable and orderly development of the transportation economy, the current demand for the inspection and analysis of traffic abnormal events is increasing continuously, and more and more traffic monitoring cameras are installed in various traffic roads. However, since most of the current traffic monitoring systems rely on human eye observation, it is difficult to monitor a large number of camera scenes simultaneously and accurately identify abnormal traffic events, resulting in low detection efficiency. Summary of the Invention

[0003] This application provides a method and device for detecting traffic abnormal events and a computer storage medium.

[0004] This application provides a method for detecting traffic abnormal events, and the method for detecting traffic abnormal events includes:

[0005] Obtain a monitoring video stream;

[0006] Estimate the key frame reference value of each video frame based on the spatio-temporal feature correlation between each video frame and its associated video frame in the monitoring video stream; and determine key video frames from the monitoring video stream based on the key frame reference value of each video frame;

[0007] Determine whether there is an abnormal event in the monitoring video stream based on the key video frames.

[0008] This application also provides a terminal device, which includes a memory and a processor, wherein the memory is coupled to the processor;

[0009] wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned method for detecting traffic abnormal events.

[0010] This application also provides a computer storage medium, which is used to store program data, and when the program data is executed by a processor, it is used to implement the above-mentioned method for detecting traffic abnormal events.

[0011] The beneficial effect of this application is that the terminal device estimates the key frame reference value of each video frame by using the spatio-temporal feature correlation between each video frame and its associated video frame in the monitoring video stream, and then obtains key video frames that can detect abnormal events in the monitoring video stream based on the key frame reference value of each video frame, avoiding the influence of non-key video frames on the detection of traffic abnormal events and improving the detection efficiency of traffic abnormal events. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:

[0013] Figure 1 is a schematic flowchart of an embodiment of the traffic anomaly event detection method provided by this application;

[0014] Figure 2 is a schematic flowchart of another embodiment of the traffic anomaly event method provided by this application;

[0015] Figure 3 is a schematic diagram of the video segment storage unit and the key frame recommendation network in the traffic anomaly event detection method provided by this application;

[0016] Figure 4 is a schematic diagram of the video segment storage unit in the traffic anomaly event detection method provided by this application;

[0017] Figure 5 is a schematic diagram of the key frame recommendation network in the traffic anomaly event detection method provided by this application;

[0018] Figure 6 is Figure 2 a schematic flowchart of an embodiment of S205 in the traffic anomaly event detection method shown;

[0019] Figure 7 is Figure 2 a schematic flowchart of an embodiment of S206 in the traffic anomaly event detection method shown;

[0020] Figure 8 is Figure 7 a schematic flowchart of an embodiment of S404 in the traffic anomaly event detection method shown;

[0021] Figure 9 is Figure 2 a schematic diagram of the pre-detection network in the traffic anomaly event detection method shown;

[0022] Figure 10 is Figure 2 a schematic flowchart of another embodiment of S206 in the traffic anomaly event method shown;

[0023] Figure 11 is Figure 1 a schematic flowchart of the process of obtaining the key frame recommendation network in the traffic anomaly event detection method shown;

[0024] Figure 12 is a schematic structural diagram of an embodiment of a terminal device provided by the present application;

[0025] Figure 13 is a schematic structural diagram of an embodiment of a computer storage medium provided by the present application. Detailed implementation manners

[0026] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0027] Please refer to Figure 1 , Figure 1 is a schematic flowchart of an embodiment of a traffic anomaly event detection method provided by the present application.

[0028] Among them, the traffic anomaly event detection method of the present application is applied to a type of terminal device. Among them, the terminal device of the present application can be a server, or a system in which the server and the terminal device cooperate with each other. Correspondingly, each part included in the terminal device, such as each unit, sub-unit, module, and sub-module, can be all set in the server, or can be respectively set in the server and the terminal device.

[0029] Furthermore, the above-mentioned server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules used to provide a distributed server, or can be implemented as a single software or software module, which is not specifically limited herein. In some possible implementation manners, the traffic anomaly event detection method in the embodiments of the present application can be implemented by a processor calling computer-readable instructions stored in a memory. Specifically, as Figure 1 shown, the traffic anomaly event detection method in the embodiments of the present application specifically includes the following steps:

[0030] S101: Obtain a monitoring video stream.

[0031] In the embodiments of the present disclosure, the terminal device obtains a monitoring video stream. In a specific embodiment, a camera installed above the traffic road surface captures the monitoring video stream and sends the monitoring video stream to the terminal device connected to the camera. In other embodiments, the terminal device can also be a device with video shooting and processing functions, and the terminal device installed above the traffic road surface directly captures the monitoring video stream.

[0032] Furthermore, considering that the monitored areas in the surveillance video stream include road surface areas and non-road surface areas, in order to achieve targeted detection of traffic abnormal events and avoid wasting resources by detecting the surveillance video stream corresponding to non-road surface areas, the terminal device in this embodiment can also filter the information of non-road surface areas. Specifically, the terminal device identifies the road surface areas in the surveillance video stream, that is, the road surface areas to be detected, and sets the detection area according to the road surface areas to be detected. Among them, the terminal device can directly set the road surface areas to be detected as the detection area.

[0033] S102: Estimate the key frame reference values of each video frame based on the spatio-temporal feature correlation between each video frame in the surveillance video stream and the front and rear video frames, and determine the key video frames from the surveillance video stream based on the key frame reference values of each video frame.

[0034] Among them, the surveillance video stream includes key video frames and non-key video frames. In order to avoid the influence of non-key video frames on the detection of key video frames, the terminal device in this embodiment filters the non-key video frames in the surveillance video stream and determines the key video frames in the surveillance video stream. Specifically, the terminal device uses the spatio-temporal feature correlation between each video frame in the surveillance video stream and its associated video frames to estimate the key frame reference values of each video frame. Then, the key video frames are determined based on the key frame reference values of each video frame.

[0035] Among them, the associated video frame of a video frame in the surveillance video stream can be the first number of video frames before a video frame in the surveillance video stream. It can also be the second number of video frames after a video frame in the surveillance video stream. It should be noted that the first number is at least one number. The second number is also at least one number. Specifically, the associated video frame of a video frame can be one video frame or multiple video frames whose frame sequence is before the current video frame in the surveillance video stream where the current video frame is located. It can also be one video frame or multiple video frames whose frame sequence is after the current video frame in the surveillance video stream where the current video frame is located.

[0036] Furthermore, considering the variability of traffic abnormal events, the terminal device in this embodiment can use a trainable network model to extract the key video frames in the surveillance video stream. Specifically, the terminal device inputs the surveillance video stream into the key frame recommendation network and obtains the key video frames output by the key frame recommendation network. Then, the redundancy of repeated analysis of adjacent video frames is reduced, and the accuracy and real-time performance of traffic abnormal event detection are improved.

[0037] Specifically, the key frame recommendation network includes three units: a 2D feature encoding unit, a bidirectional LSTM unit, and a feed-forward network unit. Among them, the 2D feature encoding unit is used to extract the spatial features of each video frame input by the video segment cache unit, and this unit multiplexes the network trained on the publicly available ImageNet dataset (such as Resnet-101). The bidirectional LSTM unit is used to associate the spatio-temporal feature correlations of N video frames in the surveillance video stream, increasing the utilization rate of video segment features. The feed-forward network unit is used to fuse the high-level semantic features associated by the bidirectional LSTM unit and output the importance score of each video frame, that is, the key frame reference value.

[0038] S103: Based on the key video frames, determine whether there are abnormal events in the surveillance video stream.

[0039] Among them, the abnormal events can be cargo spillage, traffic accidents, smoke and fog, and road construction, etc. Among them, the terminal device determines whether there is at least one of the abnormal events such as cargo spillage, traffic accidents, smoke and fog, and road construction in the surveillance video stream according to the key video frames.

[0040] In the above solution, the terminal device utilizes the spatio-temporal characteristic correlations between each video frame in the surveillance video stream and its associated video frames to estimate the key frame reference values of each video frame, and then obtains the key video frames that can detect abnormal events in the surveillance video stream according to the key frame reference values of each video frame, avoiding the influence of non-key video frames on the detection of traffic abnormal events and improving the detection efficiency of traffic abnormal events.

[0041] Please continue to refer to Figure 2 , Figure 2 which is a schematic flowchart of another embodiment of the traffic abnormal event method provided by this application. Specifically, the traffic abnormal event method of this embodiment further includes the following steps:

[0042] S201: Obtain the surveillance video stream.

[0043] For the detailed description of S201 in this embodiment, please refer to S101 in the above embodiment, and no repeated description will be given here.

[0044] S202: Preset a first cache unit and a second cache unit.

[0045] Please refer to Figure 3 , Figure 3It is a schematic diagram of the video segment storage unit and the key frame recommendation network in the traffic anomaly event detection method provided by this application. Since the traffic anomaly events in this embodiment mainly include cargo spills, traffic accidents, fireworks and fog, and road construction, etc., they have obvious suddenness and the correlation of adjacent frames. If directly detecting traffic anomaly events based on single-frame images without considering the suddenness of anomaly events and the correlation of adjacent frames, it will lead to the problem of low detection accuracy. To avoid the occurrence of the above problems, the terminal device in this embodiment first uses the video segment cache unit to store video frames in the monitoring video stream, and then uses the key frame recommendation network to associate the correlation and suddenness of adjacent frame video frames. Among them, the video segment cache unit includes a first cache unit and a second cache unit, and a total of N frames of images are stored in the video segment cache unit.

[0046] S203: Store the video frames of the preset number of frames in the monitoring video stream into the first cache unit and the second cache unit.

[0047] Continue to refer to Figure 4 . The terminal device can determine the number of video frames of the preset number of frames according to the computing power of the key frame recommendation network, store half of the video frames of the preset number of frames in the first cache unit, and store the other half of the video frames of the preset number of frames in the second cache unit. For details, refer to Figure 4 . Figure 4 A and B in are the first cache unit and the second cache unit in the video segment cache unit respectively.

[0048] Specifically, the monitoring video stream is alternately stored in the first cache unit and the second cache unit in the form of single-frame images, so that N / 2 frames of images are stored in the first cache unit and the second cache unit respectively. In practical applications, when the video segment storage unit and the key frame recommendation network are first started for calculation, it should be ensured that the first cache unit and the second cache unit are full of image frames. Then, calculate every N / 2 frames. For example, if the start calculation time is set as t, N / 2 frames of images are stored in the first cache unit at time t - 2, N / 2 frames of images are continuously stored in the second cache unit at time t - 1, and N / 2 frames of images are stored in the first cache unit in real time at time t + 1, and the second cache unit remains unchanged at this time. Repeat the above process in turn to alternately store single-frame images in the first cache unit and the second cache unit.

[0049] S204: Use the key frame recommendation network to obtain the key frame reference values of the video frames in the first cache unit and the second cache unit respectively.

[0050] Among them, the network schematic diagram of the key frame recommendation network can be referred to Figure 5。The key frame recommendation network includes a 2D feature encoding unit, a bidirectional LSTM (Long Short-Term Memory) unit, and a feed-forward network unit. Among them, the 2D feature encoding unit is used to extract the spatial features of each video frame in the first buffer unit and the second buffer unit. In a specific embodiment, the 2D feature encoding unit can be a network trained on the ImageNet public dataset, such as the fast training residual network Resnet-101. The bidirectional LSTM unit is used to associate the correlation between the spatio-temporal features before and after N video frames in the first buffer unit and the second buffer unit, so as to increase the utilization rate of video frame features. The feed-forward network unit is used to fuse the high-level semantic features associated by the bidirectional LSTM unit to output the key frame score of each video frame in the first buffer unit and the second buffer unit, that is, the key frame reference value.

[0051] S205: Screen out the key video frames based on the key frame reference values of all video frames.

[0052] Optionally, this embodiment may adopt Figure 6 the embodiment to implement S205, which specifically includes S301 to S304:

[0053] S301: Compare the key frame reference values of each video frame with the preset key frame reference value threshold in turn according to the order of each video frame in the monitoring video stream.

[0054] To ensure the accuracy of the key frame reference values of the video frames output by the key frame recommendation network, the terminal device performs inference processing on the key frame reference values of the video frames output by the key frame recommendation network. Specifically, the key frame recommendation network further includes a post-processing unit, and the post-processing unit is used to output the index number of the corresponding video frame to implement the determination of the key frame reference value among all video frames.

[0055] Furthermore, the terminal device determines whether the key frame reference value of each video frame is greater than the preset key frame reference value threshold. If so, execute S302. If not, discard the video frames whose key frame reference values are smaller than the preset key frame reference threshold.

[0056] S302: Determine the video frames with key frame reference values greater than the key frame reference value threshold as the first key video frames.

[0057] Among them, when the terminal device determines that the key frame reference value of the video frame is greater than the preset key frame reference value threshold, the current video frame is determined as the first key video frame.

[0058] S303: Calculate the similarity between a first key video frame and the previous first key video frame in turn according to the order of each first key video frame in the monitoring video stream.

[0059] To filter duplicate key video frames, the terminal device calculates the similarity between a first key video frame and the previous first key video frame in sequence according to the order of the first key video frames in the surveillance video stream.

[0060] S304: Determine whether the similarity is greater than a preset similarity threshold.

[0061] Among them, the terminal device determines whether the similarity between the current first key video frame and the previous first key video frame is greater than the preset similarity threshold. If so, execute S305. If not, execute S306.

[0062] S305: Use the previous first key video frame as the key video frame for the final output.

[0063] Among them, when the terminal device determines that the similarity between the current first key video frame and the previous first key video frame is greater than the preset similarity threshold, it uses the current first key video frame as the key video frame for the final output.

[0064] S306: Discard the previous first key video frame and use a first key video frame as the second key video frame for the final output.

[0065] Among them, when the terminal device determines that the similarity between a first key video frame and the previous first key video frame is less than or equal to the preset similarity threshold, it discards a first key video frame and uses the previous first key video frame as the key video frame for the final output.

[0066] S206: Detect abnormal events in the key video frame.

[0067] Optionally, this embodiment may adopt Figure 7 the embodiment to implement S206, specifically including S401 to S403:

[0068] S401: Extract the event features of the key video frame.

[0069] In a specific embodiment, the terminal device can determine whether there is an abnormal event in the key video frame by comparing the event similarity between the event features in the key video frame and the preset event features. Specifically, the terminal device extracts the event features of the key video frame.

[0070] S402: Obtain the event similarity between the event features and the preset abnormal event features.

[0071] Among them, the preset abnormal event features can be event features such as cargo spillage event features, traffic accident event features, smoke and fog event features, and road construction event features, etc. The terminal device calculates the event similarity between the event features and the preset abnormal event features.

[0072] S403: Determine whether the event similarity is greater than or equal to the similarity threshold.

[0073] Among them, the terminal device determines whether the event similarity between the event feature and the preset abnormal event feature is greater than or equal to the similarity threshold. If so, execute S404. If not, it is determined that there is no abnormal event in the monitored video stream.

[0074] S404: Then it is determined that there is an abnormal event in the monitored video stream.

[0075] Among them, when the terminal device determines that the event similarity is greater than or equal to the similarity threshold, it is determined that there is an abnormal event in the monitored video stream.

[0076] In the above solution, the terminal device sets the first cache unit and the second cache unit, and uses the first cache unit and the second cache unit to store video frames of a preset number of frames, avoiding waste of storage space for segments of the monitored video stream, and at the same time improving the real-time performance of traffic abnormal event detection; uses the 2D feature encoding unit in the key frame recommendation network to extract the spatial features of each video frame in the first cache unit and the second cache unit, uses the bidirectional LSTM unit to associate the correlation between the spatio-temporal features before and after N video frames in the first cache unit and the second cache unit, increasing the utilization rate of video segment features, and uses the feed-forward network unit to fuse the high-level semantic features associated by the bidirectional LSTM unit, improving the real-time performance and accuracy of abnormal event detection in the monitored video stream; performs inference processing on the scores of the video frames output by the key frame network to further determine the accuracy of traffic abnormal event detection; uses the event similarity between the event feature of the key video frame and the preset event feature to determine whether there is an abnormal event in the monitored video stream, improving the accuracy of traffic abnormal detection.

[0077] Furthermore, on the basis that the terminal device in the above embodiment uses the event similarity between the event feature of the key frame and the preset event feature to determine the abnormal event. The terminal device in this embodiment uses whether the center point of the abnormal event area of the abnormal event in the key video frame is within the detection area, and then determines the abnormal event category and the abnormal event area of the abnormal event in the key video frame. For reference, Figure 8 After S404, the following steps are further included:

[0078] S501: Detect the abnormal event category and the abnormal event area of the abnormal event in the key video frame.

[0079] In order to improve the detection accuracy of the abnormal event category and the abnormal event area of the abnormal event in the key video frame, the terminal device performs pre-detection processing on the key video to obtain a pre-detection result. The pre-detection result includes the abnormal event category and the abnormal event area. Specifically, the terminal device uses the pre-detection network to perform pre-detection processing on the key video frame. Among them, for the detailed schematic diagram of the pre-detection network, reference can be made toFigure 9 Specifically, Figure 9 the label in it is the abnormal event category. The abnormal event categories are cargo spillage, traffic accidents, fireworks and fog, and road construction, etc. Location is the abnormal event area. The abnormal event area is the coordinate of the abnormal event area in the key video frame.

[0080] In addition, the pre-detection network can be neural networks such as YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector), Faster-Rcnn (Faster-Region-Convolutional Neural Networks), CenterNet, etc. This embodiment does not make a limitation on this.

[0081] S502: Obtain a preset detection area, and determine whether the center point of the abnormal event area of the abnormal event in the key video frame is within the detection area.

[0082] To further determine the pre-detection result, the terminal device determines whether the center point of the abnormal event area of the abnormal event in the key video frame is within the detection area. If so, execute S503.

[0083] S503: Then output the abnormal event category and the abnormal event area of the abnormal event in the key video frame.

[0084] Among them, when the terminal device determines that the center point of the abnormal event area is within the detection area, it determines that the pre-detection result is the final detection result, and outputs the abnormal event category and the abnormal event area of the abnormal event in the key video frame.

[0085] In the above solution, the terminal device uses whether the center point of the abnormal event area of the abnormal event in the key video frame is within the detection area, and then outputs the abnormal event category and the abnormal event area of the abnormal event in the key video frame, thereby improving the accuracy of traffic anomaly detection.

[0086] You can continue to refer to Figure 10 , Figure 10 which Figure 2 is a schematic flowchart of another embodiment of S206 in the traffic abnormal event method shown in. Specifically, S206 further includes the following steps:

[0087] S601: Detect the abnormal event category and the abnormal event area of the abnormal event in the key video frame.

[0088] S602: Obtain a preset detection area, and determine whether the center point of the abnormal event area of the abnormal event in the key video frame is within the detection area.

[0089] Among them, for the detailed descriptions of S601 - S602 in this embodiment, reference can be made to S501 - S502 in the above - mentioned embodiment.

[0090] S603: Obtain the event confidence of the abnormal event.

[0091] To improve the accuracy of the detection result, the terminal device in this embodiment can also perform filtering processing on the pre - detection result. Specifically, the terminal device considers the influence of the event confidence on the abnormal event category and the abnormal event area, and uses the event confidence to filter the pre - detection result, so as to determine the final detection result. Specifically, the terminal device first obtains the event confidence of the abnormal event.

[0092] S604: Determine whether the event confidence is greater than or equal to the preset confidence threshold.

[0093] Among them, the terminal device determines whether the event confidence is greater than or equal to the preset confidence threshold. If so, execute S605. If not, directly filter the detection result of the current video frame.

[0094] Further, in a specific embodiment, the terminal device can also consider the influence of the perimeter area on the abnormal event category and the abnormal event area. Specifically, when the terminal device determines that the event confidence is greater than or equal to the preset confidence threshold, it obtains the previous abnormal event with the same abnormal event category as the current abnormal event. And calculate the intersection - union ratio of the abnormal event area of the current abnormal event and the abnormal event area of the previous abnormal event. Then determine whether the intersection - union ratio is less than or equal to the preset intersection - union ratio threshold. If so, execute S605.

[0095] In other embodiments, the terminal device can also consider the influence of repeated alarm filtering at the same location on the abnormal event category and the abnormal event area. Specifically, when the terminal device determines that the event confidence is greater than or equal to the preset confidence threshold, it obtains the first video frame number corresponding to the current abnormal event, and obtains the second video frame number corresponding to the previous abnormal event. Then determine whether the difference between the first video frame number and the second video frame number is greater than or equal to the preset frame number difference threshold. If so, execute S605.

[0096] It should be noted that in practical applications, the terminal device can consider the influence of at least one of the event confidence, the perimeter area, and the repeated alarm at the same location on the abnormal event category and the abnormal event area. Of course, the terminal device can also consider the influence of at least two of the event confidence, the perimeter area, and the repeated alarm at the same location on the abnormal event category and the abnormal event area. This embodiment does not limit this. In addition, the ways for the terminal device to filter the pre - detection result include but are not limited to the event confidence, the perimeter area, and the repeated alarm at the same location.

[0097] S605: Output the abnormal event category and the abnormal event area.

[0098] Among them, the terminal device outputs the abnormal event category and the abnormal event area.

[0099] In the above solution, the terminal device considers at least one of event confidence, perimeter area, and repeated alarms at the same location for the impact on the abnormal event category and the abnormal event area, and uses at least one of event confidence, perimeter area, and repeated alarms at the same location to filter the pre-detection results, improving the detection accuracy of the abnormal event category and the abnormal event area, and achieving the balance of accuracy and real-time performance in detecting the abnormal event category and the abnormal event area.

[0100] You can continue to refer to Figure 11 , Figure 11 which Figure 1 is the schematic flowchart of obtaining the key frame recommendation network in the traffic abnormal event detection method shown in

[0101] S701: Obtain a number of training images, and obtain the label score of each training image; and obtain a reference image, where the reference image includes images with a label score greater than the label score threshold.

[0102] Among them, the training set of the key frame recommendation network can be obtained by setting a camera above the traffic road surface, shooting the monitoring images in the monitoring area of the camera, and storing the monitoring images in the terminal device where the key frame recommendation network is applied. Historical monitoring images can also be extracted from the database. By using the method of setting a camera above the traffic road surface, the camera can be installed at any position above the traffic road surface so that the camera is sufficient to shoot the traffic road surface conditions, and the number of cameras can be set to one or more. In this embodiment, the camera is installed directly above the traffic road surface so that the camera can obtain a better training set. In addition, the reference image is manually selected. Specifically, the training set can include images of cargo spills, traffic accidents, smoke and fog, and road construction. The range of the label score threshold is 0 to 1. Specifically, it can be 0.95 or 1. This embodiment does not limit this. The label score of each training image can also be manually set.

[0103] S702: Input each training image into the key frame recommendation network, and obtain the cosine similarity between each training image and the reference image as the prediction score of each training image.

[0104] Among them, the cosine similarity between each training image and the reference image satisfies the following formula:

[0105]

[0106] Among them, A represents the features of the training image, B represents the features of the reference image, and cosθ / similarity is the cosine similarity between the training image and the reference image.

[0107] Specifically, the terminal device uses the cosine similarity between each training image and the reference image as the prediction score of each training image.

[0108] S703: Calculate the loss function of the key frame recommendation network based on the label score and prediction score of each training image.

[0109] In order to constrain the correlation between images and extract the event structure features between images, the terminal device in this embodiment determines the loss function between the label score and prediction score of each training image by using the mean square error of the label score and prediction score of each training image, the prediction score deviation value of adjacent frame training images, and the label score deviation value of adjacent frame training images.

[0110] Specifically, the terminal device calculates the mean square error of the label score and prediction score of each training image, the prediction score deviation value of adjacent frame training images, and the label score deviation value of adjacent frame training images respectively. And uses the mean square error, prediction score deviation value, and label score deviation value to calculate the loss function of the key frame recommendation network. Among them, the loss function includes the degree of difference between the label score and prediction score of each training image.

[0111] Specifically, the loss function of the key frame recommendation network satisfies the following formula:

[0112]

[0113] Among them, SobLoss(s, t) is the loss function between the label score and prediction score of the training image, s is the prediction score, and t is the label score of the training image. is the mean square error of the label score and prediction score of each training image, s i+1 -s i is the prediction score deviation value of adjacent frame training images, t i+1 -t i is the label score deviation value of adjacent frame training images.

[0114] S704: Train the key frame recommendation network using the loss function.

[0115] Based on the loss function obtained in S703, the terminal device trains the key frame recommendation network with the goal of reducing the loss between the label score and prediction score of each training image, and obtains a key frame recommendation network that meets the requirements. Specifically, the terminal device can determine whether the loss between the label score and prediction score of each training image is less than or equal to the loss threshold. If so, the key frame recommendation network is obtained.

[0116] In the above solution, the terminal device determines the loss function between the label score and the prediction score of each training image by using the mean square error of the label score and the prediction score of each training image, the prediction score deviation value of adjacent frame training images, and the label score deviation value of adjacent frame training images, so as to determine the key frame recommendation network based on the loss function, thereby constraining the correlation between images, extracting the event structure features between images, and realizing the direct output of key frames of abnormal events.

[0117] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0118] To implement the traffic abnormal event detection method of the above embodiment, the present application also proposes a terminal device. For details, please refer to Figure 12 , Figure 12 which is a schematic structural diagram of an embodiment of the terminal device provided by the present application.

[0119] The terminal device 120 in the embodiment of the present application includes a memory 121 and a processor 122, wherein the memory 121 and the processor 122 are coupled.

[0120] The memory 121 is used to store program data, and the processor 122 is used to execute the program data to implement the acquisition of the monitoring video stream described in the above embodiment; based on the spatio-temporal feature correlation between each video frame and its associated video frame in the monitoring video stream, estimate the key frame reference value of each video frame; and based on the key frame reference value of each video frame, determine the key video frames from the monitoring video stream; based on the key video frames, determine whether there is an abnormal event in the monitoring video stream.

[0121] In this embodiment, the processor 122 can also be called a CPU (Central Processing Unit). The processor 122 may be an integrated circuit chip with signal processing capabilities. The processor 122 may also be a general-purpose processor, a digital signal processor (DSP, Digital Signal Process), an application specific integrated circuit (ASIC, Application Specific Integrated Circuit), a field programmable gate array (FPGA, FieldProgrammable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor 122 may also be any conventional processor, etc.

[0122] The present application also provides a computer storage medium, such as Figure 13 As shown, the computer storage medium 130 is used to store program data 131. When the program data 131 is executed by a processor, it is used to implement obtaining a monitoring video stream as described in the above embodiments; estimating a key frame reference value for each video frame based on the spatio-temporal feature correlation between each video frame and its associated video frame in the monitoring video stream; and determining key video frames from the monitoring video stream based on the key frame reference value of each video frame; and determining whether there is an abnormal event in the monitoring video stream based on the key video frames.

[0123] The present application also provides a computer program product. The computer program product includes a computer program, and the computer program is operable to cause a computer to execute the traffic abnormal event detection method as described in the embodiments of the present application. The computer program product can be a software installation package.

[0124] When the traffic abnormal event detection method described in the above embodiments of the present application exists in the form of a software functional unit and is sold or used as an independent product, it can be stored in a device, such as a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0125] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A method for detecting traffic abnormal events, characterized in that, comprising obtaining a video stream; estimating the key frame reference values of each video frame based on the spatio-temporal feature correlation between each video frame and its associated video frames in the video stream; and determining key video frames from the video stream based on the key frame reference values of each video frame; determining whether there is an abnormal event in the video stream based on the key video frames; presetting a first buffer unit and a second buffer unit; storing video frames of a preset number of frames in the video stream into the first buffer unit and the second buffer unit, and the video stream is alternately stored in the first buffer unit and the second buffer unit in the form of single-frame images, so that the first buffer unit and the second buffer unit respectively store N / 2 frame images; respectively obtaining the key frame reference values of the video frames in the first buffer unit and the second buffer unit; screening out key video frames based on the key frame reference values of all video frames.

2. The method according to claim 1, characterized in that, the associated video frames of a video frame in the video stream include at least one of the following video frames: the first number of video frames before the one video frame in the video stream; the second number of video frames after the one video frame in the video stream.

3. The method according to claim 1, characterized in that, the key frame recommendation network includes a 2D feature encoding unit, a bidirectional LSTM unit and a feedforward network unit; estimating the key frame reference values of each video frame based on the spatio-temporal feature correlation between each video frame and its associated video frames in the video stream; and determining key video frames from the video stream based on the key frame reference values of each video frame, including: inputting the video stream into the key frame recommendation network and obtaining the key video frames output by the key frame recommendation network; the key frame recommendation network includes a 2D feature encoding unit, a bidirectional LSTM unit and a feedforward network unit.

4. The method according to any one of claims 1-3, characterized in that, determining key video frames from the video frames based on the key frame reference values of each video frame, including: sequentially comparing the key frame reference values of each video frame with a preset key frame reference value threshold according to the sequence of each video frame in the video stream; determining the video frames with key frame reference values greater than the key frame reference value threshold as the first key video frames.

5. The method according to claim 4, characterized in that, the method further includes: sequentially calculating the similarity between a first key video frame and the previous first key video frame according to the sequence of each first key video frame in the video stream; when the similarity is less than or equal to a preset similarity threshold, discarding the previous first key video frame and taking the one first key video frame as the finally output second key video frame.

6. The method according to claim 3, characterized in that, The key frame recommendation network is obtained in the following manner: obtaining a plurality of training images, and obtaining the label scores of each training image; and obtaining reference images, where the reference images include images with label scores greater than a label score threshold; Inputting each training image into the key frame recommendation network, and obtaining the cosine similarity between each training image and the reference images as the prediction score of each training image; Calculating the loss function of the key frame recommendation network based on the label scores and prediction scores of each training image; Adjusting the parameters of the key frame recommendation network based on the loss function.

7. The method according to claim 6, wherein, Calculating the loss function of the key frame recommendation network based on the label scores and prediction scores of each training image includes: Calculating the mean square error of the label scores and prediction scores of the plurality of training images; Calculating the prediction score deviation value of adjacent frame training images and the label score deviation value of adjacent frame training images; Calculating the loss function of the key frame recommendation network using the mean square error, the prediction score deviation value, and the label score deviation value.

8. The method according to claim 1, wherein, Determining whether there is an abnormal event in the video stream based on the key video frames includes: Extracting the event features of the key video frames; Obtaining the event similarity between the event features and preset abnormal event features; If the event similarity is greater than or equal to a similarity threshold, determining that there is an abnormal event in the video stream; If the event similarity is less than the similarity threshold, determining that there is no abnormal event in the video stream.

9. The method according to claim 8, wherein, After determining that there is an abnormal event in the video stream, it includes: Detecting the abnormal event category and abnormal event area of the abnormal event in the key video frames; Obtaining a preset detection area, and detecting whether the center point of the abnormal event area of the abnormal event in the key video frames is within the detection area; If so, outputting the abnormal event category and abnormal event area of the abnormal event in the key video frames.

10. The method according to claim 9, wherein, The method further includes: Calculating the intersection over union of the abnormal event areas in two key video frames with the same abnormal event type in sequence according to the order of the respective key video frames in the video stream; When the intersection over union is less than or equal to a preset intersection over union threshold, discarding one of the two key video frames with the same abnormal event type and outputting the other key video frame.

11. The method according to claim 9 or 10, wherein, The method further includes: Calculating the difference in video frame sequence numbers of two key video frames with the same abnormal event type in sequence according to the order of the respective key video frames in the video stream; When the difference between the video frame sequence numbers is less than or equal to a preset frame sequence number difference threshold, one of the key video frames of the two abnormal events with the same abnormal event type is discarded, and the other key video frame is output.

12. A terminal device Characterized in that The terminal device includes a memory and a processor, wherein the memory is coupled to the processor; Wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the method according to any one of claims 1-11.

13. A computer storage medium Characterized in that The computer storage medium is used to store program data, and when the program data is executed by a processor, it is used to implement the method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Self-attention video abstraction method based on distribution consistency

    CN110287374A

  • Road information monitoring and analysis detection method and intelligent traffic control system

    CN111241343A