Traffic anomaly detection method, electronic device and storage medium
By combining the dual-stream network architecture and the deep prediction loss function, accurate detection of traffic anomalies is achieved, solving the problem of insufficient detection accuracy in existing technologies and improving detection efficiency and accuracy.
Patent Information
- Application Number
- CN202410899768.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-05
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-07-05
AI Technical Summary
The accuracy of traffic anomaly detection in existing technologies still needs to be improved, especially when using video data for detection.
It adopts a dual-stream network architecture, through the encoding of RGB images and optical flow images, multi-scale attention fusion and feature reconstruction of memory modules, combined with the deep prediction loss function, to achieve accurate detection of traffic anomalies.
The accuracy and robustness of traffic anomaly detection are improved, and multiple types of traffic anomalies can be detected in real time, reducing the probability of false detection under normal circumstances.
Smart Images

Figure CN119169541B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of traffic detection, and in particular to a traffic anomaly detection method, electronic equipment, and storage medium. Background Art
[0002] With the development of transportation, traffic anomalies such as traffic accidents, congestion, and nighttime conditions are increasingly impacting road safety and traffic efficiency. Accurately detecting traffic anomalies using traffic video is an urgent issue.
[0003] There are also some methods in the existing technology that use video data to automatically detect traffic anomalies. For example, patent CN105405297B provides an automatic traffic accident detection method based on surveillance video, and patent CN113221716A provides an unsupervised traffic abnormal behavior detection method based on foreground target detection. However, the accuracy of detection still needs to be improved. Summary of the Invention
[0004] Embodiments of the present invention provide a traffic anomaly detection method, an electronic device, and a storage medium to improve the accuracy of traffic anomaly detection.
[0005] In a first aspect, an embodiment of the present invention provides a method for detecting traffic anomalies, comprising:
[0006] S110, obtaining a continuous multi-frame traffic video to be detected;
[0007] S120, respectively encode the RGB images and optical flow images of the first few frames of video using two encoders, and each layer of each encoder outputs multi-layer RGB features and multi-layer optical flow features respectively; wherein, when encoding the RGB image, the RGB features output by each layer of the RGB encoder are added to the optical flow features output by the corresponding layer of the optical flow encoder and then input into the next layer of the RGB encoder, thereby realizing shallow information interaction between RGB and optical flow;
[0008] S130, performing multi-scale attention fusion on the global RGB features and global optical flow features output by the last layer of the two encoders to obtain a deep fusion feature of RGB and optical flow;
[0009] S140. Reconstructing the deep fusion features and the global optical flow features based on the normal features recorded in the two memory modules, respectively, wherein the normal features in the two memory modules represent the characteristics of the deep fusion features and the global optical flow features of the normal traffic video, respectively;
[0010] S150, decoding the two reconstructed features respectively through two decoders to restore the RGB image and optical flow image of the last frame of video respectively;
[0011] S160 : Determine whether the last frame of video has a traffic anomaly risk based on the difference between each restored image and the original image, and the difference between each reconstructed feature and the normal feature.
[0012] In a second aspect, an embodiment of the present invention provides an electronic device, comprising:
[0013] one or more processors;
[0014] a memory for storing one or more programs,
[0015] When the one or more programs are executed by the one or more processors, the one or more processors implement the traffic anomaly detection method described in any embodiment.
[0016] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the traffic anomaly detection method described in any embodiment.
[0017] In summary, the embodiments of the present invention provide a traffic anomaly detection method, electronic device, and storage medium. These methods incorporate optical flow data and utilize inter-stream information interaction to fully fuse RGB and optical flow data to more accurately capture anomaly features. These captured dual-stream features are then subjected to multi-scale attention fusion to better reflect contextual and multi-scale information. Subsequently, a memory module, integrated with the dual-stream framework, performs feature reconstruction, further reducing the probability of false detection in normal situations. This embodiment employs an end-to-end dual-stream network architecture to train the model, significantly improving model training and prediction efficiency. Furthermore, a depth-based prediction loss function is employed to adapt to the regular shapes of vehicles and events at varying distances from surveillance cameras, enhancing the accuracy and robustness of the model's anomaly detection. Furthermore, given that datasets dedicated to traffic anomaly detection typically only include a single type of anomaly (i.e., traffic accidents), this embodiment incorporates a traffic monitoring dataset for full-class anomaly detection model validation. This allows a single model to detect all types of anomalies in traffic scenarios, further improving detection efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 This is a flow chart of a traffic anomaly detection method provided by an embodiment of the present invention;
[0020] Figure 2 is a schematic diagram of multiple frames of RGB images and optical flow images provided by an embodiment of the present invention;
[0021] Figure 3 is a flow chart of another traffic anomaly detection method provided by an embodiment of the present invention;
[0022] Figure 4 A schematic diagram of a multi-scale attention fusion module provided by an embodiment of the present invention;
[0023] Figure 5 A schematic diagram of feature reconstruction using a memory module provided by an embodiment of the present invention;
[0024] Figure 6 A comparison chart of the TSD provided by an embodiment of the present invention and the existing public transportation dataset, wherein: Figure 6 (a) is a schematic diagram of the AI City dataset. Figure 6 (b) is a schematic diagram of the DoTA dataset. Figure 6 (c) is a schematic diagram of the TSD of this embodiment;
[0025] Figure 7 A schematic diagram of a training method for a traffic anomaly detection model provided by an embodiment of the present invention;
[0026] Figure 8 A schematic diagram of updating a memory module provided by an embodiment of the present invention;
[0027] Figure 9 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0028] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are also within the scope of protection of the present invention.
[0029] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0030] In the description of the present invention, it should also be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0031] Figure 1 This is a flow chart of a traffic anomaly detection method provided by an embodiment of the present invention. The method is applicable to traffic anomaly detection on roads where video acquisition equipment is deployed, and is particularly applicable to traffic anomaly detection on highways. The method is executed by an electronic device, such as Figure 1 As shown, the method specifically includes:
[0032] S110 : Acquire multiple frames of continuous traffic video to be detected, where each frame of the video includes an RGB image and an optical flow image.
[0033] This embodiment will use RGB images and optical flow images to achieve road traffic anomaly detection. Figure 2 The RGB images and optical flow images of multiple frames of video in a traffic scene are shown as examples, where T represents the time of each frame. Optionally, RGB images and optical flow images of multiple frames of video can be obtained through roadside monitoring and used as the data source for the entire method.
[0034] S120. Encode the RGB images and optical flow images of the first few frames of video respectively through two encoders, and each layer of each encoder outputs multi-layer RGB features and multi-layer optical flow features respectively; wherein, when encoding the RGB image, the RGB features output by each layer of the RGB encoder are added to the optical flow features output by the corresponding layer of the optical flow encoder and then input into the next layer of the RGB encoder, thereby realizing shallow information interaction between RGB and optical flow.
[0035] This embodiment uses a dual-stream architecture to process the RGB images and optical flow images of the first few frames of video separately to extract traffic anomaly information in the data, and predict the RGB image and optical flow image of the last frame of video based on this information. The entire data processing flow is as follows Figure 3 As shown, this step corresponds to the encoding module in the figure. Figure 3 , perform the following operations in the encoding module:
[0036] First, the RGB images of the first few frames of video are encoded by encoder 1, and each layer of encoder 1 outputs a layer of RGB features, wherein the size of the RGB features gradually decreases as the number of layers increases. At the same time, the optical flow images of the first few frames of video are encoded by encoder 2, and each layer of encoder 2 outputs a layer of optical flow features, and similarly, the size of the optical flow features gradually decreases as the number of layers increases. Furthermore, encoder 1 and encoder 2 have the same structure, and the feature sizes output by the same layer are also the same. In order to realize shallow information interaction between RGB and optical flow, in this embodiment, before the output features of each layer of encoder 1 enter the next layer, they are first added to the output features of the same layer in encoder 2, and then the result of the addition is input to the next layer of encoder 1. In this way, shallow optical flow feature maps of different scales can be integrated into the feature extraction process of the RGB feature encoder, thereby realizing comprehensive information interaction between RGB and optical flow. The entire operation can be expressed as follows:
[0037]
[0038] in, represents the original sequence optical flow map (i.e., the optical flow images of the first few frames of video). For example, assuming that N frames of video are obtained in S110, then is the optical flow image of the first N-1 frames of video; Represents the encoding operation of encoder 2; Represents the global optical flow feature map output by the last layer of encoder 2, Before encoder 2 b Layer output Optical flow features of different sizes, b +1 is the total number of encoder layers; Represents the encoding operation of encoder 1; represents the global RGB features output by the last layer of encoder 1, ,in are the number of channels, width and height respectively. Optionally, the two encoders can adopt a 10-layer convolution and pooling mixed structure, that is, b +1=10. In practical applications, other encoder structures may also need to be selected, and this embodiment does not impose any specific limitation.
[0039] S130, perform multi-scale attention fusion on the global RGB features and global optical flow features output by the last layer of the two encoders to obtain a deep fusion feature of RGB and optical flow :
[0040]
[0041] in, Represents a multi-scale attention fusion operation. Optional, Figure 4 A schematic diagram of a multi-scale attention fusion module provided by an embodiment of the present invention is shown in FIG. Figure 4 As shown, in this module, first and They are used as initial features, and self-attention and normalization are performed respectively to enhance the RGB and optical flow feature representations; then a simple element-by-element addition is performed to achieve initial feature fusion. The details are as follows:
[0042]
[0043] in, Represents the initial fused features (or called preliminary fused features), represents the self-attention mechanism, represents batch normalization, Indicates the addition of corresponding elements.
[0044] Then, the preliminary fusion features are convolved to extract local features. In order to cope with the scale changes of different information in real traffic scenes, especially to improve the analysis ability of small objects, this step combines context information with different receptive fields to integrate the preliminary fusion features. The multi-scale information in the image is used to enhance the performance of the entire framework when processing multi-scale and small-sized objects. Optionally, local features of contextual information can be extracted through point-by-point convolution operations. :
[0045]
[0046] in, Represent two convolutional layers respectively. represent Linear activation function.
[0047] Next, the local features Perform global average pooling to extract global features Then, based on the preliminary fusion features and global features, a deep fusion feature map of RGB and optical flow is generated. :
[0048]
[0049] in, Represents the multiplication of corresponding matrix elements.
[0050] S140. Reconstruct the deep fusion features and the global optical flow features respectively according to the normal features in the two memory modules, wherein the normal features in the two memory modules respectively represent the characteristics of the deep fusion features and the global optical flow features of the normal traffic video.
[0051] This embodiment introduces a memory module to store the diverse features of normal traffic videos and prevent abnormal videos from being mistakenly identified as normal patterns. Specifically, the memory module pre-stores multiple feature representations of normal traffic videos (referred to as normal features). These features, pre-determined through training on normal traffic videos, reflect the key characteristics of normal videos while also ensuring the diversity and comprehensiveness of normal patterns. The specific training process will be described in detail in subsequent implementations.
[0052] Furthermore, to adapt to the dual-stream processing framework of RGB images and diverted images, this embodiment constructs a memory module for each of the RGB images and the diverted images. Memory module 1 corresponding to the RGB image stores normal features that can be compared with the aforementioned deep fusion features, which can be constructed based on the deep fusion features generated during the training process of normal traffic videos. Memory module 2 corresponding to the optical flow image stores normal features that can be compared with the aforementioned global optical flow features, which can be constructed based on the global optical flow features generated during the training process of normal traffic videos.
[0053] For any memory module, Figure 5 A schematic diagram of feature reconstruction through a memory module provided by an embodiment of the present invention. Figure 5 , the memory module stores M normal features, where the mth normal feature is recorded as . For the features to be reconstructed (global optical flow features or deep fusion features), they can be divided into K blocks (or K items), and the distance between each item and each normal feature in the memory module is calculated respectively; the sum of the distances between the K items and the same normal feature is taken as the distance between the feature to be reconstructed and the normal feature. Finally, according to the distance between the feature to be reconstructed and each normal feature, the weighted sum of each normal feature is performed to obtain the reconstructed feature of the global optical flow feature. Optionally, the distance can be measured by cosine similarity, and can also be normalized by a softmax function, such as Figure 5 shown.
[0054] S150 , decoding the two reconstructed features respectively through two decoders, and restoring the RGB image and optical flow image of the last frame of video respectively.
[0055] This step corresponds to Figure 3The decoding module in the decoder uses decoder 1 to decode the reconstructed features of the deep fusion features to restore the RGB image of the last frame of video; at the same time, decoder 2 is used to decode the reconstructed features of the global optical flow features to restore the optical flow image of the last frame of video.
[0056] S160 : Determine whether the last frame of video has a traffic anomaly risk based on the difference between each restored image and the original image, and the difference between each reconstructed feature and the normal feature.
[0057] The reconstructed features here still refer to the deep fusion features and the global optical flow features. As described in the memory module, the entire data processing architecture of this embodiment ( Figure 3 ) is constructed based on normal traffic videos. The above two differences corresponding to normal traffic videos will be significantly different from the above two differences of abnormal traffic videos, and thus can be used as a basis for judging whether the input video has the risk of traffic abnormality.
[0058] Optionally, this step calculates an anomaly score for each of the RGB restored image and the optical flow restored image. Taking the RGB restored image as an example, the predicted N-th frame RGB image can be calculated. Compared with the actual Nth frame RGB image Peak Signal-to-Noise Ratio (PSNR) between
[0059]
[0060] in, Represents the number of pixels in the video frame. The larger the PSNR, the greater the difference between the two images.
[0061] At the same time, the deep fusion feature and the nearest normal feature are calculated Distance:
[0062]
[0063] in, Represents the reconstructed feature, Represents the reconstructed feature The larger the distance, the greater the difference between the two features.
[0064] Finally, based on the PSNR and distance, the anomaly score of the RGB image is calculated :
[0065]
[0066] in, Normalization, represents the equilibrium parameter.
[0067] Regarding the anomaly score of the optical flow image, its calculation method is similar to that of the RGB image. The Nth frame RGB image is replaced by the Nth frame optical flow image, and the deep fusion feature is replaced by the global optical flow feature. After obtaining the two anomaly scores, the weighted average of the two or the maximum value can be taken as the final anomaly score. When the final anomaly score is higher than the set threshold, it is considered that the video data to be detected has a traffic anomaly risk. In practical applications, the method of this embodiment can be deployed on edge devices, and when the roadside monitoring device obtains the current moment t After the video data is received, the edge device t Fast prediction of several frames of video data before the moment t Video data at each moment, and with t Compare with the real video data at the moment to judge t Whether there is a risk of traffic anomaly at any time, and realize real-time detection of traffic anomalies.
[0068] Furthermore, the entire Figure 3 The data processing flow constitutes a traffic anomaly detection model, and the training process of the model is expanded below. In a specific embodiment, first, a traffic surveillance dataset (TSD) is constructed. Specifically, the video anomaly datasets in the prior art focus more on pedestrians, lack attention to vehicle anomalies in road traffic scenes, and cannot effectively verify the anomaly detection of the road traffic environment. To this end, this embodiment constructs a new TSD for full-category anomaly detection in real traffic scenes. TSD consists of high-quality, unedited videos from real monitoring angles, covering various traffic anomalies occurring in 52 real-life scenes, including accidents, traffic jams, and some night shots. Table 1 lists the comparison of TSD with commonly used datasets for full-category anomaly detection. Figure 6 The intuitive comparison between TSD and some existing public transportation datasets is shown. Figure 6 The AICity dataset in (a) contains many non-road anomalies (e.g., camera shake), which can interfere with it; Figure 6 The DoTA dataset in (b) is based on the perspective of autonomous driving and cannot integrate abnormal data from the entire road network; Figure 6 The TSD of this embodiment in (c) is constructed based on video surveillance angles and is a high-quality full-category anomaly detection dataset for traffic scenes.
[0069] Table 1
[0070]
[0071] Then, from the normal traffic video data set of TSD, multiple consecutive frames of video are extracted as training samples to obtain a training sample set for the traffic anomaly detection model. Due to the rarity of abnormal events, this embodiment uses part of the normal video data in TSD for model training. During the training process, operations S120-S150 are performed on each sample respectively, and a loss function is constructed by minimizing the difference between each restored image and the original image, as well as minimizing the difference between the reconstructed features and the normal features. Since there are large differences between normal traffic videos and abnormal traffic videos, the above training process can ensure that the trained model obtains a lower anomaly score for normal traffic videos and a higher anomaly score for abnormal traffic videos in actual detection, thereby effectively identifying whether there is a risk of traffic anomaly. After the training is completed, a part of the test samples can be collected from both normal traffic videos and abnormal traffic videos in the testing phase to verify the model's ability to recognize traffic anomalies.
[0072] Optionally, the following loss function can be constructed during training:
[0073]
[0074] in, L represents the total loss, represent the depth-based prediction losses of RGB images and optical flow images, respectively. Respectively represent the compactness loss of RGB normal features and optical flow normal features in the memory module, Represents the separation loss of RGB normal features and optical flow normal features in the memory module. These losses are separated by hyperparameters Balance. The calculation method of each loss is explained in detail below.
[0075] First, the depth-based prediction loss is explained. Generally speaking, the prediction loss penalizes the inconsistency between the prediction and the ground truth, ensuring that the decoder's prediction closely matches the ground truth, with each pixel having the same weight. However, in typical traffic monitoring scenarios, the camera has a certain depth of field, which means that for pixel blocks of the same size, pixel blocks farther from the camera contain more scene information. Therefore, this embodiment adopts a depth-based prediction loss strategy to assign greater weights to pixels farther from the camera to increase their contribution, so that the model's adjustment process is consistent with real-world monitoring scenarios, which helps the model better understand and process real-world scene data and improve detection performance.
[0076] Optional, such as Figure 7As shown in the figure, for any restored image, it is first layered along the height direction, and then the pixels of each layer are average pooled at different scales according to the distance between each layer and the camera. For example, the bottom layer image is closest to the camera, so the bottom layer image is average pooled in a 64×64 block, and the upper layer image is relatively far from the camera, so the layer image is average pooled in a 32×32 block. In general, the closer to the camera, the larger the scale of the average pooling. This is because the main subjects in real traffic scenes (especially highways) are vehicles, and average pooling of different scales can adapt to the regular shape characteristics of traffic vehicle data. Taking RGB images as an example, this operation can be expressed as:
[0077]
[0078] in, The image is divided into A layer; Indicates the a The average pooling strategy used within the layer; Indicates the height of each layer.
[0079] Then, the average pooled image is divided into blocks, and the difference between each block and the original image block is weighted averaged according to the distance from each block to the camera to construct a depth-based prediction loss function. Still taking the RGB image as an example:
[0080]
[0081] in, Represents the number of pixel blocks that the image is divided into in the width and height dimensions respectively; is assigned to the coordinate position The weight of the pixel block where it is located. The farther away from the camera, the greater the weight. Respectively represent the coordinate positions of the original image and the restored image The pixel value at . Through The L2 distance between the decoder output and the true value is scaled by the constraint Minimizing the distance between the decoder output and the true value can be minimized.
[0082] Then, combined with the update operation of the memory module during training, the feature compactness loss and feature separation loss are explained. Specifically, during training, after each sample completes the reconstruction of the global optical flow feature and the deep fusion feature, the normal features in the memory module will be updated. Figure 8 , the deep fusion features of any sample include K items, first select the normal feature farthest from the deep fusion feature Iu As the update object, then according to the deep fusion features and the normal features I u The distance of each item is weighted averaged and the I u By updating the features, various normal features can be effectively recorded, thereby ensuring the diversity and time correlation of the normal features.
[0083] Based on the reconstruction and update process of the memory module, taking the deep fusion feature as an example, the following feature compactness loss can be constructed:
[0084]
[0085] in, Represents the normal feature closest to the reconstructed feature within the memory module. Feature compactness loss can make the reconstructed feature closer to the nearest normal feature in the memory module, thereby reducing feature differences within the normal category.
[0086] At the same time, the following feature separability loss can be constructed:
[0087]
[0088] in, Represents the normal feature in the memory module that is second closest to the reconstructed feature, A hyperparameter that maintains the basic distance between classes. Feature separability loss allows the normal features in the memory module to record different normal patterns in various scenarios, thus maintaining the diversity of normal scenarios.
[0089] In the iterative training of each sample, the model parameters can be updated by minimizing the total loss L. Among them, the depth-based prediction loss can ensure that the difference between each restored image and the original image is minimized, the feature compactness loss can ensure that the difference between the optical flow features or depth fusion features corresponding to each restored image and the normal features closest to each memory module is minimized, and the feature separability loss can maximize the difference between each normal feature in each memory module.
[0090] In summary, this embodiment provides a traffic anomaly detection method, electronic device, and storage medium. It incorporates optical flow data and leverages inter-stream information interaction to fully fuse RGB and optical flow data to more accurately capture anomaly features. The captured dual-stream features are then subjected to multi-scale attention fusion to better reflect context and multi-scale information. Feature reconstruction is achieved through a memory module integrated with the dual-stream framework, further reducing the probability of false detection in normal situations. This embodiment employs an end-to-end dual-stream network architecture to train the model, significantly improving model training and prediction efficiency. It also uses a depth-based prediction loss function to accommodate the regular shapes of vehicles and events at varying distances from the surveillance camera, enhancing the accuracy and robustness of the model's anomaly detection. Furthermore, given that datasets dedicated to traffic anomaly detection typically only include a single type of anomaly (i.e., traffic accidents), this embodiment introduces a traffic monitoring dataset for full-class anomaly detection model validation. This allows a single model to detect all types of anomalies in traffic scenarios, further improving detection efficiency and accuracy.
[0091] Figure 9 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention is shown in FIG. Figure 9 As shown, the device includes a processor 60, a memory 61, an input device 62 and an output device 63; the number of processors 60 in the device can be one or more. Figure 9 In the embodiment, a processor 60 is used as an example; the processor 60, the memory 61, the input device 62 and the output device 63 in the device can be connected by a bus or other means. Figure 9 The bus connection is taken as an example.
[0092] Memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the traffic anomaly detection method in the embodiments of the present invention. Processor 60 executes the software programs, instructions, and modules stored in memory 61 to perform various functional applications and data processing of the device, thereby implementing the aforementioned traffic anomaly detection method.
[0093] The memory 61 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal. Furthermore, the memory 61 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, the memory 61 may further include memory remotely located relative to the processor 60, and these remote memories may be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0094] The input device 62 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The output device 63 may include a display device such as a display screen.
[0095] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the traffic anomaly detection method of any embodiment.
[0096] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device.
[0097] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0098] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0099] Computer program code for performing the operations of the present invention can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, Python, and conventional procedural programming languages such as C or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.
Claims
1. A traffic anomaly detection method, characterized in that: include: S110, obtaining a continuous multi-frame traffic video to be detected; S120, respectively encode the RGB images and optical flow images of the first few frames of video using two encoders, and each layer of each encoder outputs multi-layer RGB features and multi-layer optical flow features respectively; wherein, when encoding the RGB image, the RGB features output by each layer of the RGB encoder are added to the optical flow features output by the corresponding layer of the optical flow encoder and then input into the next layer of the RGB encoder, thereby realizing shallow information interaction between RGB and optical flow; S130, performing multi-scale attention fusion on the global RGB features and global optical flow features output by the last layer of the two encoders to obtain a deep fusion feature of RGB and optical flow; S140. Reconstructing the deep fusion features and the global optical flow features based on the normal features recorded in the two memory modules, respectively, wherein the normal features in the two memory modules represent the characteristics of the deep fusion features and the global optical flow features of the normal traffic video, respectively; S150, decoding the two reconstructed features respectively through two decoders to restore the RGB image and optical flow image of the last frame of video respectively; S160: Determine whether the last frame of video has a traffic anomaly risk based on the difference between each restored image and the original image, and the difference between each reconstructed feature and the normal feature; Among them, before S120, it also includes: Obtain a normal traffic video dataset and extract multiple consecutive frames of video as samples; Each encoder and decoder is trained using each sample. During the training process, operations S120-S150 are performed on each sample respectively, and a loss function is constructed by minimizing the difference between each restored image and the original image, as well as minimizing the difference between each reconstructed feature and the normal feature. Specifically, any restored image and the original image are layered along the height direction respectively; according to the distance from each layer to the camera, average pooling of pixels in each layer at different scales is performed respectively; each average pooled image is divided into blocks; according to the distance from each block to the camera, the difference between each block and the original image block is weighted averaged; and a loss function is constructed to minimize the difference after weighted averaging.
2. The method according to claim 1, characterized in that The multi-scale attention fusion of the global RGB features and the global optical flow features output by the last layer of the two encoders to obtain the deep fusion features of RGB and optical flow includes: Performing self-attention operations on the global RGB features and the global optical flow features respectively, and fusing the two self-attention results into a preliminary fused feature; Performing a convolution operation on the preliminary fusion features to extract local features; Perform global average pooling on each local feature to extract global features; Based on the preliminary fusion features and the global features, a deep fusion feature of RGB and optical flow is generated.
3. The method according to claim 1, characterized in that The reconstructing the deep fusion feature and the global optical flow feature respectively according to the normal features in the two memory modules includes: Calculating the distance between the global optical flow feature and each normal feature in the corresponding memory module respectively; A weighted sum is performed on each normal feature according to each distance to obtain a reconstructed feature of the global optical flow feature.
4. The method according to claim 1, wherein The determining whether the last frame of video has a traffic anomaly risk based on the difference between each restored image and the original image, and the difference between each reconstructed feature and the normal feature, includes: Calculate the PSNR between the RGB restored image and the original image; Calculating the distance between the deep fusion feature and the nearest normal feature; Based on the PSNR and distance, an anomaly score is calculated for the RGB image.
5. The method according to claim 1, wherein The loss function is constructed by minimizing the difference between each restored image and the original image, and minimizing the difference between each reconstructed feature and the normal feature, including: The loss function is constructed by minimizing the difference between each restored image and the original image, minimizing the difference between the optical flow features or deep fusion features corresponding to each restored image and the nearest normal features in their respective memory modules, and maximizing the difference between each normal feature in each memory module.
6. The method according to claim 1, characterized in that After performing the operations S120 to S150 on each sample during the training process, the method further includes: Decompose the deep fusion features of any sample into K items; Calculate the distance between each item and each normal feature in the corresponding memory module respectively; Determine the distance between the deep fusion feature and the same normal feature based on the distance between each item and the same normal feature; The normal feature that is farthest from the deep fusion feature is updated to a result of weighted averaging each item according to its distance from the normal feature.
7. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the traffic anomaly detection method according to any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the program is executed by a processor, the traffic anomaly detection method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
An Automatic Detection Method for Traffic Accidents Based on Surveillance Videos
CN105405297B
Unsupervised traffic abnormal behavior detection method based on foreground target detection
CN113221716A
Multi-mode two-stage unsupervised video anomaly detection method
CN114332053A