Traffic video anomaly detection method and device based on memory enhancement and time-space correlation

By adopting memory enhancement and space-time correlation methods in traffic video abnormal detection, combined with the cross-attention fusion module, the problems of insufficient utilization of space-time information and semantic deviation in the prior art are solved, and the detection performance and robustness are significantly improved.

CN120107863APending Publication Date: 2025-06-06FUDAN UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510257811.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing traffic video anomaly detection methods fail to make full use of the intrinsic connections of space-time information, and the feature fusion mechanism of memory networks is prone to introduce semantic deviations, resulting in limited detection performance.

Method used

The detection method based on memory enhancement and space-time correlation is adopted to preserve the characteristics of normal events through appearance and motion memory pools, and combined with the cross attention fusion module and space-time correlation module, the model's understanding of space-time mode and feature fusion ability are enhanced.

Benefits of technology

It significantly improves the abnormal detection performance of the model in complex traffic scenarios, enhances robustness and generalization capabilities, and improves the accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107863A_ABST
    Figure CN120107863A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a traffic video anomaly detection method and device based on memory enhancement and space-time correlation, and the method comprises the steps: obtaining traffic monitoring video stream data, processing the traffic monitoring video stream data into continuous image frame data and corresponding inter-frame difference data, and transmitting the continuous image frame data and the corresponding inter-frame difference data to a server; dividing the image frame data into a training set and a test set; an anomaly detection model is constructed based on memory enhancement and space-time association, the anomaly detection model is trained based on the inter-frame difference data and the training set, the anomaly detection model records time features and space features through memory enhancement, and the anomaly detection model establishes association between the time features and the space features; and performing anomaly detection on the test set by using the trained anomaly detection model. Therefore, the problem of how to accurately model space-time correlation in traffic video anomaly detection is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence and data processing technology, and in particular to a traffic video anomaly detection method based on memory enhancement and spatiotemporal association. Background Art

[0002] Traffic video anomaly detection refers to accurately identifying unconventional behaviors or events, such as running a red light and car accidents, from continuous traffic surveillance video clips. Due to the scarcity and diversity of such abnormal events, it is difficult to collect and annotate enough data for building a training set, which poses a great challenge to model training. Therefore, unsupervised methods are usually used in related technologies to train models using only data sets of normal samples. In the testing phase, the model regards samples that deviate from the characteristics of normal samples as abnormal, thereby realizing the detection of abnormal events.

[0003] Related technologies can generally be divided into reconstruction-based or prediction-based methods. Reconstruction-based methods usually take a single video frame as input, reconstruct the input video frame, and determine the frame with a large reconstruction error as an abnormal frame. Prediction-based methods take multiple consecutive video frames as input, predict future video frames, and determine the frame with a large prediction error as an abnormal frame. Due to the powerful generalization ability of the autoencoder, abnormal frames can sometimes be reconstructed or predicted very well, resulting in a decrease in detection performance. Related technologies also propose the use of memory networks to solve this problem, that is, using memory networks to store prototypes of normal events, thereby enhancing the model's ability to represent normal events and limiting its ability to represent abnormal events. Summary of the invention

[0004] The present application provides a method and device for traffic video anomaly detection based on memory enhancement and spatiotemporal association, aiming to solve the problem of how to accurately model spatiotemporal association in traffic video anomaly detection.

[0005] The first aspect of the present application provides a method for traffic video anomaly detection based on memory enhancement and spatiotemporal association, comprising the following steps: acquiring traffic monitoring video stream data, processing the traffic monitoring video stream data into continuous image frame data and corresponding frame difference data, and dividing the image frame data into a training set and a test set; constructing an anomaly detection model based on memory enhancement and spatiotemporal association, and training the anomaly detection model based on the frame difference data and the training set, wherein the anomaly detection model records temporal features and spatial features through memory enhancement, and the anomaly detection model establishes the correlation between temporal features and spatial features; and performing anomaly detection on the test set using the trained anomaly detection model.

[0006] Optionally, the image frame data is divided into a training set and a test set, including: classifying the image frame data according to whether an abnormal event / behavior occurs; dividing continuous image frames in which no abnormal event / behavior occurs into the training set, and dividing continuous image frames containing abnormal events / behaviors into the test set.

[0007] Optionally, the anomaly detection model includes a temporal feature network branch, a spatial feature network branch, a spatiotemporal association module and a decoder, wherein the memory-enhanced temporal features in the temporal feature network branch are input into the spatiotemporal association module; the memory-enhanced spatial features in the spatial feature network branch are input into the spatiotemporal association module; the spatiotemporal association module fuses the temporal features and the spatial features to obtain the spatiotemporal fusion features; and the decoder decodes the spatiotemporal fusion features to obtain the future frame image.

[0008] Optionally, the temporal feature network branch includes an appearance encoder, an appearance memory pool, and a cross-attention module, and the spatial feature network branch includes a motion encoder, a motion memory pool, and a cross-attention module, wherein the appearance encoder extracts spatial features in the training set, the appearance memory pool performs memory enhancement on the spatial features, and the cross-attention module of the temporal feature network branch fuses the original spatial features and the memory-enhanced spatial features; the appearance encoder extracts temporal features in inter-frame difference data, the motion memory pool performs memory enhancement on the temporal features, and the cross-attention module of the spatial feature network branch fuses the original temporal features and the memory-enhanced temporal features.

[0009] Optionally, the appearance encoder has the same structure as the motion encoder, wherein the structures of the appearance encoder and the motion encoder include: an initial convolution block and a downsampling block, the initial convolution block consists of two convolution layers, each convolution layer includes a convolution operation, batch normalization and a Leaky ReLU activation function; the downsampling block consists of a maximum pooling layer and a convolution block, the convolution block includes two convolution layers, each convolution layer applies convolution, batch normalization and a Leaky ReLU activation function in sequence.

[0010] Optionally, the appearance memory pool and the motion memory pool have the same structure, wherein the appearance memory pool and the motion memory pool are two-dimensional matrices, each row of the two-dimensional matrix is ​​a memory item, the cosine similarity of the input feature and the memory item is calculated, the weight corresponding to the memory item is determined according to the cosine similarity, and the time feature after memory enhancement is determined according to the memory item and the corresponding weight.

[0011] Optionally, the structures of the cross-attention modules of the temporal feature network branch and the spatial feature network branch are the same, wherein the execution process of the cross-attention module includes: mapping the original features to query vectors through linear operations; mapping the memory-enhanced features to key vectors and value vectors respectively through linear operations: calculating the similarity between the query vector and the key vector, and calculating the attention weight according to the similarity; multiplying the attention weight by the value vector to obtain the target feature, aggregating the value in the target feature to the corresponding query position, and mapping the aggregated features at the query position back to the target feature; fusing the target feature, the original feature and the memory-enhanced feature.

[0012] Optionally, the execution process of the spatiotemporal association module includes: concatenating the input spatial features and temporal features in the channel dimension, obtaining fused features using a depthwise separable convolution operation, and sending the fused features to two identical parallel branches, wherein each branch is composed of a Sigmoid activation function and average pooling to obtain spatial weights and temporal weights; multiplying the spatial weights and temporal weights with the original input respectively to obtain weighted gated spatial features and gated temporal features, and using the gated temporal features to weight the gated spatial features to enhance the elements in the spatial features to obtain the gated spatiotemporal features; using the channel attention mechanism to enhance the elements in the gated spatiotemporal features, and adding them to the gated spatial features to obtain the final spatiotemporal fusion features.

[0013] Optionally, anomaly detection is performed on the test set using a trained anomaly detection model, including: inputting the test set into the trained anomaly detection model, the anomaly detection model outputting a future frame image, calculating the peak signal-to-noise ratio between the future frame image and the real frame image; calculating the L2 distance between the query vector and the memory item, calculating the anomaly score according to the peak signal-to-noise ratio and the L2 distance, and setting a score threshold; if the anomaly score is greater than the score threshold, the test set is determined to be abnormal.

[0014] The second aspect of the present application provides a traffic video anomaly detection device based on memory enhancement and spatiotemporal association, including: a data processing module, used to obtain traffic monitoring video stream data, process the traffic monitoring video stream data into continuous image frame data and corresponding frame difference data, and divide the image frame data into a training set and a test set; a model training module, used to construct an anomaly detection model based on memory enhancement and spatiotemporal association, and train the anomaly detection model based on the frame difference data and the training set, wherein the anomaly detection model records the time features and spatial features through memory enhancement, and the anomaly detection model establishes the correlation between the time features and the spatial features; a detection module, used to perform anomaly detection on the test set using the trained anomaly detection model.

[0015] Therefore, the present application has at least one or more of the following beneficial effects:

[0016] The present invention uses appearance and motion memory pools to respectively store prototype features of the appearance and motion of normal events, thereby increasing the distance between normal and abnormal events in the feature space; proposes a cross-attention fusion module that can capture the fine-grained relationship between query features and features recorded in the memory pool, effectively promoting feature fusion; proposes a spatiotemporal association module that more comprehensively models the intrinsic correlation of spatiotemporal information, promotes the model to learn and understand the complex spatiotemporal interactive behaviors of normal events in the video, thereby improving the accuracy of anomaly detection; traffic video anomaly detection is performed based on the method of memory enhancement and spatiotemporal association modeling, which significantly improves the anomaly detection performance of the model in complex traffic scenes, and has higher robustness and generalization ability.

[0017] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0019] Figure 1 It is a flow chart of a traffic video anomaly detection method based on memory enhancement and spatiotemporal association according to an embodiment of the present application;

[0020] Figure 2 is a structural diagram of an anomaly detection model according to an embodiment of the present application;

[0021] Figure 3 is a flow chart of a read operation of a memory pool according to an embodiment of the present application;

[0022] Figure 4 is an execution flow chart of a cross attention fusion module according to an embodiment of the present application;

[0023] Figure 5 is an execution flow chart of the spatiotemporal association module according to an embodiment of the present application;

[0024] Figure 6 A schematic block diagram of a traffic video anomaly detection device based on memory enhancement and spatiotemporal association according to an embodiment of the present application DETAILED DESCRIPTION

[0025] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0026] Abnormal events in traffic scenes often involve the joint changes of spatial and temporal features, such as abnormal trajectories of vehicles on the road or sudden acceleration and deceleration. Therefore, how to accurately model spatiotemporal correlation is the key to traffic video anomaly detection.

[0027] Although related technologies have made certain progress in detection performance, there are still obvious deficiencies in spatiotemporal correlation modeling: ① Reconstruction-based methods usually only focus on the spatial features of a single frame and ignore temporal information, resulting in reduced detection performance; ② Prediction-based methods generally use a two-stream network to learn appearance features and motion features respectively, but lack effective exploration of the intrinsic relationship between spatiotemporal information, resulting in limited model performance; ③ Memory network-based methods usually concatenate the features aggregated in the memory pool and the original input features and directly use the decoder to decode them, but ignore the difference between the two types of feature information. This difference may introduce unexpected semantic deviations, thereby affecting the detector performance.

[0028] To this end, the main technical problem solved by the embodiments of the present application is: the current traffic video anomaly detection method fails to fully utilize the intrinsic connection of spatiotemporal information, and the feature fusion mechanism of the memory network is prone to introduce semantic deviations. A traffic video anomaly detection method and device based on memory enhancement and spatiotemporal association is designed to enhance the model's ability to understand complex spatiotemporal patterns in traffic videos and improve detection accuracy. By introducing a cross-attention fusion module, the input features are effectively fused with the features retrieved from the memory pool. By designing a spatiotemporal association module, the spatiotemporal patterns of normal events are fully learned and their intrinsic connections are modeled, thereby improving the performance of video anomaly detection.

[0029] Specifically, Figure 1 A flow chart of a traffic video anomaly detection method based on memory enhancement and spatiotemporal association provided in an embodiment of the present application.

[0030] like Figure 1 As shown, the traffic video anomaly detection method based on memory enhancement and spatiotemporal association includes the following steps:

[0031] In step S101, traffic monitoring video stream data is acquired, the traffic monitoring video stream data is processed into continuous image frame data and corresponding inter-frame difference data, and the image frame data is divided into a training set and a test set.

[0032] It can be understood that the embodiments of the present application can process traffic monitoring video data into continuous RGB image frame data, and generate RGB frame difference data by calculating the pixel difference (channel-by-channel absolute value difference) between adjacent frames, and then divide the data into training sets and test sets.

[0033] In some embodiments, the image frame data is divided into a training set and a test set, including: classifying the image frame data according to whether an abnormal event / behavior has occurred; dividing continuous image frames in which no abnormal event / behavior has occurred into the training set, and dividing continuous image frames containing abnormal events / behaviors into the test set.

[0034] In step S102, an anomaly detection model is constructed based on memory enhancement and spatiotemporal association, and the anomaly detection model is trained based on frame difference data and a training set, wherein the anomaly detection model records temporal features and spatial features through memory enhancement, and the anomaly detection model establishes the association between temporal features and spatial features.

[0035] It can be understood that the embodiment of the present application effectively fuses the input features with the features retrieved from the memory pool by introducing a cross-attention fusion module, and fully learns the spatiotemporal patterns of normal events and models their internal connections by designing a spatiotemporal association module, thereby improving the performance of video anomaly detection.

[0036] In some embodiments, Figure 2 As shown in the figure, the anomaly detection model includes a temporal feature network branch, a spatial feature network branch, a spatiotemporal association module and a decoder, wherein the memory-enhanced temporal features in the temporal feature network branch are input into the spatiotemporal association module; the memory-enhanced spatial features in the spatial feature network branch are input into the spatiotemporal association module; the spatiotemporal association module fuses the temporal features and spatial features to obtain the spatiotemporal fusion features; the decoder decodes the spatiotemporal fusion features to obtain the future frame image. The temporal feature network branch includes an appearance encoder, an appearance memory pool, and a cross-attention module, and the spatial feature network branch includes a motion encoder, a motion memory pool, and a cross-attention module.

[0037] In some embodiments, the appearance encoder extracts spatial features in the training set, the appearance memory pool performs memory enhancement on the spatial features, and the cross-attention module of the temporal feature network branch fuses the original spatial features and the memory-enhanced spatial features; the appearance encoder extracts temporal features in the inter-frame difference data, the motion memory pool performs memory enhancement on the temporal features, and the cross-attention module of the spatial feature network branch fuses the original temporal features and the memory-enhanced temporal features.

[0038] It can be understood that the cross-attention fusion module can promote the effective fusion between the query features and the features recorded in the memory pool, and improve the representation ability of the features; the spatiotemporal association module can effectively model the inherent consistency of spatiotemporal information and promote the model to understand the spatiotemporal interactions in complex scenarios.

[0039] In some embodiments, the structure of the appearance encoder is the same as that of the motion encoder, wherein the structures of the appearance encoder and the motion encoder include: an initial convolution block and a downsampling block, the initial convolution block consists of two convolution layers, each convolution layer includes a convolution operation, batch normalization and a Leaky ReLU activation function; the downsampling block consists of a maximum pooling layer and a convolution block, the convolution block includes two convolution layers, each convolution layer applies convolution, batch normalization and a Leaky ReLU activation function in sequence.

[0040] Specifically, if Figure 2 As shown in the figure, the input of the appearance encoder is continuous RGB image frames, and the input of the motion encoder is the corresponding RGB frame difference. The two learn the spatial features of normal events respectively. s and time feature F t The appearance encoder has the same structure as the motion encoder, consisting of an initial convolutional block and three downsampling blocks. The initial convolutional block consists of two convolutional layers, each of which contains a convolution operation, batch normalization, and a leaky ReLU activation function. The downsampling block consists of a max pooling layer and a convolutional block, which also contains two convolutional layers, each of which applies convolution, batch normalization, and a leaky ReLU activation function in sequence.

[0041] In some embodiments, the appearance memory pool and the motion memory pool have the same structure, wherein the appearance memory pool and the motion memory pool are two-dimensional matrices, each row of the two-dimensional matrix is ​​a memory item, the cosine similarity of the input feature and the memory item is calculated, the weight corresponding to the memory item is determined according to the cosine similarity, and the time feature after memory enhancement is determined according to the memory item and the corresponding weight.

[0042] Specifically, if Figure 3 As shown, in order to record the appearance and motion prototype patterns of normal events, the spatial and temporal features are input into the appearance memory pool and the motion memory pool respectively to obtain the memory-enhanced spatial features. and the temporal characteristics of memory enhancement The memory pool is a two-dimensional matrix M∈R N×C , each row of the matrix is ​​a memory item m j ∈R C ,j=1,2,…,N. First, calculate the cosine similarity between the input feature Q and the memory item:

[0043]

[0044] Then use the Softmax function to obtain the corresponding weight w j :

[0045]

[0046] Finally, using the weight w j Multiply it with the memory item in the memory pool to obtain the memory-enhanced feature:

[0047]

[0048] In some embodiments, the structures of the cross-attention modules of the temporal feature network branch and the spatial feature network branch are the same, wherein the execution process of the cross-attention module includes: mapping the original features to query vectors through linear operations; mapping the memory-enhanced features to key vectors and value vectors respectively through linear operations: calculating the similarity between the query vector and the key vector, and calculating the attention weight based on the similarity; multiplying the attention weight by the value vector to obtain the target feature, aggregating the value in the target feature to the corresponding query position, and mapping the aggregated features at the query position back to the target feature; fusing the target feature, the original feature, and the memory-enhanced feature.

[0049] Specifically, if Figure 4 As shown in Fig. 1, in order to avoid the direct concatenation of features introducing additional semantic deviations and reducing the fusion performance, the present invention proposes a cross-attention fusion module, which deeply fuses the original input features and the memory-enhanced features through a cross-attention mechanism to obtain the fused feature F′. s . With spatial characteristics F s and its memory-enhancing characteristics Take this as an example, the specific execution process is:

[0050] (1) First, the input spatial feature F s ∈R C×H×W Mapped into query vector Q∈R through linear operation N×C , N = H × W, and the memory-enhanced spatial features Through linear operations, they are mapped into key vectors K∈R C×N The sum vector V∈R C ×N :

[0051]

[0052] Where W q ,W k ,W v are the coefficients that the model needs to learn.

[0053] (2) Then the scaled dot product is used to calculate the similarity between Q and K, and then normalized by the Softmax function to obtain the attention weight A∈R N×N :

[0054] A=Softmax(Q T K)

[0055] (3) Then, the attention weight is multiplied by the value vector, the values ​​with higher correlation are aggregated to the corresponding query positions, and the obtained features are then mapped back to F o ∈R C×H×W :

[0056] F o =AV T

[0057] (4) Finally, the weighted output features are added to the original input features and fused through a 1×1 convolution to obtain the final output features:

[0058]

[0059] The memory pool features and original query features are fused through the cross-attention mechanism, which enables the model to discover deeper association patterns and improve the overall expressiveness of the features.

[0060] In some embodiments, the execution process of the spatiotemporal association module includes: splicing the input spatial features and temporal features in the channel dimension, obtaining the fused features using a deep separable convolution operation, and sending the fused features to two identical parallel branches, wherein each branch is composed of a Sigmoid activation function and average pooling to obtain spatial weights and temporal weights; multiplying the spatial weights and temporal weights with the original input respectively to obtain weighted gated spatial features and gated temporal features, and using the gated temporal features to weight the gated spatial features to enhance the elements in the spatial features to obtain the gated spatiotemporal features; using a channel attention mechanism to enhance the elements in the gated spatiotemporal features, and adding them to the gated spatial features to obtain the final spatiotemporal fusion features.

[0061] Specifically, if Figure 5 As shown in the figure, in order to encourage the model to learn the inherent spatiotemporal correlation of normal events and accurately understand the complex spatiotemporal interaction behaviors in traffic videos, thereby improving the anomaly detection performance, the present invention proposes a spatiotemporal correlation module. This module semantically aligns deep spatial features and temporal features to capture their inherent spatiotemporal consistency. The specific execution process is as follows:

[0062] (1) First, the input spatial feature F′ s and time feature F′ t The concatenation is performed on the channel dimension, and the depth-wise convolution operation is used to obtain the fused features, and then the fused features are sent to two identical parallel branches, each of which is composed of a Sigmoid activation function and average pooling to obtain the spatial weight g.s and time weight g t :

[0063] g=AvgPool(Sigmoid(DWCConv(Concat(F′ s ,F′ t ))))

[0064] (2) Then the spatial weight and temporal weight are multiplied with the original input to obtain the weighted gated spatial feature G′ s And the gated temporal feature G′ t :

[0065]

[0066] (3) Then use the gated temporal features to weight the gated spatial features to enhance the significant elements in the spatial features and obtain the gated spatiotemporal features G′ t :

[0067]

[0068] (4) Finally, the channel attention mechanism is used to enhance the spatiotemporal features of the gate G′ t The important elements in the fusion function are added to the gated spatial feature G′ to obtain the final spatiotemporal fusion feature F st ∈R C×H×w :

[0069]

[0070] In some embodiments, the decoder is responsible for combining the spatiotemporal fusion features F st Decode and predict future frame RGB images Specifically, it consists of three upsampling blocks and a final convolution block. Each upsampling block is responsible for gradually restoring the size of the feature to the size of the input image and reducing the channel dimension of the feature. The final convolution block is responsible for mapping the channel dimension of the feature to the number of channels of the RGB image. Each upsampling block consists of an upsampling layer and a convolution block. The convolution block contains two convolution layers, each of which applies convolution, batch normalization, and Leaky ReLU activation functions in sequence. The final convolution block consists of a convolution layer that maps the number of channels to the number of channels of the corresponding RGB image, and applies the Tanh activation function to limit the output value to the range of [-1,1].

[0071] Furthermore, the k consecutive RGB images I in the training set are t-k:t and k-1 frames of RGB frame difference image Y t-k:t-1 Input into the model, the model predicts the future frame RGB image In order to better optimize the model, we combine the prediction loss Feature compactness loss of memory pool and feature diversity loss The following loss function is constructed, where λ 1 and λ 2 is a configurable hyperparameter:

[0072]

[0073] (1) is calculated by calculating the real frame I t+1 and future frame images The L2 distance between them is:

[0074]

[0075] (2) In order to make the query vector and the most similar memory item As similar as possible, we calculate the L2 distance between the two as the feature compactness loss

[0076]

[0077] Where N is the number of query vectors.

[0078] (3) To ensure that the memory items in the memory pool can store diverse normal patterns, we propose a feature diversity loss To ensure that there is a certain degree of difference between the memory items:

[0079]

[0080] in and Represent the most similar and second most similar memory items to the query vector respectively.

[0081] In step S103, anomaly detection is performed on the test set using the trained anomaly detection model.

[0082] It is understandable that the embodiments of the present application can input test data into a trained model and determine whether the test frame is abnormal based on the anomaly score output by the model.

[0083] In some embodiments, anomaly detection is performed on a test set using a trained anomaly detection model, including: inputting the test set into the trained anomaly detection model, the anomaly detection model outputting a future frame image, calculating the peak signal-to-noise ratio between the future frame image and the real frame image; calculating the L2 distance between the query vector and the memory item, calculating the anomaly score based on the peak signal-to-noise ratio and the L2 distance, and setting a score threshold; if the anomaly score is greater than the score threshold, the test set is determined to be abnormal.

[0084] The score threshold may be specifically set to represent a critical value of an event / behavior in a traffic video.

[0085] Specifically, the k consecutive RGB images I in the test set are t-k:t and k-1 frames of RGB frame difference image Y t-k:t-1 Input into the model, the model predicts the future frame RGB image The anomaly score of the predicted frame is calculated jointly from the image domain and feature space. The specific calculation process is as follows:

[0086] (1) Use peak signal-to-noise ratio (PSNR) to measure the true frame I t+1 and future frame images The difference between is used to measure the degree of abnormality in the image domain:

[0087]

[0088] Where N is the total number of pixels in the video image, is the maximum pixel value.

[0089] (2) Since the memory pool only stores normal patterns, when the query vector of an abnormal event is retrieved from the memory pool, there will be a large distance between it and the memory item. Therefore, the query vector and the most similar memory item are calculated. The L2 distance between them is used to measure the degree of abnormality in the feature space:

[0090]

[0091] in represents the memory item that is most similar to the query vector, and N is the number of query vectors.

[0092] (3) Finally, the anomaly score can be obtained as shown below:

[0093]

[0094] where τ represents an adjustable hyperparameter, Represents the maximum and minimum normalization function. A score threshold can be set. If the anomaly score of the predicted frame is greater than the threshold, it can be judged as abnormal, otherwise it is judged as normal.

[0095] According to the traffic video anomaly detection method based on memory enhancement and spatiotemporal association proposed in the embodiment of the present application, the prototype features of the appearance and motion of normal events are respectively saved through the appearance and motion memory pools, thereby increasing the distance between normal and abnormal events in the feature space; a cross-attention fusion module is proposed, which can capture the fine-grained relationship between the query features and the features recorded in the memory pool, and effectively promote the fusion of features; a spatiotemporal association module is proposed, which more comprehensively models the intrinsic correlation of spatiotemporal information, promotes the model to learn and understand the complex spatiotemporal interactive behaviors of normal events in the video, thereby improving the accuracy of anomaly detection; traffic video anomaly detection is performed based on the method of memory enhancement and spatiotemporal association modeling, which significantly improves the anomaly detection performance of the model in complex traffic scenes, and has higher robustness and generalization ability.

[0096] Next, a traffic video anomaly detection device based on memory enhancement and spatiotemporal association proposed in accordance with an embodiment of the present application is described with reference to the accompanying drawings.

[0097] Figure 6 It is a block diagram of a traffic video anomaly detection device based on memory enhancement and spatiotemporal association according to an embodiment of the present application.

[0098] like Figure 6 As shown, the traffic video anomaly detection device 10 based on memory enhancement and spatiotemporal association includes: a data processing module 100, a model training module 200 and a detection module 300.

[0099] Among them, the data processing module 100 is used to obtain traffic monitoring video stream data, process the traffic monitoring video stream data into continuous image frame data and corresponding frame difference data, and divide the image frame data into a training set and a test set; the model training module 200 is used to construct an anomaly detection model based on memory enhancement and spatiotemporal association, and train the anomaly detection model based on the frame difference data and the training set, wherein the anomaly detection model records the temporal features and spatial features through memory enhancement, and the anomaly detection model establishes the correlation between the temporal features and the spatial features; the detection module 300 is used to perform anomaly detection on the test set using the trained anomaly detection model.

[0100] It should be noted that the aforementioned explanation of the embodiment of the traffic video anomaly detection method based on memory enhancement and spatiotemporal association is also applicable to the traffic video anomaly detection device based on memory enhancement and spatiotemporal association of this embodiment, and will not be repeated here.

[0101] According to the traffic video anomaly detection device based on memory enhancement and spatiotemporal association proposed in the embodiment of the present application, the prototype features of the appearance and motion of normal events are respectively saved through the appearance and motion memory pools, thereby increasing the distance between normal and abnormal events in the feature space; a cross-attention fusion module is proposed, which can capture the fine-grained relationship between the query features and the features recorded in the memory pool, and effectively promote the fusion of features; a spatiotemporal association module is proposed, which more comprehensively models the intrinsic correlation of spatiotemporal information, promotes the model to learn and understand the complex spatiotemporal interactive behaviors of normal events in the video, thereby improving the accuracy of anomaly detection; traffic video anomaly detection is performed based on the method of memory enhancement and spatiotemporal association modeling, which significantly improves the anomaly detection performance of the model in complex traffic scenes, and has higher robustness and generalization ability.

[0102] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0103] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0104] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

[0105] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array, a field programmable gate array, etc.

[0106] A person of ordinary skill in the art may understand that all or part of the steps of the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one of the steps of the method embodiment or a combination thereof.

Claims

1. A traffic video anomaly detection method based on memory enhancement and spatiotemporal association, characterized in that: The following steps are involved: Acquire traffic monitoring video stream data, process the traffic monitoring video stream data into continuous image frame data and corresponding inter-frame difference data, and divide the image frame data into a training set and a test set; An anomaly detection model is constructed based on memory enhancement and spatiotemporal association, and the anomaly detection model is trained based on the inter-frame difference data and the training set, wherein the anomaly detection model records the temporal features and the spatial features through memory enhancement, and the anomaly detection model establishes the association between the temporal features and the spatial features; The trained anomaly detection model is used to perform anomaly detection on the test set.

2. The traffic video anomaly detection method based on memory enhancement and spatiotemporal association according to claim 1 is characterized in that: The step of dividing the image frame data into a training set and a test set comprises: Classifying the image frame data according to whether an abnormal event / behavior occurs; The continuous image frames without abnormal events / behaviors are divided into the training set, and the continuous image frames containing abnormal events / behaviors are divided into the test set.

3. The traffic video anomaly detection method based on memory enhancement and spatiotemporal association according to claim 1 is characterized in that: The anomaly detection model includes a temporal feature network branch, a spatial feature network branch, a spatiotemporal association module and a decoder, wherein: The memory-enhanced temporal features in the temporal feature network branch are input into the spatiotemporal association module; The memory-enhanced spatial features in the spatial feature network branch are input into the spatiotemporal association module; The spatiotemporal association module fuses the temporal features and spatial features to obtain spatiotemporal fusion features; The decoder decodes the spatiotemporal fusion features to obtain a future frame image.

4. The traffic video anomaly detection method based on memory enhancement and spatiotemporal association according to claim 3 is characterized in that: The temporal feature network branch includes an appearance encoder, an appearance memory pool, and a cross attention module, and the spatial feature network branch includes a motion encoder, a motion memory pool, and a cross attention module, wherein: The appearance encoder extracts the spatial features in the training set, the appearance memory pool performs memory enhancement on the spatial features, and the cross attention module of the temporal feature network branch fuses the original spatial features and the memory-enhanced spatial features; The appearance encoder extracts the temporal features in the inter-frame difference data, the motion memory pool performs memory enhancement on the temporal features, and the cross-attention module of the spatial feature network branch fuses the original temporal features with the memory-enhanced temporal features.

5. The traffic video anomaly detection method based on memory enhancement and spatiotemporal association according to claim 4 is characterized in that: The appearance encoder has the same structure as the motion encoder, wherein the structures of the appearance encoder and the motion encoder include: an initial convolution block and a downsampling block, The initial convolution block consists of two convolutional layers, each of which includes a convolution operation, batch normalization, and a LeakyReLU activation function; The downsampling block consists of a maximum pooling layer and a convolution block, the convolution block contains two convolution layers, and each convolution layer applies convolution, batch normalization and Leaky ReLU activation functions in sequence.

6. The traffic video anomaly detection method based on memory enhancement and spatiotemporal association according to claim 4 is characterized in that: The appearance memory pool and the motion memory pool have the same structure, wherein the appearance memory pool and the motion memory pool are two-dimensional matrices, each row of the two-dimensional matrix is ​​a memory item, the cosine similarity of the input feature and the memory item is calculated, the weight corresponding to the memory item is determined according to the cosine similarity, and the time feature after memory enhancement is determined according to the memory item and the corresponding weight.

7. The traffic video anomaly detection method based on memory enhancement and spatiotemporal association according to claim 4 is characterized in that: The structures of the cross attention modules of the temporal feature network branch and the spatial feature network branch are the same, wherein the execution process of the cross attention module includes: Map the original features into query vectors through linear operations; The memory-enhanced features are mapped into key vectors and value vectors through linear operations: Calculate the similarity between the query vector and the key vector, and calculate the attention weight according to the similarity; Multiplying the attention weight by the value vector to obtain a target feature, aggregating the value in the target feature to a corresponding query position, and mapping the aggregated feature at the query position back to the target feature; The target feature, the original feature and the memory-enhanced feature are fused.

8. The traffic video anomaly detection method based on memory enhancement and spatiotemporal association according to claim 4 is characterized in that: The execution process of the spatiotemporal association module includes: The input spatial features and temporal features are concatenated in the channel dimension, and the fused features are obtained by using the depthwise separable convolution operation. The fused features are sent to two identical parallel branches, where each branch is composed of a Sigmoid activation function and average pooling to obtain spatial weights and temporal weights. The spatial weight and the temporal weight are multiplied by the original input to obtain the weighted gated spatial features and the gated temporal features. The gated temporal features are used to weight the gated spatial features to enhance the elements in the spatial features and obtain the gated spatiotemporal features. The channel attention mechanism is used to enhance the elements in the gated spatiotemporal features and add them to the gated spatial features to obtain the final spatiotemporal fusion features.

9. The traffic video anomaly detection method based on memory enhancement and spatiotemporal association according to claim 4 is characterized in that: The performing anomaly detection on the test set using the trained anomaly detection model includes: Inputting the test set into the trained anomaly detection model, the anomaly detection model outputs a future frame image, and calculating a peak signal-to-noise ratio between the future frame image and the real frame image; Calculate the L2 distance between the query vector and the memory item, calculate the anomaly score according to the peak signal-to-noise ratio and the L2 distance, and set the score threshold; If the anomaly score is greater than the score threshold, the test set is determined to be abnormal.

10. A traffic video anomaly detection device based on memory enhancement and spatiotemporal association, characterized in that: include: A data processing module, used for acquiring traffic monitoring video stream data, processing the traffic monitoring video stream data into continuous image frame data and corresponding inter-frame difference data, and dividing the image frame data into a training set and a test set; A model training module, used for constructing an anomaly detection model based on memory enhancement and spatiotemporal association, and training the anomaly detection model based on the inter-frame difference data and the training set, wherein the anomaly detection model records the temporal features and the spatial features through memory enhancement, and the anomaly detection model establishes the association between the temporal features and the spatial features; The detection module is used to perform anomaly detection on the test set using the trained anomaly detection model.

Citation Information

Cited By

  • Traffic behavior anomaly detection method and system, computer equipment and medium

    CN122116013A