A video surveillance anomaly detection method based on memory neural network

By introducing memory neural networks into video surveillance abnormal detection, the shortcomings of the prior art in identifying abnormal events in complex video data are solved, and more efficient and accurate abnormal detection effects are achieved.

CN116310960BActive Publication Date: 2025-05-06TIBET HONGZHI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310155405.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-23
Publication Date
2025-05-06
Estimated Expiration
2043-02-23

AI Technical Summary

Technical Problem

When existing video surveillance abnormal detection technology processes complex real video data, it is difficult to effectively identify abnormal events, especially abnormal data points near the boundaries of normal areas, and insufficient detection accuracy due to limited label training data.

Method used

A memory neural network is introduced to hash the normal video stream information through unsupervised learning and store it in the memory storage unit. During testing, the video encoding feature information required for reconstruction is retrieved from the memory unit to enhance the abnormal detection ability.

Benefits of technology

The memory neural network enhances the deep autoencoder, which improves the performance and accuracy of video surveillance abnormal detection, and can more effectively identify abnormal events in video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310960B_ABST
    Figure CN116310960B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of video anomaly monitoring, and specifically relates to a video monitoring anomaly detection method based on a memory neural network, comprising: obtaining a video frame segment to be detected, extracting a feature map for the video frame segment to be detected; performing matrix transformation on the feature map of the frame segment to obtain query information of the video frame segment; equally dividing the query information and performing matrix transformation on each of them to obtain a pre-search address and weight information; fusing two groups of address information to obtain a final search address and weight; retrieving video frame feature information stored in a storage unit from a memory module according to the address, and weightedly splicing the video frame feature information into a decoding feature map; inputting the decoding feature map into a decoder, outputting a reconstructed decoded video frame, and obtaining a video monitoring anomaly detection result; the present invention enhances a deep autoencoder by a memory neural network, so that the deep autoencoder can learn the feature information of a normal video frame segment, and improves the detection performance of abnormal video monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of video anomaly monitoring and pattern recognition, and in particular relates to a video monitoring anomaly detection method based on a memory neural network. Background Art

[0002] Video surveillance uses computer vision technology to analyze and understand long video streams, and plays an irreplaceable role in public safety. Abnormal event detection, as an important part of intelligent video surveillance, automatically discovers and identifies anomalies in the ever-changing scenes under surveillance, and takes timely measures to deal with emergencies. Anomaly detection in video refers more to identifying events that deviate from expected behavior, which is an important task in video analysis and plays a vital role in video surveillance. However, video anomaly detection is a very challenging task for the following reasons: First, real video data is complex, and some abnormal data points may be close to the boundary of normal areas. For example, skateboarders and walking people are similar in appearance, but scooters are abnormal objects that are prohibited from driving on sidewalks. Second, the labeled training data for anomaly detection is limited. Although normal patterns are usually relatively easy to collect, abnormal samples are rare and expensive to obtain.

[0003] Deep autoencoders have been widely used for anomaly detection in video surveillance. When the autoencoder is trained only with normal video stream data, the normal video frame segments can be reconstructed and restored, but the abnormal video frame segments cannot be restored, resulting in large errors and being identified as abnormal video frame segments. However, this assumption does not always hold true in practice. In practical applications, sometimes the generalization of the deep autoencoder is so strong that it can also reconstruct abnormal video frames well, resulting in missed detection of anomalies. To solve the above problems, we introduce memory neural networks and propose a video surveillance anomaly detection method using memory neural networks. Through unsupervised learning, the normal video stream information is hashed and widely stored in the memory storage unit. During testing, the video coding feature information required for reconstruction is retrieved from the memory unit. The reconstruction error between the input of the abnormal video frame and the output that tends to be normal video frames will be amplified, thereby enhancing the anomaly detection capability. Summary of the invention

[0004] In order to solve the problems existing in the above prior art, the present invention proposes a video surveillance anomaly detection method based on a memory neural network, which method includes: obtaining video stream data, and preprocessing the video stream data to obtain video numerical information; obtaining a video frame segment to be detected from the video numerical information, and inputting the video frame segment to be detected into an encoder to extract a feature map of the video frame segment; performing a matrix transformation on the feature map of the frame segment to obtain query information of the video frame segment; dividing the query information equally into two parts, and performing a matrix transformation on the divided query information respectively to obtain the pre-search address and weight information of two groups of memory modules; fusing the two groups of address information to obtain the final search address and weight; retrieving the video frame feature information stored in the storage unit from the memory module according to the address, and weighting and splicing the video frame feature information into a decoding feature map; inputting the decoding feature map into a decoder, and outputting a reconstructed decoded video frame; subtracting the reconstructed decoded video frame from the input video frame information, judging whether it is abnormal according to the value of the subtraction, and obtaining a video surveillance anomaly detection result.

[0005] Preferably, the process of preprocessing the video stream data includes: obtaining a video stream from a video surveillance terminal; extracting jpg format images of all frames from the video stream, using the numpy tool library to save all frame images as 5D numerical matrix information, and from the numerical matrix information, intercepting the information of every 12 images according to the first dimension into a frame segment, and collecting all the intercepted frame segment information.

[0006] Preferably, the process of performing matrix transformation on the frequency frame segment feature map includes: setting a query information generator The query information generator is initialized; the encoded feature map of the video frame segment is input into the query information generator for matrix transformation to obtain the query information

[0007] Preferably, the process of performing matrix transformation on the split query information includes: performing matrix transformation on the two sub-query information through two index generators respectively. The matrix is ​​converted into the front retrieval address ind of the two memory modules. 1 and ind 2 and weight information

[0008] Preferably, the process of fusing the two sets of address information includes: fusing the two sets of address information using a cross-sum method.

[0009] Preferably, the process of retrieving the video frame feature information stored in the storage unit from the memory module according to the address includes: setting the memory module And initialize the memory module; each row in the memory module is set with a number, and the stored video feature information is retrieved from the memory module M according to the fused index number.

[0010] Preferably, the video frame feature information is weighted and spliced ​​into a decoding feature map: the retrieved video feature information is matched one-to-one with the obtained weight value and multiplied; the multiplied data is spliced ​​into a matrix of a specified dimension to obtain a decoding feature map.

[0011] Preferably, the process of inputting the decoded feature map into the decoder for decoding includes: initializing the video frame segment decoder module, inputting the decoded feature map into the decoder, and obtaining the output result through forward propagation calculation of the neural network.

[0012] Beneficial effects of the present invention:

[0013] The present invention enhances the deep autoencoder through the memory neural network, so that the deep autoencoder can learn the characteristic information of normal video frame segments and improve the detection performance of abnormal video monitoring. The present invention can integrate the memory network module in the deep autoencoder and apply it to the abnormal detection task of monitoring video, which can better help the network learn the normal mode and enhance the detection performance of video abnormalities. At the same time, compared with other memory module methods, our method has higher detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a schematic diagram of the network structure of the present invention. DETAILED DESCRIPTION

[0015] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0016] A video surveillance anomaly detection method based on a memory neural network, the method comprising acquiring video stream data, preprocessing the video stream to obtain video frame information; encoding the video frame to obtain a feature map of the video frame segment; retrieving the video frame encoding information according to the feature map of the video frame segment, and encoding the retrieved video frame encoding information; using a decoder to decode the encoded information, and determining whether the decoded video frame is abnormal.

[0017] A specific implementation of a video surveillance anomaly detection method based on a memory neural network includes: obtaining video stream data; video stream preprocessing: using a Python tool to make the video stream into numerical information and save it in a computer; encoding video frame information: obtaining a video frame segment to be detected from the saved video numerical information, and inputting the video frame segment to be detected into an encoder to extract a feature map of the video frame segment; retrieving video frame encoding information: performing matrix transformation on the feature map of the frame segment to obtain query information of the video frame segment; equally dividing the query information into two parts, performing matrix transformation on the divided query information respectively, and obtaining the pre-search address and weight information of two groups of memory modules; fusing the two groups of address information to obtain the final search address and weight; retrieving the video frame feature information stored in the storage unit from the memory module according to the address, and weightedly splicing the video frame feature information into a decoding feature map; decoding the video frame information: inputting the decoding feature map into a decoder, and outputting a reconstructed decoded video frame; identifying the anomaly of the video frame segment: subtracting the reconstructed decoded video frame from the input video frame information, judging whether it is abnormal according to the value of the subtraction, and obtaining a video surveillance anomaly detection result.

[0018] A specific implementation method of a video surveillance anomaly detection method introducing a memory neural network, such as Figure 1 As shown, the method comprises the following steps:

[0019] Step 1: Design a video frame segment encoder and initialize it;

[0020] Step 2: Select samples from the set of video frames to be detected and input them into the encoder according to the input format, and encode them into feature maps containing video feature information. The conversion formula is:

[0021] Z eout =f e (X;θ e ) (1)

[0022] Among them, f e is the designed video frame segment encoder structure, θ e is the encoder parameter, X is the video frame segment input, Z eout It is a feature map containing video feature information.

[0023] Step 3: Perform a matrix transformation on the feature map of the video frame segment to obtain the query information of the video frame segment, divide the query information into two equal parts, perform matrix transformation again on each part to obtain the pre-retrieval address and weight information of the two sets of memory modules, and then fuse the two sets of address information to obtain the final retrieval address and weight, and then retrieve the video frame feature information stored in the storage unit from the memory module according to the address, and then weight these feature information into a decoding feature map.

[0024] Step 301: Design a query information generator Design two address index generators Designing a memory module And initialize, where C represents the dimension of the storage information item, that is, the dimension of the encoding feature map of the video frame segment, D represents the dimension of the query information, and nkeys represents the storage capacity;

[0025] Step 302: The query information feature graph Z obtained in step 2 eout Query Information Generator Convert to get a query information

[0026] Step 303: Divide the query information into two sub-query information

[0027] Step 304: Pass the two subquery information through two index generators respectively. The matrix is ​​converted into the front retrieval address ind of the two memory modules. 1 and ind 2 and weight information

[0028] Step 305: Cross-sum the two weight information to obtain the storage unit weight used for retrieval At the same time, the two search addresses are cross-added to obtain the search address ind corresponding to the storage unit weight;

[0029] Step 306: Normalize the storage unit weights using Softmax and select the largest K retrieval weights And the corresponding search address

[0030] Step 307: Retrieve the video frame feature information m stored in the corresponding K rows of storage units from the memory storage module according to the address. i , concatenate the K video frame feature information according to the retrieval weight to obtain the decoded feature map Z din ;

[0031] Step 4: Design a video frame segment decoder and initialize it;

[0032] Step 5: Input the decoded information vector obtained in step 307 into the decoder in the decoding format, and obtain the output sample through forward propagation calculation. The conversion formula is:

[0033]

[0034] in, is the output video frame segment, f d is the video frame segment decoder structure, θ dis the decoder parameter, Z din It is the feature map obtained in step 307.

[0035] In this embodiment, video anomaly detection is to detect and identify video frames or clips of abnormal behavior in the video stream. This experiment will be conducted on the monitoring video stream datasets of pedestrian behavior from two universities, UCSD-Ped2 and CUHK Avenue. The set comparison methods are 3D convolutional autoencoder model method 1, video anomaly detection model TSC method 2, direct index memory network method 3, and method 4 of the present invention. In view of the multi-dimensional complexity of video data, 4 3D convolutional layers and 4 deconvolutional layers are used to implement the encoder and decoder.

[0036] Considering the complexity of the data, the design The coded information is stored by pixel, where K is set to k=2048. In this experiment, reconstruction error is used to represent the anomaly score. When an abnormal frame is inserted into a small video clip, the anomaly score of this clip will be high. Regarding the evaluation indicators, since most models use the form of abnormal scoring, the AUC indicator is used to measure performance, that is, the ROC value is calculated as the AUC indicator through a variable threshold, 0 is the worst and 1 is the best.

[0037] Table 1 shows the test results on the database. It can be seen that in terms of the AUC index, the neural network based on the present invention performs better on both application data sets.

[0038] Table 1 AUC indicators of different methods on application data

[0039]

[0040] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation modes of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A video surveillance anomaly detection method based on memory neural network, characterized in that: include: Acquire video stream data, and pre-process the video stream data to obtain video numerical information; The video frame segment to be detected is obtained from the video numerical information, and the video frame segment to be detected is input into the encoder to extract the feature map of the video frame segment; the feature map of the video frame segment is subjected to matrix transformation to obtain the query information of the video frame segment; the query information is equally divided into two parts, and the divided query information is subjected to matrix transformation respectively to obtain the pre-search address and weight information of two groups of memory modules; the pre-search address and weight information of the two groups of memory modules are integrated to obtain the final search address and weight; the video frame feature information stored in the storage unit is retrieved from the memory module according to the address, and the video frame feature information is weighted and spliced ​​into a decoding feature map; the decoding feature map is input into the decoder, and a reconstructed decoded video frame is output; the reconstructed decoded video frame is subtracted from the video frame segment to be detected, and whether it is abnormal is determined according to the subtraction value, so as to obtain the video surveillance abnormality detection result.

2. The video surveillance anomaly detection method based on memory neural network according to claim 1 is characterized in that: The process of preprocessing video stream data includes: obtaining video stream from the video surveillance terminal; extracting jpg format images of all frames from the video stream, using the numpy tool library to save all frame images as 5D numerical matrix information, and intercepting the information of every 12 images in the first dimension from the numerical matrix information as a frame segment, and collecting all the intercepted frame segment information.

3. The video surveillance anomaly detection method based on memory neural network according to claim 1 is characterized in that: The process of matrix transformation of the feature map of the video frame segment includes: setting the query information generator The query information generator is initialized; the encoded feature map of the video frame segment is input into the query information generator for matrix transformation to obtain the query information Among them, C represents the storage information item dimension, and D represents the query information dimension.

4. The video surveillance anomaly detection method based on memory neural network according to claim 1 is characterized in that: The process of performing matrix transformation on the split query information includes: passing the two sub-query information through two index generators respectively The matrix is ​​converted into the front retrieval address ind of the two memory modules. 1 and ind 2 and weight information Where nkeys represents the storage capacity.

5. The video surveillance anomaly detection method based on memory neural network according to claim 1 is characterized in that: The process of fusing the pre-retrieval addresses of the two groups of memory modules includes: fusing the pre-retrieval addresses of the two groups of memory modules by adopting a cross-summation method.

6. The video surveillance anomaly detection method based on memory neural network according to claim 1 is characterized in that: The process of retrieving the video frame feature information stored in the storage unit from the memory module according to the address includes: setting the memory module And initialize the memory module; each row in the memory module is set with a number, and the stored video frame feature information is retrieved from the memory module M according to the fused index number.

7. The video surveillance anomaly detection method based on memory neural network according to claim 1 is characterized in that: The video frame feature information is weighted and spliced ​​into a decoding feature map: the retrieved video frame feature information is matched with the obtained weight value one by one and multiplied; the multiplied data is spliced ​​into a matrix of a specified dimension to obtain a decoding feature map.

8. The video surveillance anomaly detection method based on memory neural network according to claim 1 is characterized in that: The process of inputting the decoded feature map into the decoder for decoding includes: initializing the video frame segment decoder module, inputting the decoded feature map into the decoder, and obtaining the output result through the forward propagation calculation of the neural network.