Video monitoring image abnormity identification method, device and system and medium
By combining the image abnormality recognition model of visual self-attention transformation network and convolutional neural network, image features are extracted and converted, the problems of strong subjectivity and high complexity of image abnormality recognition methods in the prior art are solved, and more efficient and accurate abnormality recognition is achieved.
Patent Information
- Application Number
- CN202311688642.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-06
AI Technical Summary
The existing video surveillance image abnormality recognition methods have problems such as strong subjectivity and weak anti-interference ability. Traditional image processing methods require manual setting of feature parameters, and traditional machine learning methods require setting different models for each exception type, which is more complex to implement.
An image anomaly recognition model based on a combination of visual self-attention transformation network (Transformer) and convolutional neural network (CNN) is used to extract two-dimensional features in sequence through the image feature extraction module, and convert them into three-dimensional features in image form through the abnormal feature extraction module to perform abnormal detection.
It improves the accuracy and efficiency of abnormal recognition of video surveillance images, can better capture global and local feature information in the image, and improves the recognition effect.
Smart Images

Figure CN120107840A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to a method, device, system and medium for identifying abnormalities in video surveillance images. Background Art
[0002] In the related technologies, video surveillance image anomaly recognition methods are mainly divided into traditional image processing methods and traditional machine learning methods. Among them, traditional image processing methods require manual setting of feature parameters, which are highly subjective and have weak anti-interference capabilities, while traditional machine learning methods require setting different models for each anomaly type, which is more complicated to implement. Summary of the invention
[0003] Based on this, it is necessary to provide a video surveillance image anomaly recognition method, device, system and medium that can improve the accuracy and efficiency of video surveillance image anomaly recognition in response to the above technical problems.
[0004] A method for identifying abnormalities in video surveillance images, comprising:
[0005] Acquire the image frame to be detected;
[0006] An image feature extraction module based on an image anomaly recognition model extracts a first feature of the image frame; the first feature is a two-dimensional feature in a sequence form; wherein the image feature extraction module is built based on a visual self-attention transformation network;
[0007] Based on the abnormal feature extraction module of the image abnormality recognition model, the first feature is converted into a second feature, and abnormality detection is performed according to the second feature to determine the abnormal type of the image frame; the second feature is a three-dimensional feature in image form; the abnormal feature extraction module is built based on a convolutional neural network.
[0008] In the above solution, the abnormal feature extraction module based on the image abnormality recognition model converts the first feature into the second feature, including:
[0009] Performing linear processing on the first feature to obtain a first processing result;
[0010] Performing conversion processing on the first processing result to obtain a second processing result;
[0011] Perform convolution processing on the second processing result to obtain the second feature.
[0012] In the above solution, performing anomaly detection according to the second feature to determine the abnormality type of the image frame includes:
[0013] The second feature is subjected to feature extraction processing, and the abnormal type of the image frame is determined according to the extracted feature.
[0014] In the above solution, before the image feature extraction module based on the image anomaly recognition model extracts the first feature of the image frame, the method includes:
[0015] Build an image anomaly recognition model to be trained;
[0016] Constructing a sample data set, and dividing the sample data set into a training set and a test set; wherein the sample data set includes image frames of various abnormal types and non-abnormal types;
[0017] Using the training set to train the constructed image anomaly recognition model, and determining the loss value corresponding to each training cycle;
[0018] Inputting the test set into the model after each training cycle to determine the accuracy corresponding to each training cycle;
[0019] The loss value, accuracy and image anomaly recognition model corresponding to each training cycle are stored in the database;
[0020] When the training cycle reaches E times, an image anomaly recognition model whose accuracy and loss value meet the set conditions is selected in the database as the final image anomaly recognition model.
[0021] In the above scheme, when the constructed image anomaly recognition model is trained using the training set and the loss value corresponding to each training cycle is determined, the method includes:
[0022] In a training cycle, the training set is divided into B data sets;
[0023] Inputting a copy of the data set into the constructed image anomaly recognition model to obtain a prediction result corresponding to the data set;
[0024] Calculate the loss value corresponding to the data set according to the prediction result corresponding to the data set, and input the next data set into the constructed image anomaly recognition model;
[0025] After completing the loading of B portions of the data set, the loss value corresponding to the current training cycle is calculated based on the loss value corresponding to each portion of the data set.
[0026] In the above solution, the step of calculating the loss value corresponding to the data set according to the prediction result corresponding to the data set includes:
[0027] According to the prediction result corresponding to each image frame in the data set, the loss value corresponding to each image frame is calculated;
[0028] The loss value corresponding to the data set is determined according to the loss value corresponding to each image frame and the number of image frames contained in the data set.
[0029] In the above solution, the image feature extraction module based on the image anomaly recognition model extracts the first feature of the image frame, including:
[0030] comparing the image frame with a previous image frame;
[0031] In the case that the image frame is different from the previous image frame, the first feature is extracted based on an image feature extraction module of an image anomaly recognition model.
[0032] A video surveillance image anomaly recognition device, comprising:
[0033] An acquisition module, used for acquiring an image frame to be detected;
[0034] A feature extraction module, which is used to extract a first feature of the image frame based on an image feature extraction module of an image anomaly recognition model; the first feature is a two-dimensional feature in a sequence form; wherein the image feature extraction module is built based on a visual self-attention transformation network;
[0035] A detection module is used to convert the first feature into a second feature based on the abnormal feature extraction module of the image abnormality recognition model, and perform abnormality detection according to the second feature to determine the abnormal type of the image frame; the second feature is a three-dimensional feature in the form of an image; the abnormal feature extraction module is built based on a convolutional neural network.
[0036] A system includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the above-mentioned video surveillance image anomaly recognition method are implemented.
[0037] A computer-readable storage medium stores a computer program, which implements the steps of the above-mentioned video surveillance image anomaly recognition method when executed by a processor.
[0038] The above-mentioned video surveillance image anomaly recognition method, device, system and medium extract the two-dimensional features of the image frame to be detected through the image feature extraction module of the image anomaly recognition model, which can improve the detection performance in the image anomaly recognition task, and then convert the two-dimensional features of the image frame into three-dimensional features through the abnormal feature extraction module of the image anomaly recognition module, and identify the abnormal type of the image frame based on the three-dimensional features, which can retain the local and hierarchical feature information of the image frame, thereby improving the recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A schematic diagram of the structure of an image anomaly recognition model in one embodiment;
[0040] Figure 2 A schematic diagram of a flow chart of a method for identifying abnormalities in video surveillance images in one embodiment;
[0041] Figure 3 is a schematic diagram of the structure of the Transformer layer in one embodiment;
[0042] Figure 4 Schematic diagram of the structure of a self-attention module (Att1 module) in one embodiment;
[0043] Figure 5 is a schematic diagram of the structure of an MLP module in an embodiment;
[0044] Figure 6 is a schematic diagram of a processing flow of a position encoding module in one embodiment;
[0045] Figure 7 is a schematic diagram of the processing flow of TransBlock in one embodiment;
[0046] Figure 8 Schematic diagram of data processing flow of Att2 module in TransBlock in one embodiment;
[0047] Fig. 9 A schematic diagram of a process of converting a first feature into a second feature by an abnormal feature extraction module in an embodiment;
[0048] Fig.10 A schematic diagram of a training process of an image anomaly recognition module in one embodiment;
[0049] Fig.11 is a schematic diagram of a process for determining a loss value in an embodiment;
[0050] Fig.12 A calculation process of the loss value corresponding to a data set in an embodiment;
[0051] Fig.13 A schematic diagram of a process of image abnormality recognition in one embodiment;
[0052] Fig.14 The structure block diagram of a device for identifying abnormalities in video surveillance images in one embodiment. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0054] Before describing the technical solution of the embodiment of the present application in detail, a brief introduction to the image anomaly recognition solution in the related art is first given.
[0055] Accidents caused by power line failures in video surveillance systems, resulting in fires and serious casualties and property losses, often occur. The probability of fire in such systems is low, but with dense deployment, the probability of fire increases. When different types of electrical failures occur in video surveillance systems, the surveillance images will show different abnormalities such as no picture, horizontal stripes, etc. The image abnormalities caused by power line failures in video surveillance systems can be divided into 5 different types: no video, horizontal stripes, noise, color abnormalities, and deformation and dislocation.
[0056] The types of image abnormalities caused by power line failure in video surveillance systems, such as no video, horizontal stripes, noise, color abnormalities, deformation and dislocation, are similar to video image quality diagnosis. At present, the evaluation methods of video image quality are mainly divided into traditional image processing methods and traditional machine learning methods.
[0057] Traditional image processing methods mainly look for image features that differ between different types of anomalies, and manually set feature detection parameters and thresholds for judgment. This method requires a diagnostic method to be set for each type of anomaly, which is relatively complex to implement and requires manual setting of feature parameters, which is highly subjective and has weak anti-interference capabilities.
[0058] Traditional machine learning mainly evaluates various quality types by training models, which requires training multiple types and is relatively complex to implement.
[0059] Based on this, the present application provides a video surveillance image anomaly recognition method, device, system and medium that can improve detection accuracy and efficiency.
[0060] The implementation details of the technical solution of the embodiment of the present application are described in detail below.
[0061] like Figure 1 As shown, Figure 1The image anomaly recognition model used in the video surveillance image anomaly recognition method of the embodiment of the present application is shown. The image anomaly recognition model consists of two modules, namely, an image feature extraction module and an abnormal feature extraction module. Among them, the image feature extraction module is built based on the visual self-attention transformation network (Transformer). In practical applications, the image feature extraction module is essentially a Tokens-to-Token Vision Transformer (T2TViT). T2T ViT is a visual Transformer model based on the Transformer architecture. Compared with the ordinary Transformer model, T2T ViT can be directly applied to image data, which simplifies the processing flow of the model, improves the efficiency and performance of the model, and is also better able to capture global features in the image.
[0062] The abnormal feature extraction module is built based on convolutional neural networks (CNN). CNN is a commonly used model in image tasks and can extract local features in images well.
[0063] In this embodiment, the Transformer network structure is combined with the CNN network structure to build an image anomaly recognition model, which can combine the advantages of both Transformer and CNN, obtain a more comprehensive and rich feature representation, and improve the abnormal feature recognition of the image and perform accurate classification.
[0064] In practical applications, the image feature extraction module and the abnormal sign extraction module also include different processing modules respectively. The following is a method for identifying abnormalities in video surveillance images in combination with an embodiment of the present application. Figure 1 The various structures of the anomaly recognition model in are explained in detail.
[0065] This application provides a method for identifying abnormalities in video surveillance images, such as Figure 2 As shown, Figure 2 A schematic flow chart of a method for identifying anomalies in video surveillance images is shown, wherein the method for identifying anomalies in video surveillance images may include:
[0066] Step S201, obtaining an image frame to be detected.
[0067] Here, the image frame is an image frame captured by a video surveillance system, wherein the image frame to be detected may be captured by the video surveillance system in real time or in non-real time.
[0068] In this embodiment, the image frame to be detected is taken as an example of a real-time captured image frame. The video image stream is captured in real time by a surveillance camera. For a digital video stream, a decoding operation is required to parse the video data into image frames, so that the acquired image frames can be further processed.
[0069] Step S202: extracting a first feature of the image frame based on an image feature extraction module of the image anomaly recognition model.
[0070] Here, the acquired image frame is input into the image anomaly recognition model. First, the image feature extraction module in the image anomaly recognition model extracts features from the input image frame to obtain a first feature, wherein the first feature here is a two-dimensional feature of the image frame, represented in the form of a sequence. In practical applications, the extracted first feature is a global feature of the image frame.
[0071] The following is a detailed description of the image feature extraction module's processing of image frames. Figure 1 As shown in the figure, the image feature extraction module includes an unfolding module (Unfold), a Transformer layer, a position encoding module (Cls-Pos), a Transformer block (Transblock), and a layer normalization module (Layer Norm). The image feature extraction module processes the input image frame as follows:
[0072] Step 1: The input image frame will first be processed by the expansion module, which divides the image frame into fixed-size patches. These patches can be regarded as representations of different perspectives or scales of the input data. Each patch contains local information in the image. The processing here is to convert the image information into sequence data for input into the Transformer layer.
[0073] Step 2: The Transformer layer processes the multiple image blocks obtained by the expansion module. The Transformer layer includes a self-attention mechanism and a multi-layer perceptron (MLP). In the self-attention mechanism, the dependencies between different positions in the sequence are calculated to capture the long-range dependencies in the sequence, which can help the model understand the importance of different positions in the sequence and better capture the characteristics of the sequence data. In the MLP, the features of each position are nonlinearly transformed to enhance the expressiveness of the model.
[0074] Step 3: Use the expansion module again to process the data output by the Transformer layer in step 2.
[0075] Step 4: Use the Transformer layer again to process the data obtained in step 3.
[0076] Step 5, use the expansion module again to process the data processed in step 4.
[0077] Among them, the reason for performing two expansion module-Transformer layer processing on the input image in the image feature extraction module is mainly to extract feature information of different scales and levels in the image. This multi-scale and multi-level feature extraction helps the model to better understand the image content, thereby improving the performance of the model.
[0078] In the first expansion module-Transformer layer processing, the image frame is divided into larger image blocks, which helps the model capture the overall features and global information in the image. In the second expansion module-Transformer layer processing, the image is divided into medium-sized image blocks, which helps the model capture medium-scale features and structural information in the image, so as to understand the image content more comprehensively. This multi-scale and multi-level feature extraction helps improve the model's ability to express image data and helps the model better cope with features of different scales and levels.
[0079] Step 6, input the data output in step 5 into the position encoding module. The function of the position encoding module can be understood as position encoding the input data.
[0080] Step 7. In this embodiment, there are 7 TransBlocks in total. TransBlock is a basic component unit. TransBlock is usually stacked in multiple layers to build a deeper network structure, so as to better learn data feature representation.
[0081] Step 8: Use the layer normalization module to normalize the data output by TransBlock to obtain the first feature of the image frame.
[0082] In practical applications, reference Figure 3 As shown, Figure 3 The schematic diagram of the structure of the Transformer layer is shown. The Transformer layer consists of a layer normalization module, a self-attention module (Att1), an MLP module, a regularization module, and an addition module (Add). The data processing flow in the Transformer layer is as follows:
[0083] Step 1: The layer normalization module normalizes the input data, which helps alleviate the gradient vanishing and gradient exploding problems and improves the model training effect. The input data is normalized by the layer normalization module to obtain data F 1.
[0084] Step 2: Use the Att1 module to process the data F 1 The Att1 module is used to capture the correlation information between different positions in the input sequence and weight the input feature representation based on this correlation information. Through the self-attention module, the model can focus on the information related to the current position in the input sequence, so as to better understand the global dependency of the sequence data and achieve better representation learning and feature extraction. 1 After processing, the self-attention module can output data F 2 .
[0085] Step 3: Get the data F 2 Input to the layer normalization module for normalization, and then the normalized data F 2 Input to the MLP module, and the data F 2 The processing can further help the model to extract features and learn representations of input data, improve the model's representation and generalization capabilities, alleviate the overfitting problem, and help improve the performance of the model.
[0086] Step 4: Use the regularization module to process the data output by the MLP module, where the regularization module here is Drop Path. Drop Path is a network regularization method proposed for branch networks. It selects the network structure by generating a series of masks. Where mask = 1, the corresponding network structure is retained, and where mask = 0, the part of the network structure is invalidated. This can randomly delete the multi-branch structure in the deep learning model, which can prevent overfitting, improve model performance, and overcome the problem of network degradation. Here, after Drop Path processing, the data F is obtained. 3 .
[0087] Step 5: The data F output by the regularization module 3 And the data F output by the self-attention module 2 The residual connection is introduced by adding the data. Residual connection is a common technique used to alleviate the gradient vanishing problem in deep neural networks and help information propagate better in the network. 2 With data F 3 By adding the data, the data F 2 The information is directly passed to the subsequent layers, which helps to avoid gradient disappearance and speed up the convergence of the model.
[0088] In practical applications, reference Figure 4 As shown, Figure 4The structure diagram of the self-attention module (Att1 module) is shown. The Att1 module consists of a linear transformation module (Linear), a regularization module and an addition module (Add). The process of Att1 module processing data is described in detail below.
[0089] Step 1: Input the data into the linear transformation module. The linear transformation module is used to perform linear changes on the input data. Through linear transformation, the input data can be mapped into a query vector (query, hereinafter referred to as q), a key keyword vector (key, hereinafter referred to as k) and a value vector (value, hereinafter referred to as v). These vectors are used to calculate the attention weight and weighted sum process respectively.
[0090] Step 2, perform a dot product operation on q and k. The dot product operation obtains the similarity score between them, and then calculates the correlation between each q and all k through a softmax operation, and then normalizes them to obtain the attention weight. The purpose of the softmax operation is to convert the similarity score into a probability distribution in order to represent the weight corresponding to each key.
[0091] Step 3, use the regularization module to process the attention weight, where the regularization module here refers to Drop Out. The idea of Drop Out is similar to Drop Path. The difference is that Drop Path acts on the network branch, while Drop Out acts on the neural network unit. Through Drop Out, the neural network unit is temporarily discarded from the network with a certain probability, thereby eliminating the weakened joint adaptability between neuron nodes and enhancing the generalization ability of the model.
[0092] Step 4, perform weighted summation of the attention weight and v to obtain the weighted average result. This step can be understood as performing a weighted average of v according to the attention weight to obtain a representation weighted according to the degree of attention.
[0093] Step 5: Perform a linear transformation on the result obtained by weighted summation, and process the regularization module (Drop Out). Linear transformation can help the model learn more complex feature representations, while Drop Out helps reduce the risk of overfitting.
[0094] Step 6, add the result obtained by linear transformation and the second regularization process to the original v. This step can be understood as residually linking the result after attention weighting and linear transformation with the original v to obtain the final self-attention representation. The main purpose of residual connection is to solve the gradient vanishing and gradient exploding problems in deep neural networks. In deep neural networks, after multiple layers of nonlinear transformation, the gradient will gradually become smaller, leading to the gradient vanishing problem during training. Residual connection can help the gradient back propagate better, thereby alleviating the gradient vanishing problem. In addition, residual connection also helps in model training and optimization.
[0095] Based on the above steps 1 to 6, the data processing flow of the Att1 module is completed.
[0096] In practical applications, reference Figure 5 As shown, Figure 5 The schematic diagram of the structure of the MLP module is shown, and the MLP module includes a linear transformation module (Linear), an activation function module and a regularization module. The following is a detailed description of the data processing process of the MLP module.
[0097] Step 1: The linear transformation module linearly changes the data input to the MLP module so that the model can learn the linear relationship of the input data and map the input data to a higher-dimensional space, which helps to extract the feature representation of the data.
[0098] Step 2: The activation function module processes the data output by the linear transformation module. The activation function module uses the GELU activation function, which is a nonlinear activation function that performs well in deep learning and can help the model learn nonlinear relationships and improve the model's representation ability.
[0099] Step 3, the regularization module in the MLP module uses Drop Out. The idea of Drop Out is similar to that of Drop Path. The difference is that Drop Path acts on the network branches, while Drop Out acts on the neural network units. Through Drop Out, the neural network units are temporarily discarded from the network with a certain probability, thereby eliminating the weakened joint adaptability between neuron nodes and enhancing the generalization ability of the model.
[0100] Step 4: Use the linear transformation module again to perform a second linear transformation on the data output by the regularization module. The second linear transformation can further extract the feature representation of the data, which helps the model learn more complex features.
[0101] Step 5: Input the data after the second linear transformation into the regularization module again for the second DropOut processing, which can further reduce the overfitting of the model and improve the generalization ability of the model.
[0102] In summary, the data processing process of the MLP module is completed through steps 1 to 5. The MLP module can help the model extract features and learn representations of input data, improve the model's representation and generalization capabilities, and alleviate the overfitting problem, which helps improve the performance of the model.
[0103] In practical applications, reference Figure 6 As shown, Figure 6 A schematic diagram of the processing flow of the position encoding module is shown, and the data processing process of the position encoding module is described in detail below.
[0104] Step 1: Concatenate the input data with the Cls tokens, so that the input data can be concatenated with a special token used to represent the entire sequence. The Cls tokens here are usually used to represent the special mark of the entire sequence to provide summary information of the entire sequence.
[0105] Step 2, adding the result obtained by the connection operation to the position code (Pos embed), the position code is used to represent the position information of each position in the input data, so that the position information can be introduced into the data.
[0106] In step 3, the data carrying the position information will be processed by the regularization module, where DropOut is used for regularization. DropOut is a regularization technique that can randomly set the output of some neurons to 0 to prevent overfitting.
[0107] Based on the above steps 1 to 3, the data processing flow of the position encoding module is completed.
[0108] In practical applications, reference Figure 7 As shown, Figure 7 The processing flow diagram of TransBlock is shown, and the data processing process of TransBlock is described in detail below.
[0109] Step 1: The input data is normalized by the layer normalization module. By normalizing the input data, the numerical ranges of different features can be made relatively consistent, which helps to improve the training stability and convergence speed of the model.
[0110] In step 2, the normalized data is processed by the self-attention module (Att2 module) to obtain the attention representation of the input data. Through the Att2 module, the model can perform weighted integration of information at different positions in the input data, thereby better capturing the key information and contextual relationships in the input sequence.
[0111] Step 3: The result processed by the Att2 module is then processed by the regularization module to obtain the first processing result, wherein the regularization module here adopts Drop Patch.
[0112] Step 4, adding the input data and the first processing result to obtain the second processing result. Adding the input data and the result processed by the Att2 module and the regularization module can make the model combine different information at different levels, which helps to improve the expression ability of the model.
[0113] In step 5, the second processing results are normalized by the layer normalization module, and then nonlinearly transformed by MLP. The results after MLP processing are then processed by the regularization module to obtain the third processing result. Here, normalizing the second processing result can ensure the stability and convergence speed of the model; nonlinear transformation by MLP can help the model learn more complex feature representations and improve the expressiveness of the model; regularization again in step 5 can prevent the model from overfitting and increase the generalization ability of the model.
[0114] Step 6: Add the third processing result to the second processing result to further integrate information at different levels to obtain the output result of TransBlock.
[0115] It should be noted that the Att2 module in TransBlock is slightly different from the Att1 module in the Transformer layer. The Att1 module in the Transformer layer contains residual connections, while the Att2 module in the Transformer layer contains residual connections. Figure 8 Schematic diagram of the data processing flow of the Att2 module in the TransBlock shown in FIG. 1 . The Att2 module in the TransBlock does not include a residual connection processing flow, but other data processing flows are consistent with the Att1 module in the Transformer layer.
[0116] Based on the above steps 1 to 6, TransBlock completes the data processing flow.
[0117] Step S203: Based on the abnormal feature extraction module of the image abnormality recognition model, the first feature is converted into a second feature, and abnormality detection is performed according to the second feature to determine the abnormal type of the image frame.
[0118] Here, the first feature obtained after the image feature extraction module extracts features from the image frame is in a sequence form, represented as a two-dimensional feature. In order to further extract significant difference features between different categories and retain the feature information of the image block, it is necessary to process the sequence form represented in two dimensions into an image form represented in three dimensions in the spatial dimension. In this embodiment, the abnormal feature extraction module of the image abnormality recognition model is used to convert the first feature extracted by the image feature extraction module into a second feature, wherein the second feature is a three-dimensional feature of the image frame, represented in the form of an image. In practical applications, the second feature here is a local feature of the image frame.
[0119] After the abnormal feature extraction module converts the first feature into the second feature, the second feature is optimized to further determine the significant difference features between different categories and finally determine the abnormal category of the image frame.
[0120] In one embodiment, reference Fig. 9 As shown, Fig. 9 The schematic diagram of the process of converting the first feature into the second feature by the abnormal feature extraction module is shown below. Figure 1 and Fig. 9 The process of converting the first feature into the second feature by the abnormal feature extraction module is described in detail.
[0121] Step S901, performing linear processing on the first feature to obtain a first processing result.
[0122] The first feature is input into the fully connected layer (Linear) for linear processing. Specifically, the first feature is flattened into a one-dimensional vector and input into the fully connected layer for linear transformation and feature extraction to obtain a first processing result. Through linear transformation of the fully connected layer r, the high-dimensional features of the image can be mapped to a lower-dimensional feature space to extract more representative features.
[0123] Assume that the first feature T∈(L+1)×S, after linear processing, the first processing result T′=Linear(T), T′∈L×S, where represents the real number domain, that is, T∈(L+1)×S means that the dimension of feature T is (L+1)×S, and each element is a real value, not a discrete category or label.
[0124] Step S902: convert the first processing result to obtain a second processing result.
[0125] The first processing result is input to a subsequent view processing (View) module. The view processing module here is usually used to change the function of the shape of the tensor. The size or dimension of the tensor can be changed without changing the data through the view processing module.
[0126] Suppose there is a 2D feature tensor x with shape (batch_size, feature_dim), which means there are batch_size samples and each sample has feature_dim features. Select a new 3D feature shape, which includes a new height (rows) and width (columns). Use the view processing module to convert the 2D feature tensor to a new 3D feature tensor, making sure the new shape is consistent with the number of elements in the original tensor. x_3d is a new 3D feature tensor with shape (batch_size, new_height, new_width)
[0127] Based on this, the second processing result output by the view processing module has been converted into a three-dimensional feature, wherein the second processing result F=View(T′), F∈H×w×C.
[0128] Step S903, performing convolution processing on the second processing result to obtain a second feature.
[0129] Here, the convolution processing is performed by a convolutional layer (Convolutional Layer), a batch normalization (BN) unit, and a linear rectifier unit (ReLU).
[0130] First, the second processing result is processed using a convolutional layer. The convolutional layer is an important component of a convolutional neural network. The convolutional layer extracts the features of the input data by applying a convolution operation, which can effectively capture the local patterns and spatial relationships of the input data, thereby achieving feature extraction and feature mapping.
[0131] Secondly, batch normalization is performed. Batch normalization makes the network training more stable and faster by normalizing the input data of each batch, which helps to solve problems such as gradient disappearance and gradient explosion.
[0132] Finally, ReLU is used to process the data. ReLU is one of the commonly used activation functions. The ReLU function returns the input value when the input is greater than 0, and returns 0 when the input is less than or equal to 0. Its nonlinear characteristics help the network learn nonlinear function relationships, and compared with traditional activation functions, ReLU can alleviate the gradient vanishing problem.
[0133] In summary, the processing process of the second processing result can be expressed as Among them, X represents the second feature, Denotes convolution processing, including convolution operation, batch normalization and linear rectifier unit processing. By performing convolution processing on the second processing result, more significant difference features between different categories and more complete feature information of the image block can be obtained.
[0134] In one embodiment, after the abnormal feature extraction module converts the first feature into the second feature, it will identify the abnormal type of the image frame based on the second feature. Based on the second feature, the abnormal feature extraction module will further extract the significant difference features between different categories and perform optimization processing, and finally convert them into specific abnormal categories. The specific calculation process is:
[0135] Y=Φ(X)
[0136] In the above formula, Φ(·) represents the feature extraction and conversion to abnormal category operations, including batch normalization, regularization module, flattening operation (Flatten), fully connected layer (Linear), layer normalization (L2 Normalization). Figure 1 The processing flow of the second feature will be described in detail.
[0137] First, batch normalization is usually applied to the convolutional layer or fully connected layer of a convolutional neural network. For the second feature, batch normalization will perform standardization at the same position of each sample, which helps alleviate the internal covariate shift in the training process, improve the training speed of the model, and accelerate convergence.
[0138] Secondly, the regularization module is used to process the batch normalized data. The regularization module here adopts the DropOut mechanism, which can randomly set a part of neurons to zero to prevent overfitting.
[0139] Then, the data output by the regularization module is flattened into a one-dimensional vector. In the transition layer of the neural network, the flattening operation can convert the three-dimensional features into one dimension to prepare for the input of the fully connected layer.
[0140] Next, the flattened one-dimensional features are connected to the fully connected layer. The fully connected layer maps the input to the output through the linear transformation of the weight matrix, combines and nonlinearly transforms the features, and prepares for the final classification.
[0141] Next, the data output by the fully connected layer is batch normalized again to help stabilize the training process.
[0142] Next, the batch normalized data is subjected to layer normalization. Layer normalization will standardize the data in the channel dimension, which helps the stability and generalization ability of the network.
[0143] Finally, the fully connected layer maps the previously processed features to the output of the number of categories. The fully connected layer completes the final classification operation and obtains the probability distribution of each abnormal category, thereby obtaining the abnormal category of the image frame.
[0144] The following is a detailed description of the training process of the image anomaly recognition module. Fig.10 As shown, Fig.10 A schematic diagram of the training process of the image anomaly recognition module is shown.
[0145] Step S1001, constructing an image anomaly recognition model to be trained.
[0146] First, it is necessary to select the architecture of the image anomaly recognition model. In this embodiment, the architecture of the image anomaly recognition model is a combination of the Transformer architecture and the CNN architecture. The image feature extraction module and the abnormal feature extraction module in the image anomaly recognition model are respectively constructed to complete the construction of the image anomaly recognition model.
[0147] Step S1002: construct a sample data set, and divide the sample data set into a training set and a test set.
[0148] Here, image data sets of 5 abnormal types and other non-abnormal types are collected and obtained from surveillance cameras or other sources to construct a sample data set. That is, the sample data set contains 6 types of image data.
[0149] For supervised learning, it is also necessary to provide labels for each image in the sample dataset. The labels can be divided into normal and specific abnormal categories (type labels C = 0, 1, 2, 3, 4, 5). The labeling of each image can be completed by manual labeling or using automatic labeling tools.
[0150] In practical applications, images in the sample dataset also need to be preprocessed, including but not limited to:
[0151] (1) Resize images: Make sure all images have the same size.
[0152] (2) Regularization: Normalize pixel values to a fixed range, such as [0, 1].
[0153] (3) Data enhancement: Increase data diversity through operations such as rotation, flipping, and scaling.
[0154] For the constructed sample data set, the sample data set is divided into a training set and a test set in a 4:1 ratio. The training set is used for model training, and the test set is used to finally evaluate the generalization ability of the model.
[0155] Step S1003, using the training set to train the constructed image anomaly recognition model, and determining the loss value corresponding to each training cycle.
[0156] After the processing of the sample data set is completed, the divided training set is loaded into the image anomaly recognition model constructed in step S1001 for training. After a training cycle is completed, the loss value of the model corresponding to this training cycle needs to be calculated.
[0157] Computing loss plays a key role in model training. It is an objective function in the training process that is used to measure the performance of the model on the current training batch or data, thereby guiding model updates.
[0158] In one embodiment, Fig.11 As shown, Fig.11 A schematic diagram of the process of determining the loss value is shown.
[0159] Step S1101: In one training cycle, the training set is divided into B data sets.
[0160] In practical applications, directly inputting the entire training set into the model for training may occupy a large amount of memory. Based on this, in a training cycle, the training set is divided into B data sets, where each data set contains N image frames, and each divided data set contains images and corresponding type labels. Using the divided B data sets, the data can be divided into small batches for processing, which reduces memory usage, accelerates model convergence, and improves the generalization ability of the model.
[0161] It should be noted that a training cycle refers to the process in which the model performs forward propagation and back propagation on the entire training set. In a complete training cycle, the model will input the entire training data set into the model in turn for forward propagation, calculate the loss function, and then perform back propagation to update the model parameters. This iterative process continues until all images in the entire training set have been used once, completing a complete training cycle.
[0162] Step S1102: input a data set into the constructed image anomaly recognition model to obtain a prediction result corresponding to the data set.
[0163] Here, a data set is input into the image anomaly recognition model and the corresponding prediction results are obtained. This process is usually called forward propagation. In this process, the data set will pass through the input layer of the model, then through the hidden layer, and finally reach the output layer. Each layer will perform certain transformations and processing on the data set to finally obtain the prediction results. Among them, the prediction results can be used to calculate the loss function, evaluate the model performance, etc.
[0164] Step S1103, calculate the loss value corresponding to the data set according to the prediction result corresponding to the data set, and input the next data set into the constructed image anomaly recognition model.
[0165] Here, the loss value is used to measure the difference between the model's prediction results and the actual labels. The calculation of the loss value is achieved through the loss function, which is determined at the beginning of the model training.
[0166] Reference Fig.12 As shown, Fig.12 The calculation process of the loss value corresponding to the data set is shown.
[0167] Step S1201, based on the prediction result corresponding to each image frame in the data set, calculate the loss value corresponding to each image frame.
[0168] First, calculate the loss value corresponding to each image frame in the dataset according to the following formula.
[0169]
[0170] Among them, b refers to one of the B data sets, n is one of the N image frames, indicating the prediction value of the cth category, and C refers to the number of abnormal categories. In this embodiment, C=6, that is, the trained image anomaly recognition model can recognize six different images, including normal images and five different abnormal images.
[0171] Step S1202, determining the loss value corresponding to the data set according to the loss value corresponding to each image frame and the number of image frames contained in the data set.
[0172] Next, calculate the loss value corresponding to the data set according to the following formula.
[0173]
[0174] It can be understood that for the b-th dataset, the loss values of the N image frames in the b-th dataset are calculated, and then the loss values of the N image frames in the b-th dataset are averaged to obtain the loss value corresponding to the b-th dataset. The loss value corresponding to the dataset calculated here can better reflect the loss of the entire dataset and help evaluate the overall performance of the model.
[0175] After determining the loss value corresponding to the data set, the next data set is loaded into the image anomaly recognition model to further determine the loss value corresponding to the next data set.
[0176] Step S1104, after completing the loading of B data sets, the loss value corresponding to the current training cycle is calculated based on the loss value corresponding to each data set.
[0177] Here, completing the loading of B sets of data means that all image frames in the entire training set have been used once, that is, a training cycle of the model has been completed. At this time, it is necessary to determine the loss value corresponding to this training cycle. In practical applications, the loss value corresponding to the training cycle can be calculated by the following formula:
[0178]
[0179] Among them, represents the loss value corresponding to a training cycle. It can be determined from the above formula that the loss value corresponding to the training cycle is obtained by adding the loss values corresponding to each data set.
[0180] Step S1004, input the test set into the model of each training cycle to determine the accuracy corresponding to each training cycle.
[0181] After the model completes a training cycle, the test set is used to determine the accuracy of the model after completing this training cycle. The test set is input into the model that has completed a training cycle. The model will infer the abnormal type of the test set, compare the model's prediction results with the true label of the test set, and calculate the accuracy of the model.
[0182] Step S1005, storing the loss value, accuracy rate and image anomaly recognition model corresponding to each training cycle into a database.
[0183] Here, the loss value can reflect the degree of model fitting and guide parameter updates, the accuracy rate can reflect the overall performance of the model, and the image anomaly recognition model refers to the parameters of the image anomaly recognition model after completing a training cycle. By saving the loss value, accuracy rate, and image anomaly recognition model of each training cycle, the image anomaly recognition model for the final application can be screened from the image anomaly recognition models trained in multiple training cycles.
[0184] Step S1006, when the training cycle reaches E times, an image anomaly recognition model whose accuracy and loss value meet the set conditions is selected from the database as the final image anomaly recognition model.
[0185] Here, before starting model training, a maximum training cycle E is set to control the training duration or resource consumption to avoid overfitting or long training time.
[0186] In the process of model training, the model needs to be trained for E cycles, that is, the above steps S1101 to S1104 are repeated E times. When the training cycle reaches E times, it indicates that the condition for stopping training is met. At this time, the final image anomaly recognition model is selected in the set database according to the accuracy and loss value corresponding to each training cycle, wherein the accuracy and loss value of the determined image anomaly recognition model can meet the set conditions, thereby ensuring the detection accuracy of the image anomaly recognition model.
[0187] In practical applications, the model with the highest accuracy and the lowest loss value will be selected from the database as the final image recognition model.
[0188] In one embodiment, before the acquired image frame is input into the image anomaly recognition model, it is necessary to compare the acquired image frame with the previous image frame to determine whether the acquired image frame is different from the previous image frame. It can be understood that when the acquired image frame is the same as the previous image frame, it means that the abnormal recognition result of the acquired image frame is the same as the abnormal recognition result of the previous image frame. Based on this, the detection process of the acquired image frame can be omitted, and the new image frame can be directly extracted for processing, thereby avoiding repeated recognition processing of the same image. When the acquired image frame is different from the previous image frame, it is considered that there is a possibility that the abnormal recognition result of the acquired image frame is inconsistent with the abnormal recognition result of the previous image frame, and the image anomaly recognition model needs to be used to identify the image frame.
[0189] In practical applications, image similarity measurement methods (such as mean square error, structural similarity index, etc.) can be used to compare the similarity between the acquired image frame and the previous image frame. If the similarity is lower than the threshold, the acquired image frame is input into the model for recognition. If the similarity is higher than the threshold, a new image frame is directly extracted.
[0190] refer to Fig.13 As shown, Fig.13 A schematic diagram of a process of image anomaly recognition is shown.
[0191] Step 1: Obtain real-time monitoring image frames.
[0192] Step 2, determine whether the image frame has changed. If the image frame has changed, execute step 3; if the image frame has not changed, execute step 1.
[0193] Step 3: Input the image frame into the image anomaly recognition model.
[0194] Step 4, determine whether the image frame is abnormal. If the image frame is changed, execute step 5; if the image frame is not changed, execute step 1.
[0195] Step 5: Output the abnormal type of the image frame.
[0196] In the above embodiment, the image feature extraction module in the image anomaly recognition model is built based on the Transformer network, which can recognize images and improve the detection performance; on the basis of the original image feature module in the image anomaly recognition model, an abnormal feature extraction module based on the CNN network is added, which can retain the local and hierarchical feature information of the image block, improve the recognition effect, and improve the accuracy of abnormality recognition.
[0197] In one embodiment, a video surveillance image abnormality recognition device is provided, referring to Fig.13 As shown, the video surveillance image anomaly recognition device 1300 may include: an acquisition module 1301, a feature extraction module 1302, a detection module 1303 and a training module 1304.
[0198] Among them, the acquisition module 1301 is used to acquire the image frame to be detected; the feature extraction module 1302 is used to extract the first feature of the image frame based on the image feature extraction module of the image anomaly recognition model; the first feature is a two-dimensional feature in the form of a sequence; wherein the image feature extraction module is built based on a visual self-attention transformation network; the detection module 1303 is used to convert the first feature into a second feature based on the abnormal feature extraction module of the image anomaly recognition model, and perform abnormality detection based on the second feature to determine the abnormal type of the image frame; the second feature is a three-dimensional feature in the form of an image; the abnormal feature extraction module is built based on a convolutional neural network.
[0199] In one embodiment, the detection module 1303 is specifically used to perform linear processing on the first feature to obtain a first processing result; perform conversion processing on the first processing result to obtain a second processing result; and perform convolution processing on the second processing result to obtain a second feature.
[0200] In one embodiment, the detection module 1303 is specifically used to perform feature extraction processing on the second feature, and determine the abnormal type of the image frame according to the extracted feature.
[0201] In one embodiment, the video surveillance image anomaly recognition device 1300 also includes a training module 1304. Before the image feature extraction module based on the image anomaly recognition model extracts the first feature of the image frame, the training module 1304 is specifically used to construct an image anomaly recognition model to be trained; construct a sample data set, and divide the sample data set into a training set and a test set; wherein the sample data set contains image frames of various abnormal types and non-abnormal types; use the training set to train the constructed image anomaly recognition model to determine the loss value corresponding to each training cycle; input the test set into the model after each training cycle to determine the accuracy rate corresponding to each training cycle; store the loss value, accuracy rate and image anomaly recognition model corresponding to each training cycle into a database; when the training cycle reaches E times, select the image anomaly recognition model whose accuracy rate and loss value meet the set conditions in the database as the final image anomaly recognition model.
[0202] In one embodiment, the training module 1304 is specifically used to, in a training cycle, divide the training set into B data sets; input one data set into the constructed image anomaly recognition model to obtain a prediction result corresponding to the data set; calculate the loss value corresponding to the data set according to the prediction result corresponding to the data set, and input the next data set into the constructed image anomaly recognition model; after completing the loading of B data sets, calculate the loss value corresponding to this training cycle according to the loss value corresponding to each data set.
[0203] In one embodiment, the training module 1304 is specifically used to calculate the loss value corresponding to each image frame according to the prediction result corresponding to each image frame in the data set; and determine the loss value corresponding to the data set according to the loss value corresponding to each image frame and the number of image frames contained in the data set.
[0204] In one embodiment, the feature extraction module 1302 is specifically used to compare the image frame with the previous image frame; when the image frame is different from the previous image frame, the image feature extraction module based on the image anomaly recognition model extracts the first feature.
[0205] For the specific definition of the video surveillance image anomaly recognition device, please refer to the definition of the video surveillance image anomaly recognition method above, which will not be repeated here. Each module in the above-mentioned video surveillance image anomaly recognition device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0206] In one embodiment, a system is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements a method for identifying abnormalities in video surveillance images when executing the computer program.
[0207] In one embodiment, a computer storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, a method for identifying anomalies in video surveillance images is implemented.
[0208] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0209] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0210] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0211] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0212] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. A method for identifying abnormalities in video surveillance images. It is characterized in that include: Acquire the image frame to be detected; An image feature extraction module based on an image anomaly recognition model extracts a first feature of the image frame; The first feature is a two-dimensional feature in the form of a sequence; wherein the image feature extraction module is built based on a visual self-attention transformation network; Based on the abnormal feature extraction module of the image abnormality recognition model, the first feature is converted into a second feature, and abnormality detection is performed according to the second feature to determine the abnormal type of the image frame; the second feature is a three-dimensional feature in image form; the abnormal feature extraction module is built based on a convolutional neural network.
2. The method according to claim 1, It is characterized in that The abnormal feature extraction module based on the image abnormality recognition model converts the first feature into a second feature, including: Performing linear processing on the first feature to obtain a first processing result; Performing conversion processing on the first processing result to obtain a second processing result; Perform convolution processing on the second processing result to obtain the second feature.
3. The method according to claim 1, It is characterized in that The performing abnormality detection according to the second feature to determine the abnormality type of the image frame includes: The second feature is subjected to feature extraction processing, and the abnormal type of the image frame is determined according to the extracted feature.
4. The method according to claim 1, It is characterized in that Before the image feature extraction module based on the image anomaly recognition model extracts the first feature of the image frame, the method includes: Build an image anomaly recognition model to be trained; Constructing a sample data set, and dividing the sample data set into a training set and a test set; wherein the sample data set includes image frames of various abnormal types and non-abnormal types; Using the training set to train the constructed image anomaly recognition model, and determining the loss value corresponding to each training cycle; Inputting the test set into the model after each training cycle to determine the accuracy corresponding to each training cycle; The loss value, accuracy and image anomaly recognition model corresponding to each training cycle are stored in the database; When the training cycle reaches E times, an image anomaly recognition model whose accuracy and loss value meet the set conditions is selected in the database as the final image anomaly recognition model.
5. The method according to claim 4, It is characterized in that When the constructed image anomaly recognition model is trained using the training set to determine the loss value corresponding to each training cycle, the method includes: In a training cycle, the training set is divided into B data sets; Inputting a copy of the data set into the constructed image anomaly recognition model to obtain a prediction result corresponding to the data set; Calculate the loss value corresponding to the data set according to the prediction result corresponding to the data set, and input the next data set into the constructed image anomaly recognition model; After completing the loading of B portions of the data set, the loss value corresponding to the current training cycle is calculated based on the loss value corresponding to each portion of the data set.
6. The method according to claim 5, It is characterized in that The calculating the loss value corresponding to the data set according to the prediction result corresponding to the data set includes: According to the prediction result corresponding to each image frame in the data set, the loss value corresponding to each image frame is calculated; The loss value corresponding to the data set is determined according to the loss value corresponding to each image frame and the number of image frames contained in the data set.
7. The method according to claim 1, It is characterized in that The image feature extraction module based on the image anomaly recognition model extracts the first feature of the image frame, including: comparing the image frame with a previous image frame; When the image frame is different from a previous image frame, an image feature extraction module based on an image anomaly recognition model extracts the first feature.
8. A video surveillance image anomaly recognition device, It is characterized in that include: An acquisition module, used for acquiring an image frame to be detected; A feature extraction module, which is used to extract a first feature of the image frame based on an image feature extraction module of an image anomaly recognition model; the first feature is a two-dimensional feature in a sequence form; wherein the image feature extraction module is built based on a visual self-attention transformation network; A detection module is used to convert the first feature into a second feature based on the abnormal feature extraction module of the image abnormality recognition model, and perform abnormality detection according to the second feature to determine the abnormal type of the image frame; the second feature is a three-dimensional feature in the form of an image; the abnormal feature extraction module is built based on a convolutional neural network.
9. A video surveillance system, comprising a memory and a processor, wherein the memory stores a computer program, It is characterized in that When the processor executes the computer program, the steps of the video surveillance image anomaly identification method described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the method for identifying anomalies in video surveillance images described in any one of claims 1 to 7 are implemented.