Artificial Intelligence-Based Video Anomaly Recognition and Processing Method and System
By adopting artificial intelligence-based methods in video abnormality recognition, including video compression, keyframe extraction, spatiotemporal feature extraction and motion feature extraction, the problems of large calculation volume, high false alarm rate and single spatiotemporal information processing in the prior art are solved, and efficient and accurate video abnormality recognition is achieved.
Patent Information
- Application Number
- CN202510388932.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The prior art has a large amount of calculation and high false alarm rate in video abnormality recognition, so it is impossible to effectively handle complex scenarios, and the processing of space-time information is single, so the correlation between time and space dimensions cannot be considered at the same time.
Using the video abnormality recognition and processing method based on artificial intelligence, we use the video to obtain and compress the monitoring video, extract keyframes for object detection, locate abnormal video clips, extract spatiotemporal features using appearance extraction model and spatiotemporal attention module, combine TV-L1 algorithm and motion extraction model to obtain motion features, and feature fusion is performed through the gated fusion mechanism, and abnormal detection is performed using normal mode memory.
It greatly reduces the storage and transmission burden of video data, improves the accuracy and efficiency of abnormal identification, can identify abnormal events in the space-time dimension, and improves the reliability of abnormal detection.
Smart Images

Figure CN119904785B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video anomaly recognition, and in particular to a video anomaly recognition processing method and system based on artificial intelligence. Background Art
[0002] Traditional methods usually rely on frame-by-frame processing of the entire video frame without selective analysis of key frames or key areas in the video. This approach results in a very large amount of calculation, especially when processing long videos or high-resolution videos, the computing resource consumption is extremely serious; and traditional methods often rely on rule-based detection algorithms or simple statistical methods, which may not be able to handle very complex scenes and are prone to more false positives or omissions. For example, ordinary motion detection algorithms may not be able to distinguish between normal background motion and abnormal events, resulting in a large number of false positives; and traditional methods often process spatiotemporal information in a relatively simple manner and cannot simultaneously consider the association between time and space dimensions. For example, traditional motion detection usually relies on simple optical flow methods to analyze the motion of objects, but does not have the ability to understand the pattern and change of motion from a timing perspective like modern deep learning methods; and traditional methods of anomaly detection often rely on fixed background modeling methods, such as background subtraction or difference map methods, which perform poorly in environments with complex or dynamically changing backgrounds. For example, in a multi-object background, background subtraction methods may misjudge the motion of normal objects as abnormal. Summary of the invention
[0003] The technical problem to be solved by the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a video anomaly recognition and processing method and system based on artificial intelligence.
[0004] The technical solution adopted to solve the above technical problems is: a video anomaly recognition and processing method based on artificial intelligence, including:
[0005] Acquire a surveillance video in a target area, and compress the surveillance video to obtain a compressed surveillance video;
[0006] Extracting key frames from the compressed surveillance video to obtain a key frame sequence corresponding to the compressed surveillance video, and performing target detection on each key frame in the key frame sequence according to a pre-trained image recognition model to obtain an abnormal key frame sequence;
[0007] Locating the abnormal monitoring video segment according to the abnormal key frame sequence and the compressed monitoring video;
[0008] Extract the appearance of the abnormal monitoring video clip according to a pre-trained appearance extraction model to obtain the spatial appearance features corresponding to the abnormal monitoring video clip, and perform spatio-temporal extraction on the spatial appearance features according to a spatio-temporal attention module to obtain the spatio-temporal features corresponding to the abnormal monitoring video clip;
[0009] Extract the optical flow field of the abnormal monitoring video clip according to the TV-L1 algorithm to obtain the optical flow field corresponding to the abnormal monitoring video clip, and perform motion extraction on the optical flow field according to a pre-trained motion extraction model to obtain the motion features of the abnormal monitoring video clip;
[0010] Perform feature fusion on the spatio-temporal features and the motion features according to a gated fusion mechanism to obtain fusion features, and perform anomaly detection on the fusion features according to a preset normal mode memory bank to obtain the anomaly score of the abnormal monitoring video clip.
[0011] Preferably, compressing the monitoring video to obtain a compressed monitoring video includes:
[0012] Divide the monitoring video into multiple GOPs, where the first frame in the GOP is a key frame and the remaining frames in the GOP are non-key frames;
[0013] Perform secondary reconstruction on the non-key frames according to the prior of the non-key frame self-measurement value to obtain the residuals of the non-key frames;
[0014] Obtain output features through denoising based on the residuals of the non-key frames;
[0015] Perform hierarchical downsampling on the reference frame features and the non-key frame features to obtain a multi-scale feature space, fuse the aligned features of the previous level according to the inner product attention in the time domain space, and adaptively refine the non-key frame features using dynamic domain prior knowledge to obtain refined non-key frame features;
[0016] Reconstruct the refined non-key frame features to obtain compressed non-key frames, and combine the compressed non-key frames and key frames to obtain a compressed monitoring video.
[0017] Preferably, the expression of the reconstruction process is as follows: ; where represents the secondary reconstruction of the non-key frame, represents the primary reconstruction of the non-key frame, represents the non-key frame measurement matrix, represents the reference frame measurement matrix, represents the reference frame measurement value, represents the non-key frame measurement value, represents at Extract the static domain residual under the guidance of context information;
[0018] The expression of the output feature is as follows: ; Where represents the th extracted reconstruction feature, represents the denoising network, and represent the measurement operators, and represent the transpose of the measurement operator, represents the feature transformation, represents the inverse feature transformation, and are obtained by repeatedly expanding the measurements of the key frame and non-key frames, represents the th extraction of the static domain residual under the guidance of context information, represents the th extraction of the denoising function implemented by the convolutional network.
[0019] Preferably, the expression of the non-key frame refinement feature is as follows: ; where represents the non-key frame refinement feature, represents being composed of two residual blocks to remove the noise introduced by the process, represents the non-key frame feature, represents being composed of 3 convolutional layers with ReLU, represents the upsampling result of the aligned feature, and represent being implemented by 1×1 convolution, represents the Hadamard product operation.
[0020] Preferably, the image recognition model uses a convolutional neural network, the appearance extraction model uses 3D-ResNet, and the motion extraction model uses a convolutional neural network.
[0021] Preferably, the expression of the spatio-temporal attention module is as follows: ; where represents the spatio-temporal attention weight, represents the activation function, , , represent the query, key, and value matrices, represents the key vector dimension for scaling the dot product, represents the spatial appearance feature.
[0022] Preferably, the expression of the fusion feature is as follows: ; where represents the fusion feature, represents the Sigmoid function, , represent learnable parameters, represents the motion feature.
[0023] Preferably, the calculation formula for the anomaly score of the abnormal monitoring video segment is as follows: ; where represents the anomaly score of the abnormal monitoring video segment, represents the total number of video frames in the abnormal monitoring video segment, represents the normal mode memory bank, represents the similarity smoothing factor.
[0024] The technical solution adopted to solve the above technical problems is: a video anomaly recognition and processing system based on artificial intelligence, which is applicable to the above-mentioned video anomaly recognition and processing method based on artificial intelligence, including:
[0025] A video compression unit, which is used to obtain the monitoring video in the target area and compress the monitoring video to obtain a compressed monitoring video;
[0026] An image recognition unit, which is used to extract key frames from the compressed monitoring video to obtain a key frame sequence corresponding to the compressed monitoring video, and perform target detection on each key frame in the key frame sequence according to a pre-trained image recognition model to obtain an abnormal key frame sequence;
[0027] A video positioning unit, which is used to locate the abnormal monitoring video segment according to the abnormal key frame sequence and the compressed monitoring video;
[0028] A spatio-temporal extraction unit, which is used to perform appearance extraction on the abnormal monitoring video segment according to a pre-trained appearance extraction model to obtain the spatial appearance feature corresponding to the abnormal monitoring video segment, and perform spatio-temporal extraction on the spatial appearance feature according to a spatio-temporal attention module to obtain the spatio-temporal feature corresponding to the abnormal monitoring video segment;
[0029] A motion extraction unit, which is used to perform optical flow field extraction on the abnormal monitoring video segment according to the TV-L1 algorithm to obtain the optical flow field corresponding to the abnormal monitoring video segment, and perform motion extraction on the optical flow field according to a pre-trained motion extraction model to obtain the motion feature of the abnormal monitoring video segment;
[0030] Anomaly scoring unit, which is used to perform feature fusion on the spatio-temporal feature and the motion feature according to the gating fusion mechanism to obtain a fusion feature, and perform anomaly detection on the fusion feature according to a preset normal mode memory bank to obtain the anomaly score of the anomaly monitoring video segment.
[0031] The beneficial effects of the present invention are as follows: (1) By compressing the monitoring video, the present invention can greatly reduce the storage and transmission burden of video data. The compressed video not only saves storage space but also speeds up the subsequent processing steps, which is particularly important in real-time or large-scale processing scenarios of monitoring videos. And through key frame extraction, the computational complexity of video analysis can be effectively reduced because key frames usually contain the main information of the video. By extracting key frames and performing object detection, this method can identify abnormal scenes in the shortest time without processing the entire video stream, reducing the waste of computing resources; (2) By using an appearance extraction model, the present invention can extract key spatial features from the anomaly monitoring video segment. These features help the system identify and understand the spatial structure of abnormal events. Combining with the spatio-temporal attention module, it can not only analyze spatial information but also enhance the model's understanding of temporal information, and can consider the appearance changes and motion changes of objects in both the spatial and temporal dimensions. By using the TV-L1 algorithm to extract the optical flow field, the motion pattern of objects in the video can be accurately captured, and further the motion feature can be obtained through the motion extraction model, so as to be able to identify the motion changes in abnormal events, whether it is the sudden movement of an object, the change in speed or the dynamic performance of other abnormal behaviors; (3) By fusing the spatio-temporal feature and the motion feature, the present invention makes full use of the complementarity of the two to improve the detection accuracy of the model. The spatio-temporal feature mainly describes the changes in the scene, while the motion feature emphasizes the dynamic behavior of the object. By fusing these two features, abnormal events can be identified more accurately. And for anomaly detection, by using a preset normal mode memory bank to compare the fusion feature, it can achieve intelligent judgment based on historical data, quickly and accurately locate abnormal behaviors, and improve the reliability of anomaly detection. Description of the Drawings
[0032] Figure 1 It is a schematic flowchart of the steps of the overall method in an embodiment proposed by the present invention;
[0033] Figure 2 It is a schematic system architecture diagram of the overall system in an embodiment proposed by the present invention.
[0034] Reference numerals: 1, video compression unit; 2, image recognition unit; 3, video positioning unit; 4, spatio-temporal extraction unit; 5, motion extraction unit; 6, anomaly scoring unit. Detailed Embodiments
[0035] Embodiment 1, as Figure 1 shown, the video anomaly recognition and processing method based on artificial intelligence proposed by the present invention includes:
[0036] S1. Obtain the surveillance video within the target area, and compress the surveillance video to obtain a compressed surveillance video;
[0037] S2. Extract key frames from the compressed surveillance video to obtain a key frame sequence corresponding to the compressed surveillance video, and perform object detection on each key frame in the key frame sequence according to a pre-trained image recognition model to obtain an abnormal key frame sequence;
[0038] S3. Locate the abnormal surveillance video segment according to the abnormal key frame sequence and the compressed surveillance video;
[0039] S4. Extract the spatial appearance features corresponding to the abnormal surveillance video segment according to a pre-trained appearance extraction model, and perform spatio-temporal extraction on the spatial appearance features according to a spatio-temporal attention module to obtain the spatio-temporal features corresponding to the abnormal surveillance video segment;
[0040] S5. Extract the optical flow field of the abnormal surveillance video segment according to the TV-L1 algorithm to obtain the optical flow field corresponding to the abnormal surveillance video segment, and perform motion extraction on the optical flow field according to a pre-trained motion extraction model to obtain the motion features of the abnormal surveillance video segment;
[0041] S6. Perform feature fusion on the spatio-temporal features and the motion features according to a gated fusion mechanism to obtain fusion features, and perform anomaly detection on the fusion features according to a preset normal mode memory bank to obtain the anomaly score of the abnormal surveillance video segment.
[0042] In the present invention, monitoring video compression is to compress the original video, usually to reduce the costs of storage and processing. In video compression, redundant data parts are removed, making the video file smaller. The compressed monitoring video retains sufficient information for subsequent analysis; key frames refer to the frames in the video that represent changes in the video content. These frames usually capture important information in the video or the start and end of events. By selecting key frames, the computational amount for processing the video can be reduced, and the main temporal information can be retained; key frame extraction is to extract the most representative frames from the video, usually achieved through image content change detection, clustering methods, or content changes based on the video stream; object detection refers to the process of identifying and locating specific objects from images or video frames. For example, people, vehicles, or other objects of interest in the video frames can be detected. In video anomaly detection, object detection is usually used to identify unusual events or objects in the scene; appearance features refer to the visual features extracted from images to describe target objects, such as color, texture, shape, edges, etc. Appearance extraction is a key step for surface description of objects, used to identify and distinguish different objects or scenes; appearance extraction model refers to a deep learning or traditional machine learning model used to extract appearance features from images, usually trained based on convolutional neural network (CNN); spatio-temporal attention refers to that in the process of processing videos or consecutive frames, the spatio-temporal attention mechanism focuses on information at certain temporal and spatial positions in the video. When extracting spatio-temporal features, the model not only needs to consider the spatial features of the image (such as objects in the image), but also the temporal changes (i.e., motion changes between frames). The spatio-temporal attention module helps to extract important spatio-temporal information by adaptively weighting temporal and spatial features; spatio-temporal features are features that contain both spatial information (static features of the image, such as the appearance of objects) and temporal information (changes or motions of objects in different time frames). In anomaly video detection, spatio-temporal features can help distinguish normal behaviors from abnormal behaviors; the TV-L1 algorithm is an optimized algorithm for optical flow estimation. Optical flow refers to the motion of pixel points between consecutive video frames. The TV-L1 algorithm combines total variation (TV) regularization and L1 norm optimization, and can better handle noise and boundary problems in motion estimation. In video anomaly detection, the TV-L1 algorithm is used to extract the optical flow field in video segments, thereby analyzing the motion patterns of targets; the normal mode memory bank is a database storing normal behavior patterns. When detecting anomalies, the system will compare the currently extracted features with the normal mode. If the difference from the normal mode is large, it is determined to be an anomaly.
[0043] Embodiment 2. The video anomaly recognition and processing method based on artificial intelligence proposed by the present invention. Compared with Embodiment 1, this embodiment further includes: compressing the monitoring video to obtain a compressed monitoring video, including:
[0044] A1. Divide the surveillance video into multiple GOPs, where the first frame in the GOP is a key frame and the remaining frames in the GOP are non-key frames;
[0045] A2. Perform secondary reconstruction on the non-key frames according to the prior of the non-key frame's own measurement values to obtain the residuals of the non-key frames;
[0046] A3. Obtain the output features by denoising the residuals of the non-key frames;
[0047] A4. Perform hierarchical downsampling on the reference frame features and the non-key frame features to obtain a multi-scale feature space. Fuse the aligned features of the previous level according to the inner product attention in the time domain space, and adaptively refine the non-key frame features using the dynamic domain prior knowledge to obtain the refined non-key frame features;
[0048] A5. Reconstruct the refined non-key frame features to obtain the compressed non-key frames, and combine the compressed non-key frames and the key frames to obtain the compressed surveillance video.
[0049] In this embodiment, the terms in GOP video coding refer to a group of image frames. Usually, it consists of a key frame (I frame) and several non-key frames (P frames and B frames). The first frame within the GOP is the key frame, and the other frames are encoded based on the differences from the previous frame, which are usually called non-key frames; secondary reconstruction refers to the process of reconstructing non-key frames through some prior information, which is usually to restore the details of non-key frames and reduce the loss caused by compression; the dynamic domain refers to the dynamic information area that changes over time in video processing. In time-series data, the changes of objects or scenes occur over time, and the dynamic domain emphasizes the dynamic change characteristics at different moments; prior knowledge refers to the knowledge based on experience or existing knowledge, which helps the model make better predictions during training and inference. In video analysis, the dynamic domain prior knowledge may include information such as the motion patterns and behavior patterns of objects, which helps to analyze and understand the video more accurately; feature reconstruction refers to the process of restoring the extracted feature information to the original image or video. In this process, the model attempts to restore the visual effect of the original frame through features. The reconstruction process is usually used for video compression and denoising tasks to minimize the information lost during compression.
[0050] In an optional embodiment, the expression of the reconstruction process is as follows: ; where represents the secondary reconstruction of the non-key frame, represents the primary reconstruction of the non-key frame, represents the non-key frame measurement matrix, represents the reference frame measurement matrix, represents the reference frame measurement value, represents the non-key frame measurement value, It means to extract the static domain residual under the guidance of context information;
[0051] The expression of the output feature is as follows: ; where means the th reconstructed feature extracted, means the denoising network, and mean the measurement operator, and mean the transpose of the measurement operator, means the feature transformation, means the inverse feature transformation, and mean being repeatedly expanded from the measurement values of key frames and non-key frames, means the th extraction of the static domain residual under the guidance of context information, means the th extraction of the denoising function implemented by the convolutional network.
[0052] It should be noted that one-time reconstruction usually refers to the preliminary process of restoring non-key frames from compressed data (such as video streams or frame encodings). Different from two-time reconstruction, one-time reconstruction is often based on the original information stored during encoding, aiming to restore a reasonable image or video frame; the non-key frame measurement matrix is a matrix that maps the features of non-key frames to a low-dimensional space, and this matrix is used to compress or represent non-key frames. During the process of compressing videos, the measurement matrix is used to extract or measure certain features or information in the frames; the reference frame measurement matrix is similar to the non-key frame measurement matrix but is used for reference frames. A reference frame usually refers to a frame that has been fully reconstructed in video coding or a frame used for predicting non-key frames. The measurement matrix of the reference frame is used to extract the features in these frames and compare and align them with non-key frames; the reference frame measurement values are the values extracted through the measurement matrix of the reference frame, and they represent the compressed data or feature information of the reference frame, which is very important in video coding and reconstruction and is used to extract information from the reference frame to assist in restoring non-key frames; context information refers to using prior knowledge or the background information of the current video frame to assist the processing process; the measurement operator is a linear transformation used to map the information in an image or video to a new space or dimension. During the compression process, the measurement operator compresses the data into a low-dimensional representation, and during the restoration process, the measurement operator can be used to restore the original information; the feature transformation refers to performing a certain transformation on the input features (such as the features of video frames) to make them adapt to specific processing or analysis requirements.
[0053] In an alternative embodiment, the expression of the non-critical frame refinement feature is as follows: ; where represents the non-critical frame refinement feature, represents being composed of two residual blocks for removing the noise introduced during the process, represents the non-critical frame feature, represents being composed of 3 convolutional layers with ReLU, represents the upsampling result of the aligned feature, and represents being implemented by 1×1 convolution, represents the Hadamard product operation.
[0054] It should be noted that upsampling is an operation used to expand a low-resolution feature map to a higher resolution. Common upsampling methods include transposed convolution or interpolation methods. Upsampling operations are often used in image generation or feature recovery processes with the aim of restoring low-resolution features to a higher resolution state.
[0055] In an alternative embodiment, the image recognition model uses a convolutional neural network, the appearance extraction model uses 3D-ResNet, and the motion extraction model uses a convolutional neural network.
[0056] It should be noted that 3D-ResNet is an extension of ResNet (residual network) designed for videos and 3D data. Traditional ResNet is used for image recognition tasks. By introducing skip connections, the problem of vanishing gradients in deep networks is solved, enabling the network to be deeper and easier to train; the motion extraction model refers to a model that extracts motion information from videos or continuous image sequences, and the goal of motion extraction is to analyze and capture the motion trajectories or changes of objects in images or videos.
[0057] In an alternative embodiment, the expression of the spatio-temporal attention module is as follows: ; where represents the spatio-temporal attention weight, represents the activation function, 、 、 represent the query, key, and value matrices, represents the key vector dimension for scaling the dot product, represents the spatial appearance feature.
[0058] In an alternative embodiment, the expression of the fused feature is as follows: ; where represents the fused feature, denotes the Sigmoid function, , denotes the learnable parameters, denotes the motion features.
[0059] In an optional embodiment, the formula for calculating the anomaly score of the anomaly-monitoring video segment is as follows: ; where, denotes the anomaly score of the anomaly-monitoring video segment, denotes the total number of video frames in the anomaly-monitoring video segment, denotes the normal mode memory bank, denotes the similarity smoothing factor.
[0060] Embodiment 3, as Figure 2 shown, the video anomaly recognition and processing system based on artificial intelligence proposed by the present invention, which is applicable to the video anomaly recognition and processing method based on artificial intelligence, includes:
[0061] Video compression unit 1, which is used to obtain the monitoring video within the target area and compress the monitoring video to obtain the compressed monitoring video;
[0062] Image recognition unit 2, which is used to extract key frames from the compressed monitoring video to obtain the key frame sequence corresponding to the compressed monitoring video, and perform object detection on each key frame in the key frame sequence according to the pre-trained image recognition model to obtain the anomaly key frame sequence;
[0063] Video positioning unit 3, which is used to locate the anomaly-monitoring video segment according to the anomaly key frame sequence and the compressed monitoring video;
[0064] Spatio-temporal extraction unit 4, which is used to perform appearance extraction on the anomaly-monitoring video segment according to the pre-trained appearance extraction model to obtain the spatial appearance features corresponding to the anomaly-monitoring video segment, and perform spatio-temporal extraction on the spatial appearance features according to the spatio-temporal attention module to obtain the spatio-temporal features corresponding to the anomaly-monitoring video segment;
[0065] Motion extraction unit 5, which is used to perform optical flow field extraction on the anomaly-monitoring video segment according to the TV-L1 algorithm to obtain the optical flow field corresponding to the anomaly-monitoring video segment, and perform motion extraction on the optical flow field according to the pre-trained motion extraction model to obtain the motion features of the anomaly-monitoring video segment;
[0066] Anomaly scoring unit 6, which is used to perform feature fusion on the spatio-temporal features and the motion features according to the gated fusion mechanism to obtain the fused features, and perform anomaly detection on the fused features according to the preset normal mode memory bank to obtain the anomaly score of the anomaly-monitoring video segment.
[0067] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited thereto, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those skilled in the art to which the present invention pertains.
Claims
1. A video anomaly recognition and processing method based on artificial intelligence, characterized in that: include: Acquire a surveillance video in a target area, and compress the surveillance video to obtain a compressed surveillance video; Extracting key frames from the compressed surveillance video to obtain a key frame sequence corresponding to the compressed surveillance video, and performing target detection on each key frame in the key frame sequence according to a pre-trained image recognition model to obtain an abnormal key frame sequence; Locating the abnormal monitoring video segment according to the abnormal key frame sequence and the compressed monitoring video; Performing appearance extraction on the abnormal monitoring video clip according to the pre-trained appearance extraction model to obtain the spatial appearance features corresponding to the abnormal monitoring video clip, and performing spatiotemporal extraction on the spatial appearance features according to the spatiotemporal attention module to obtain the spatiotemporal features corresponding to the abnormal monitoring video clip; Performing optical flow field extraction on the abnormal monitoring video clip according to the TV-L1 algorithm to obtain the optical flow field corresponding to the abnormal monitoring video clip, and performing motion extraction on the optical flow field according to a pre-trained motion extraction model to obtain the motion features of the abnormal monitoring video clip; The spatiotemporal features and the motion features are fused according to a gated fusion mechanism to obtain fused features, and anomaly detection is performed on the fused features according to a preset normal pattern memory library to obtain an anomaly score of the abnormal monitoring video clip.
2. The method for video anomaly recognition and processing based on artificial intelligence according to claim 1 is characterized in that: Compressing the surveillance video to obtain a compressed surveillance video includes: dividing the surveillance video into a plurality of GOPs, wherein the first frame in the GOP is a key frame, and the remaining frames in the GOP are non-key frames; Reconstructing the non-key frame a second time according to the non-key frame's own measurement value a priori to obtain the residual of the non-key frame; Denoising the residual of the non-key frame to obtain output features; The reference frame features and non-key frame features are downsampled step by step to obtain a multi-scale feature space. The alignment features of the previous level are fused according to the temporal spatial inner product attention. The non-key frame features are adaptively refined using the dynamic domain prior knowledge to obtain non-key frame refined features. The non-key frame refinement features are reconstructed to obtain compressed non-key frames, and the compressed non-key frames are combined with key frames to obtain compressed surveillance videos.
3. The method for video anomaly recognition and processing based on artificial intelligence according to claim 2 is characterized in that: The expression of the reconstruction process is as follows: ; in, represents the secondary reconstruction of non-keyframes, represents a reconstruction of a non-keyframe, represents the non-keyframe measurement matrix, represents the reference frame measurement matrix, represents the reference frame measurement, represents non-keyframe measurements, Indicated in The static domain residual is extracted under the guidance of context information; the expression of the output feature is as follows: ;in, Indicates The reconstructed features extracted are represents the denoising network, and represents the measurement operator, and represents the transpose of the measurement operator, represents feature transformation, represents the inverse feature transform, and It is obtained by repeatedly expanding the key frame and non-key frame measurements. Indicates Extracted in Extract static domain residuals guided by context information, Indicates Extract the denoising function implemented by the convolutional network again.
4. The method for video anomaly recognition and processing based on artificial intelligence according to claim 3 is characterized in that: The expression of the non-keyframe refinement feature is as follows: ;in, represents the non-keyframe refinement feature, It is composed of two residual blocks to remove the noise introduced by the process. represents non-keyframe features, It means that it consists of 3 convolutional layers with ReLU. represents the upsampling result of the aligned features, and It means that it is realized by 1×1 convolution. Represents the Hadamard product operation.
5. The method for video anomaly recognition and processing based on artificial intelligence according to claim 1 is characterized in that: The image recognition model adopts a convolutional neural network, the appearance extraction model adopts 3D-ResNet, and the motion extraction model adopts a convolutional neural network.
6. The method for video anomaly recognition and processing based on artificial intelligence according to claim 1 is characterized in that: The expression of the spatiotemporal attention module is as follows: ;in, represents the spatiotemporal attention weight, represents the activation function, , , represents the query, key, and value matrix, represents the key vector dimension, used to scale the dot product, Represents the appearance characteristics of the space.
7. The method for video anomaly recognition and processing based on artificial intelligence according to claim 1 is characterized in that: The expression of the fusion feature is as follows: ;in, represents the fusion feature, represents the Sigmoid function, , represents the learnable parameters, Indicates motion characteristics.
8. The method for video anomaly recognition and processing based on artificial intelligence according to claim 1 is characterized in that: The calculation formula of the abnormal score of the abnormal monitoring video clip is as follows: ;in, represents the anomaly score of the abnormal surveillance video clip, Indicates the total number of video frames in the abnormal monitoring video clip. Indicates the normal mode memory bank, Represents the similarity smoothing factor.
9. An artificial intelligence-based video anomaly recognition and processing system, which is applicable to the artificial intelligence-based video anomaly recognition and processing method according to any one of claims 1 to 8, characterized in that: include: A video compression unit (1), the video compression unit (1) is used to obtain a surveillance video in a target area, and compress the surveillance video to obtain a compressed surveillance video; An image recognition unit (2), the image recognition unit (2) being used to extract key frames from the compressed surveillance video to obtain a key frame sequence corresponding to the compressed surveillance video, and to perform target detection on each key frame in the key frame sequence according to a pre-trained image recognition model to obtain an abnormal key frame sequence; A video positioning unit (3), the video positioning unit (3) being used to locate the abnormal monitoring video segment according to the abnormal key frame sequence and the compressed monitoring video; A spatiotemporal extraction unit (4), the spatiotemporal extraction unit (4) being used to perform appearance extraction on the abnormal monitoring video segment according to a pre-trained appearance extraction model to obtain a spatial appearance feature corresponding to the abnormal monitoring video segment, and to perform spatiotemporal extraction on the spatial appearance feature according to a spatiotemporal attention module to obtain a spatiotemporal feature corresponding to the abnormal monitoring video segment; A motion extraction unit (5), the motion extraction unit (5) being used to perform optical flow field extraction on the abnormal monitoring video segment according to the TV-L1 algorithm to obtain an optical flow field corresponding to the abnormal monitoring video segment, and to perform motion extraction on the optical flow field according to a pre-trained motion extraction model to obtain motion features of the abnormal monitoring video segment; An anomaly scoring unit (6) is used to perform feature fusion on the spatiotemporal features and the motion features according to a gated fusion mechanism to obtain a fusion feature, and perform anomaly detection on the fusion feature according to a preset normal pattern memory library to obtain an anomaly score for the abnormal monitoring video clip.
Citation Information
Patent Citations
Video stream acquisition method based on deep learning
CN119299703A
Monitoring video information intelligent processing system based on cloud computing
CN119672613A
Cited By
Video monitoring abnormal behavior identification method and system based on edge AI
CN120877371A