Unmanned aerial vehicle video anomaly detection method in dynamic background

CN121564425BActive Publication Date: 2026-09-15ANHUI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511800360.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-09-15
Estimated Expiration
2045-12-02

AI Technical Summary

Technical Problem

[0004]鉴于以上现有技术的缺陷,本发明提供一种动态背景下的无人机视频异常检测方法,以解决现有技术中无人机拍摄视频检测复杂和准确率低的技术问题

Benefits of technology

(1)实现了复杂动态背景下的有效运动解耦,显著降低了误报率。针对无人机视频因自身飞行运动导致背景不断变化、难以区分背景流与前景运动的核心技术难题,本发明通过引入频率解耦时空关联模块,创新性地在频率域对复杂的混合运动进行分析与解耦,从而显著降低了由动态背景引起的误报。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121564425B_ABST
    Figure CN121564425B_ABST
Patent Text Reader

Abstract

The application provides a kind of unmanned aerial vehicle video anomaly detection method in dynamic background, the method comprises: obtaining the first video frame sequence to be detected;Any frame in the first video frame sequence is regarded as the current frame, and a predetermined number of video frames before the current frame is extracted to obtain the second video frame sequence corresponding to the current frame;The second video frame sequence corresponding to the current frame is input into the trained video frame prediction model to obtain the predicted frame;According to all frames in the first video frame sequence and the corresponding predicted frame, the score of each frame in the first video frame sequence is obtained;According to the score of each frame in the first video frame sequence and the threshold, the anomaly detection result of each frame in the first video frame sequence is obtained. The current frame is predicted using the historical frame to model the space-time dynamics, and the anomaly is accurately captured by the difference between the predicted frame and the real frame, effectively solving the problem of dynamic background and multi-source motion coupling. Without anomaly annotation, high-precision and strong-generalization unmanned aerial vehicle video anomaly detection can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a method for detecting anomalies in drone videos under dynamic backgrounds. Background Technology

[0002] With the rapid development of drone technology, drone-based video surveillance is playing an increasingly important role in urban security, traffic management, and environmental monitoring. Drones can provide flexible perspectives and acquire large-scale, high-precision video data, demonstrating great potential in tasks such as emergency response, traffic flow monitoring, and natural disaster assessment. However, unlike videos captured by fixed ground cameras, drone video has unique and challenging characteristics, particularly the problem of multi-source motion coupling. The drone's own motion introduces global background motion, which intertwines with the local motion of foreground objects, causing a deviation between the observed object trajectory and the actual trajectory. This multi-source motion coupling problem makes anomaly detection in drone video extremely complex.

[0003] Existing video anomaly detection methods are primarily designed for videos captured by fixed ground cameras. These methods mostly employ prediction or reconstruction paradigms, learning normal patterns through spatiotemporal modeling and identifying regions with high reconstruction or prediction errors as anomalies. However, these methods have two main limitations when processing drone videos: First, existing methods lack decoupling capabilities, failing to effectively separate background flow caused by the drone's own motion from the motion of real foreground objects. This deficiency may lead to misjudging drastic visual changes generated by normal drone flight as anomalies, while masking genuine anomalous events (such as sudden traffic accidents or intruding unrelated targets) by the dynamically changing background. Existing technologies often fail to consider the complex interaction between background and foreground, making them unsuitable for handling frequently changing background environments in practical applications such as dynamic cruise. Second, drone videos typically exhibit continuous, smooth global motion, especially during tasks such as cruise surveillance and long-duration flights, where the overall background of the video shows continuous, gradual changes. Therefore, to effectively capture these global dynamics, the model needs a relatively long time window to track changes in background motion. However, local objects (such as moving vehicles and pedestrians) often exhibit irregular, rapidly changing motion, requiring a shorter time window to capture these fine-grained changes. Understanding the motion dynamics of drone videos requires modeling across multiple time scales to comprehensively capture both global and local changes. However, existing methods typically use fixed time scales, neglecting the joint modeling of temporal continuity and local spatial correlations across different time scales, which significantly reduces their ability to detect anomalous events. Summary of the Invention

[0004] In view of the above-mentioned deficiencies of the prior art, the present invention provides a method for detecting anomalies in drone videos under dynamic backgrounds, so as to solve the technical problems of complexity and low accuracy in drone video detection in the prior art.

[0005] To achieve the above and other related objectives, the present invention provides a method for detecting anomalies in UAV videos under dynamic backgrounds, comprising: acquiring a first video frame sequence to be detected; taking any frame in the first video frame sequence as the current frame, extracting a preset number of video frames preceding the current frame to obtain a second video frame sequence corresponding to the current frame; inputting the second video frame sequence corresponding to the current frame into a trained video frame prediction model to obtain a predicted frame corresponding to the current frame; obtaining a score for each frame in the first video frame sequence based on all frames in the first video frame sequence and their corresponding predicted frames; and obtaining an anomaly detection result for each frame in the first video frame sequence based on the score of each frame in the first video frame sequence and a threshold.

[0006] In one embodiment of the present invention, the video frame prediction model includes an encoder, a feature enhancement module, and a decoder; inputting the second video frame sequence corresponding to the current frame into the trained video frame prediction model to obtain the predicted frame corresponding to the current frame includes: preprocessing the second video frame sequence corresponding to the current frame; processing the preprocessed second video frame sequence using the encoder to obtain hierarchical features; processing the hierarchical features using the feature enhancement module to obtain enhanced features; and processing the enhanced features using the decoder to obtain the predicted frame.

[0007] In one embodiment of the present invention, the preprocessing includes size adjustment and normalization; the encoder includes an M-stage pyramid visual transformer, each stage of which outputs a hierarchical feature of a scale.

[0008] In one embodiment of the present invention, the feature enhancement module includes a frequency domain decoupled spatiotemporal correlation unit and a multi-scale spatiotemporal Mamba unit; the feature enhancement module processes the hierarchical features to obtain enhanced features, including: processing the hierarchical features using the frequency domain decoupled spatiotemporal correlation unit to obtain a first feature; processing the hierarchical features using the multi-scale spatiotemporal Mamba unit to obtain a second feature; and concatenating the first feature and the second feature in the channel dimension and projecting them back to the original dimension to obtain the enhanced features.

[0009] In one embodiment of the present invention, the hierarchical features are processed using the frequency domain decoupling spatiotemporal correlation unit to obtain a first feature, including: performing a one-dimensional fast Fourier transform on the hierarchical features along the time dimension to obtain complex-valued frequency domain features and frequency amplitude; obtaining frequency correlation weights based on the frequency amplitudes and a preset normalized frequency vector; weighting the complex-valued frequency domain features using the frequency correlation weights and then performing an inverse fast Fourier transform to obtain frequency-weighted features; performing a two-dimensional fast Fourier transform on the frequency-weighted features along the spatiotemporal dimension to obtain a power spectral density; performing a two-dimensional inverse Fourier transform on the power spectral density to obtain a spatiotemporal autocorrelation matrix; and using the spatiotemporal autocorrelation matrix as attention weights to process the frequency-weighted features to obtain the first feature.

[0010] In one embodiment of the present invention, the multi-scale spatiotemporal Mamba unit can be expressed by the following formula: , In the formula, For an η-fold expansion and recombination transformation, This is an inverse recombination transformation with an η-fold expansion, where η ∈ {1, 2, ..., n}, and n is the maximum expansion factor; f η-1 The output of the η-fold dilation inverse recombination transform, f0 is the hierarchical feature, and STMamba is the spatiotemporal Mamba block. This is the second feature corresponding to the hierarchical feature.

[0011] In one embodiment of the present invention, the processing procedure of the spatiotemporal Mamba block is as follows: The input features of the spatiotemporal Mamba block are sequentially normalized and linearly projected to obtain core features and gated features; each pixel position of the core features is flattened in row-major order and column-major order to obtain flattened vectors, and the flattened vectors are concatenated in chronological order to obtain a first sequence and a second sequence; each frame of the core features is divided into non-overlapping blocks, and for each block position, flattened in row-major order and column-major order to obtain flattened vectors, and the flattened vectors are concatenated in chronological order to obtain a third sequence and a fourth sequence; the first sequence, the second sequence, the third sequence, and the fourth sequence are scanned forward and backward to obtain eight scan sequences; each scan sequence is sequentially subjected to causal one-dimensional convolution, state-space model (SSM) processing, and gating with the gated features; the gating results of all scan sequences are summed, and then linear mapping and residual connection are performed to obtain the output features of the spatiotemporal Mamba block.

[0012] In one embodiment of the present invention, the decoder can be expressed by the formula: , , In the formula, f i For the i-th hierarchical feature, F i f i The corresponding enhancement features, δ represents the intermediate features of the decoder, P represents the predicted frame, Up2 represents the ConvTranspose2d operation, BN represents batch normalization, δ represents the ReLU activation function, and Up represents the upsampling operation.

[0013] In one embodiment of the present invention, the score of each frame in the first video frame sequence is obtained based on all frames in the first video frame sequence and their corresponding predicted frames, including: obtaining the peak signal-to-noise ratio (PSNR) of each frame in the first video frame sequence based on all frames in the first video frame sequence and their corresponding predicted frames; obtaining the maximum and minimum values ​​of the PSNR based on the PSNR of all frames in the first video frame sequence; and obtaining the score of each frame in the first video frame sequence based on the PSNR of each frame in the first video frame sequence and the maximum and minimum values ​​of the PSNR.

[0014] In one embodiment of the present invention, obtaining the anomaly detection result of each frame in the first video frame sequence based on the score and threshold of each frame in the first video frame sequence includes: generating multiple optional thresholds uniformly distributed within the [0,1] interval; for each optional threshold, generating predicted labels based on the anomaly score, and calculating the F1 score of these predicted labels and the true labels; selecting the optional threshold with the largest F1 score as the optimal threshold; and obtaining the anomaly detection result of each frame in the first video frame sequence based on the score of each frame in the first video frame sequence and the optimal threshold.

[0015] The beneficial effects of this invention: The present invention proposes a method for detecting anomalies in UAV videos under dynamic backgrounds, which has the following advantages: (1) Effective motion decoupling under complex dynamic backgrounds is achieved, significantly reducing the false alarm rate. In response to the core technical challenge of UAV video where the background is constantly changing due to its own flight motion and it is difficult to distinguish between background flow and foreground motion, this invention introduces a frequency decoupling spatiotemporal correlation module, which innovatively analyzes and decouples complex mixed motions in the frequency domain, thereby significantly reducing false alarms caused by dynamic backgrounds.

[0016] (2) Possesses powerful multi-scale spatiotemporal modeling capabilities, accurately capturing various abnormal events. This invention combines the multi-scale spatiotemporal Mamba module and utilizes the advantages of the state-space model in long sequence modeling to achieve accurate capture of long-term continuous changes and short-term sudden events in videos, further improving the detection accuracy and robustness of the model in challenging scenarios such as multi-source motion coupling.

[0017] (3) It adopts an unsupervised learning paradigm, which has strong generalization ability and broad application prospects. The present invention adopts an unsupervised learning paradigm with future frame prediction as the core. It does not require manual annotation of massive amounts of data and has good universality and scalability for various dynamic perspectives. It is not only suitable for UAV security monitoring, but can also be directly applied to other video anomaly detection fields that include dynamic perspectives, such as vehicle monitoring and robot inspection. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. The accompanying drawings are incorporated in and constitute a part of this specification, illustrating embodiments consistent with this application, and are used together with the description to explain the principles of this application. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative effort.

[0019] Figure 1 A flowchart of a method for detecting video anomalies in unmanned aerial vehicles (UAVs) according to an embodiment of the present invention; Figure 2 This is an architecture diagram of a video frame prediction model provided in an embodiment of the present invention; Figure 3 A flowchart illustrating the processing of a video frame prediction model provided in an embodiment of the present invention; Figure 4 This is a flowchart of the feature enhancement module provided in an embodiment of the present invention; Figure 5 A flowchart of the processing of a frequency domain decoupling spatiotemporal correlation unit provided in an embodiment of the present invention; Figure 6 This is an architectural diagram of a frequency domain decoupling spatiotemporal correlation unit provided in an embodiment of the present invention; Figure 7 This is an architectural diagram of a multi-scale spatiotemporal Mamba unit provided in an embodiment of the present invention; Figure 8 A flowchart illustrating the processing of a spatiotemporal Mamba block according to an embodiment of the present invention; Figure 9 This is a flowchart of frame score calculation provided in an embodiment of the present invention; Figure 10 A flowchart for frame anomaly detection provided in an embodiment of the present invention Figure 11 This invention provides a visualization of anomaly curves for four test video segments across three datasets, as provided in one embodiment of the invention. Figure 12The present invention provides a visualization of prediction results across three datasets according to an embodiment of the invention. Detailed Implementation

[0020] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. It should be noted that, unless otherwise specified, the following embodiments and features can be combined with each other. In addition to the specific methods, equipment, and materials used in the embodiments, based on the knowledge of the prior art and the description of the present invention by those skilled in the art, any prior art methods, equipment, and materials similar to or equivalent to the methods, equipment, and materials in the embodiments of the present invention can be used to implement the present invention.

[0021] It should be understood that the terminology used in the embodiments of this invention is for describing specific implementations and not for limiting the scope of protection of this invention. Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art.

[0022] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In some embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.

[0023] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functions, and operations that may be implemented in the methods and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0024] Please see Figure 1 , Figure 1 An embodiment of the present invention provides a method for detecting anomalies in UAV video under dynamic background, including steps S101 to S105.

[0025] Step S101: Obtain the first video frame sequence to be detected. This step is the foundation for all subsequent processing. Its core task is to extract a continuous video frame segment from the raw video stream captured by the drone as the input unit to be detected. Specifically, by connecting to the drone's image transmission system or reading a stored video file, sampling is performed at a fixed frame rate to ensure the temporal continuity of the sequence, providing a stable and regular data foundation for subsequent feature extraction and anomaly analysis.

[0026] Step S102: Take any frame in the first video frame sequence as the current frame, and extract a preset number of video frames preceding the current frame to obtain the second video frame sequence corresponding to the current frame. In subsequent steps, the prediction frame corresponding to the current frame will be predicted based on the second video frame sequence; therefore, this step requires extracting a preset number of video frames preceding the current frame. For example... Figure 2 In the middle, T is recorded N Given the current frame, with a preset number of frames T=N-1, then the video frames T1, T2, ..., T... N-1 This constitutes the second video frame sequence.

[0027] Understandably, when any frame in the first video frame sequence is taken as the current frame, if T... N T N+1 The current frame is defined as T1, T2, ..., T..., which are preceded by a preset number of video frames. N-1 If any frame in the sequence is taken as the current frame, and the number of frames preceding it in the video is less than T, there are several ways to handle this: for example, several frames preceding the first video frame sequence can be acquired together to ensure that when T1 is taken as the current frame, the preceding T frames can also be extracted; or, in subsequent processing, T1~T can be skipped. N-1 These frames are not used as the current frame (i.e., no anomaly detection is performed on them).

[0028] Step S103: Input the second video frame sequence corresponding to the current frame into the trained video frame prediction model to obtain the predicted frame corresponding to the current frame. This step is the core of the model's inference process. The second video frame sequence, containing historical information, is input into the trained prediction model. The model generates a reasonable prediction for the current frame by learning the spatiotemporal dynamic features of the normal pattern. This predicted frame represents the "normal" state inferred by the model based on the historical context. Subsequently, the difference between the predicted frame and the real frame will be compared to identify possible anomalies.

[0029] Please see Figure 2 In a specific embodiment of the present invention, the video frame prediction model includes an encoder, a feature enhancement module, and a decoder, wherein the encoder corresponds to... Figure 2 In the encoding stage, the feature enhancement module corresponds to Figure 2The part enclosed in the middle dashed box in the image corresponds to the decoder. Figure 2 The decoding stage in the process.

[0030] Please see Figure 3 In a specific embodiment of the present invention, when the video frame prediction model adopts Figure 2 When the architecture shown is used, step S103 includes steps S301 to S304.

[0031] Step S301: Preprocess the second video frame sequence corresponding to the current frame. Preprocessing generally includes operations such as resizing and normalization. In this step, for example, each frame in the second video frame sequence can be uniformly adjusted to a resolution of 256×256 to adapt to the input of the video frame prediction model. Then, the pixel values ​​are normalized so that their numerical range becomes [-1,1]. The shape of each frame after processing is C×H×W. Taking RGB format as an example, C=3 represents the number of channels, H=256 represents the height of the frame, and W=256 represents the width of the frame.

[0032] Step S302: Use the encoder to process the preprocessed second video frame sequence to obtain hierarchical features.

[0033] In one specific embodiment of the present invention, the encoder includes an M-stage pyramid visual transformer, where each stage of the pyramid visual transformer outputs a hierarchical feature at a specific scale. Taking M=4 as an example, the hierarchical feature obtained after encoder processing can be denoted as... ,in , It refers to the batch size. It is the input frame number. It is the number of feature channels. and These represent the height and width of the feature, respectively. To reduce computational complexity, this invention reduces the number of video feature channels at different levels. The standard value is 384.

[0034] Step S303: Use the feature enhancement module to process the hierarchical features to obtain enhanced features.

[0035] Please see Figure 2In one specific embodiment of the present invention, the feature enhancement module includes a frequency domain decoupled spatiotemporal correlation unit and a multi-scale spatiotemporal Mamba unit (also known as a multi-scale spatiotemporal Mamba unit). Through the synergistic effect of these two units, the hierarchical features extracted by the encoder are deeply enhanced. Specifically, the frequency domain decoupled spatiotemporal correlation unit effectively separates background motion from foreground anomalies and suppresses interference; the multi-scale spatiotemporal Mamba unit jointly models spatiotemporal continuity at multiple scales, capturing long- and short-range dependencies. The features output by both units are fused to form enhanced features with richer information and less noise, laying the foundation for the decoder to generate accurate predictions.

[0036] Understandably, the feature enhancement module processes each hierarchical feature separately to obtain their corresponding enhanced features. For hierarchical features... In other words, the enhanced features obtained after processing are .

[0037] Please see Figure 4 In a specific embodiment of the present invention, when the feature enhancement module adopts... Figure 2 When the architecture shown is used, step S303 includes steps S401 to S403.

[0038] Step S401: Process the hierarchical features using the frequency domain decoupling spatiotemporal correlation unit to obtain the first feature.

[0039] Please see Figure 5 and Figure 6 In a specific embodiment of the present invention, step S401 includes steps S501 to S506.

[0040] Step S501: Perform a one-dimensional Fast Fourier Transform (FFT) along the time dimension on the hierarchical features to obtain complex-valued frequency domain features and frequency amplitudes. In this step, the input features f∈{f1,f2,f3,f4} are denoted as... For each batch b, channel c, and spatial location (h, w) of the input feature f, apply a one-dimensional FFT along the time dimension and calculate the complex-valued frequency domain features according to the following formula. and frequency amplitude : , , In the formula, k = 0, 1, ..., T-1, j is the imaginary unit, and Re() and Im() represent the real part and imaginary part, respectively.

[0041] Step S502: Obtain the frequency-related weights based on the frequency amplitude and the preset normalized frequency vector. Wherein, the normalized frequency vector... Its element l k It can be calculated using the following formula: .

[0042] After obtaining the normalized frequency vector l, the frequency-related weight w k It can be calculated using the following formula: .

[0043] Step S503: Utilize frequency-related weights w k For complex-valued frequency domain features After weighting, an inverse fast Fourier transform is performed to obtain the frequency-weighted feature f'. This step can be expressed by the formula: , , In this way, for each input feature f, a corresponding f' can be obtained.

[0044] Step S504: Perform a two-dimensional fast Fourier transform on the frequency-weighted features along the spatiotemporal dimension to obtain the power spectral density. In this step, the spatial dimension is reshaped as follows: For each batch b and channel c, apply a two-dimensional FFT in the spatiotemporal dimension and calculate the power spectral density. The calculation formula is as follows: , , In the formula, and These represent time and spatial frequency indices, respectively. It is a complex conjugate.

[0045] Step S505: Perform a two-dimensional inverse Fourier transform on the power spectral density to obtain the spatiotemporal autocorrelation matrix. The specific calculation formula is as follows: .

[0046] Step S506: Use the spatiotemporal autocorrelation matrix as attention weights. The frequency-weighted feature f' is processed to obtain the first feature. The specific calculation formula is as follows: , In the formula, ⊙ represents element-wise multiplication.

[0047] In the above steps, the input features of the frequency domain decoupling spatiotemporal correlation unit are: Its output features are corresponding to Since steps S501 to S506 above involve more subscripts, to keep the formula clear and concise, we use the input feature f, frequency-weighted feature f', and first feature... The subscript i is omitted in the text.

[0048] See also Figure 4 Step S402: Process the hierarchical features using multi-scale spatiotemporal Mamba units to obtain the second feature.

[0049] Please see Figure 7 In the upper part of a specific embodiment of the present invention, the multi-scale spatiotemporal Mamba unit can be expressed by the formula: , In the formula, For an η-fold expansion and recombination transformation, This is an inverse recombination transformation with an η-fold expansion, where η ∈ {1, 2, ..., n}, and n is the maximum expansion factor; f η-1 The output of the η-fold dilation inverse recombination transformation is given, f0 is the hierarchical feature f, and STMamba is the spatiotemporal Mamba block (also known as the spatiotemporal Mamba block). This is the second feature corresponding to the hierarchical feature.

[0050] In this embodiment, the hierarchical feature f is also... Therefore, the second feature after processing by the multi-scale spatiotemporal Mamba unit corresponds to... In this embodiment, the subscript i is also hidden. The process in this embodiment can be expanded as follows (note that the subscript of f in the following formula does not represent the i-th hierarchical feature, but corresponds to the output of different multiplier dilation inverse recombination transformations): , , ... , Finally, for f1~f n After summing, the data is processed by a spatiotemporal Mamba block to obtain the final output of the multi-scale spatiotemporal Mamba unit.

[0051] In this embodiment, the η-fold dilation and recombination transformation The processing procedure is as follows: , In the formula, the dilation rate η is used to extract frame sequences at different time scales. Through η-fold dilation and recombination transformation, the input features can be recombined into ηB samples, each sample having a sequence length of T / η. These ηB samples are then concatenated in the channel dimension and processed through a spatiotemporal Mamba block, followed by η-fold dilation and inverse recombination transformation. Figure 7 (Not shown in the image) is mapped to the original feature dimension.

[0052] Please see Figure 8 In a specific embodiment of the present invention, the processing procedure of the spatiotemporal Mamba block includes steps S801 to S806, and its architecture diagram is as follows. Figure 7 The lower half is shown.

[0053] Step S801: Input features of the spatiotemporal Mamba block Normalization and linear projection operations are performed sequentially to obtain the core feature X and the gated feature Z. In this step, the input feature f is first processed... in Perform normalization (i.e.) Figure 7 (Regularization in the model), and then generate core features X and gated features Z through two linear projections respectively.

[0054] Step S802: Flatten each pixel position of the core feature X according to row-major order and column-major order to obtain a flattened vector. Then, concatenate the flattened vectors in chronological order to obtain the first sequence and the second sequence. In this step, the spatial dimension of each frame is flattened according to row-major order and column-major order, and concatenated along the time axis to form the first sequence S1 and the second sequence S2. .

[0055] Step S803: Divide each frame of the core feature X into non-overlapping blocks, for example, into P×P blocks. For each block position, flatten it according to row-major order and column-major order to obtain a flattened vector. Then, concatenate the flattened vectors in chronological order to obtain the third and fourth sequences. In this step, by extracting the temporal sequence of each block position in T frames, The block sequences are concatenated in row-major order and column-major order, ultimately yielding the third sequence S3 and the fourth sequence S4, respectively. .

[0056] Step S804: Perform forward and backward scans on the first, second, third, and fourth sequences respectively to obtain 8 scan sequences, which can be denoted as... .

[0057] Step S805: Perform causal one-dimensional convolution, state space model SSM processing, and gating operation with gating features on each scan sequence in sequence.

[0058] Step S806: After summing the gating results of all scan sequences, perform linear mapping and residual connection to obtain the output feature f of the spatiotemporal Mamba block. out .

[0059] The processes of steps S805 and S806 can be expressed by the following formula: , In the formula, Conv1d is a causal one-dimensional convolution, which acts like a short-time focusing lens to perform fast local preprocessing; SSM is the state-space model SSM processing, which is very suitable for processing long sequences. Like the human brain, it can maintain the "memory" of previous information while reading current information, thus understanding long-range dependencies. It performs deep sequence modeling along each scanning path; reshape is the reshaping operation, σ is the Sigmoid function, and Z is the gated feature obtained in step S801; Linear is the linear mapping operation.

[0060] Please see Figure 4 Step S403: Concatenate the first and second features along the channel dimension and project them back to the original dimension to obtain the enhanced features. The hierarchical features after encoder processing are... Its corresponding first feature is The corresponding second feature is During splicing and projection processing, first... and The features are concatenated and then projected back to the original dimensions to obtain the enhanced features F1. Similarly, the same process is applied to i=2, 3, 4, resulting in four enhanced features, which can be denoted as F1. .

[0061] Please see Figure 3 Step S304: Use the decoder to process the enhanced features to obtain the predicted frame.

[0062] In a specific embodiment of the present invention, the decoder can be expressed by the formula: , , In the formula, f i For the i-th hierarchical feature, F i f i The corresponding enhancement features, Here, P represents the intermediate features of the decoder, Up2 is the ConvTranspose2d operation, BN is batch normalization, δ is the ReLU activation function, and Up is the upsampling operation. In this embodiment, the decoder adopts a typical pyramid-style upsampling structure, gradually restoring spatial resolution by fusing the hierarchical features of the encoder with the corresponding scale-enhanced features layer by layer. Specifically, deep features are upsampled through transposed convolution and added to shallow features, combined with batch normalization and activation functions to enhance nonlinear expressive power. Finally, two upsampling operations output a predicted frame with the same size as the input, completing the accurate mapping from the feature space to the pixel space.

[0063] For the video frame prediction model described above, it is necessary to construct a dataset and a model, and then train the model using the dataset to obtain a trained video frame prediction model. During the training phase, the loss function used may include, for example, intensity loss L. int and gradient loss L grl Among them, the strength loss L int The L2 norm is used to calculate the pixel-wise difference between the predicted frame P and the ground frame G, and the gradient loss L is applied. grl The L1 norm is used to calculate the gradient differences between the predicted frame P and the ground truth frame G in the horizontal and vertical directions. Their calculation formulas are as follows: , .

[0064] Back Figure 1 Step S104: Based on all frames in the first video frame sequence and their corresponding predicted frames, obtain the score for each frame in the first video frame sequence. This step is achieved by quantifying the difference between the predicted frame and the real frame and processing it into anomaly score. The higher the score, the greater the difference and the higher the probability of an anomaly.

[0065] Please see Figure 9 In a specific embodiment of the present invention, step S104 includes steps S901 to S903.

[0066] Step S901: Based on all frames in the first video frame sequence and their corresponding predicted frames, obtain the peak signal-to-noise ratio (PSNR) of each frame in the first video frame sequence. This step can be expressed by the formula: , In the formula, P is the predicted frame of the current frame G, M represents the total number of pixels in the frame, and maxP is the maximum value in the predicted frame P. After calculating the peak signal-to-noise ratio of each frame in the first video frame sequence, it can be normalized to the interval [0,1].

[0067] Step S902: Obtain the maximum peak signal-to-noise ratio (PSNR) based on the peak signal-to-noise ratio (PSNR) of all frames in the first video frame sequence. max and minimum PSNR min .

[0068] Step S903: Based on the peak signal-to-noise ratio (PSNR) of each frame in the first video frame sequence, and the maximum and minimum PNR values, obtain the score for each frame in the first video frame sequence. The formula for calculating the score is as follows: , In the formula, S(t) is the score of the t-th frame, and PSNR is... t Let be the peak signal-to-noise ratio of the t-th frame.

[0069] Step S105: Based on the score and threshold of each frame in the first video frame sequence, obtain the anomaly detection result for each frame in the first video frame sequence. Based on the score calculated for each frame, a threshold can be set to distinguish between normal frames and abnormal frames.

[0070] Please see Figure 10 In a specific embodiment of the present invention, due to the complexity and diversity of video scenes, a fixed threshold is difficult to adapt to all situations. Therefore, an adaptive threshold method can be used to dynamically determine the anomaly judgment criteria. In this case, step S105 includes steps S1001 to S1004.

[0071] Step S1001: Generate multiple optional thresholds from the [0,1] interval using a uniform distribution. Considering that the outlier scores have been normalized to the [0,1] interval, covering all possible threshold ranges, 100 optional thresholds can be generated using a uniform distribution in this step.

[0072] Step S1002: For each optional threshold, generate predicted labels based on the anomaly score, and calculate the F1 score (i.e., the harmonic mean of precision and recall) of these predicted labels and the true labels.

[0073] For example, with an optional threshold of 0.35, we first iterate through each frame in the first video frame sequence, comparing the anomaly score S(t) of frame t with 0.35. If S(t) ≥ 0.35, the frame is determined to be abnormal, and its predicted label is 1; if S(t) < 0.35, the frame is determined to be normal, and its predicted label is 0. Through this process, a binary predicted label (0 or 1) is generated for each frame in the first video frame sequence, thus obtaining a set of prediction results based on the current threshold (0.35).

[0074] Next, this set of prediction results will be precisely compared with the previously known true labels. First, by comparing the predicted labels and true labels frame by frame, four basic indicators will be statistically determined: True Positive (TP): The number of frames that the model predicts as anomalous (1) and that are also anomalous (1) in reality; False positive (FP): The number of frames that the model predicts as abnormal (1) but are actually normal (0) (false alarms); False negative (FN): The number of frames that the model predicted as normal (0) but were actually abnormal (1) (missed). True negative (TN): The number of frames that the model predicts as normal (0) and that are also true as normal (0).

[0075] Secondly, calculate precision and recall: Precision (P) = TP / (TP+FP), which represents how many of the frames predicted as anomalous by the model are actually anomalous, and measures the accuracy of the model's alarm. Recall (R) = TP / (TP+FN) represents how many frames of true anomalies were successfully detected by the model. It measures the comprehensiveness of the model's anomaly detection.

[0076] Finally, the F1 score = 2 × (P × R) / (P + R) is the harmonic mean of precision and recall, which comprehensively considers the model's accuracy and comprehensiveness; a high F1 score means that the model has achieved a good balance between false positives (FP) and false negatives (FN). Understandably, the corresponding F1 score can be calculated for each optional threshold.

[0077] Step S1003: Select the optional threshold with the largest F1 score as the optimal threshold. For example, if the F1 score is the largest when the optional threshold is 0.75, then 0.75 is selected as the optimal threshold.

[0078] Step S1004: Based on the score of each frame in the first video frame sequence and the optimal threshold, obtain the anomaly detection result of each frame in the first video frame sequence.

[0079] The optimal threshold is typically calculated only once for a specific scenario. Within the same scenario, this optimal threshold can be used as a fixed threshold for identifying abnormal frames. Employing an adaptive threshold method broadens the applicability of this invention, while also increasing the accuracy of anomaly detection for a given scenario.

[0080] It should be noted that the steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they contain the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this patent.

[0081] To verify the effectiveness of this invention, simulation experiments were conducted on the method. All experiments were based on the PyTorch deep learning framework and performed on two RTX 3090 GPUs. The area under the receiver operating characteristic curve (AUC) was used as the evaluation metric. The three unsupervised drone video datasets used in the experiments were Drone-Anomaly, UIT-ADrone, and MUVAD. The training sets of all three datasets contained only normal events, while the test sets contained both normal and abnormal events.

[0082] The experiment compared the method of this invention with several existing video anomaly detection methods, including: ASTT (Transformer-based Spatio-Temporal Attention Network), MAPDM (Motion and Appearance Guided Patch Diffusion Model), VADMamba (Mamba-based Video Anomaly Detection Framework), and LGN-Net (Local-Global Normality Network). The experimental results are shown in Table 1 below. Figure 11 as well as Figure 12 .

[0083] Table 1: Comparison of AUC evaluation metrics between the method of this invention and other methods on different datasets.

[0084]

[0085] Figure 11The diagram illustrates the anomaly curves for four test video segments across three datasets. The dashed line represents an adaptive threshold; values ​​above the threshold indicate anomalies, and values ​​below indicate normal conditions. The diagram shows that the anomaly score curve remains below the threshold during normal periods; however, when anomalies occur, the model responds rapidly and maintains a high score during the anomaly persistence interval, demonstrating its excellent anomaly detection capability.

[0086] Figure 12 The visualization results of the present invention's predictions on three datasets are shown. The error map is generated by calculating the pixel-level differences between the ground truth frames and the predicted frames, where high-brightness areas represent larger prediction errors. The map reveals that for normal regions, the error map is predominantly black, indicating that the pixel-level differences between the predicted and ground truth frames are minimal. Conversely, the model fails to accurately predict abnormal regions, which appear as high-brightness areas in the error map.

[0087] Combined with Table 1, Figure 11 as well as Figure 12 It can be seen that the method of the present invention can significantly improve the detection accuracy compared with other unsupervised video anomaly detection methods.

[0088] In summary, the present invention has the following advantages: (1) Effective motion decoupling: By introducing frequency domain decoupling spatiotemporal correlation units, complex mixed motions can be analyzed in the frequency domain, effectively distinguishing between the motion of the UAV itself and the real motion of the foreground objects, significantly reducing false alarms caused by dynamic backgrounds, and enhancing sensitivity to real anomalies; (2) Powerful spatiotemporal modeling capability: By utilizing multi-scale spatiotemporal Mamba units, combining the sequential modeling advantages of state space models and multi-scale strategies, it can simultaneously capture long-term continuous changes and short-term sudden events, achieving a comprehensive understanding of video dynamics; (3) The method of the present invention has good versatility and scalability, and is not only applicable to UAV video surveillance, but can also be extended to other dynamic perspective video anomaly detection scenarios, such as vehicle monitoring, robot inspection, and other fields.

[0089] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A method for detecting abnormality of UAV video in dynamic background, characterized in that, include: Obtain the sequence of the first video frames to be detected; Take any frame in the first video frame sequence as the current frame, extract a preset number of video frames before the current frame, and obtain the second video frame sequence corresponding to the current frame; The second video frame sequence corresponding to the current frame is input into the trained video frame prediction model to obtain the predicted frame corresponding to the current frame. The video frame prediction model includes an encoder, a feature enhancement module, and a decoder. The feature enhancement module includes a frequency domain decoupled spatiotemporal correlation unit and a multi-scale spatiotemporal Mamba unit. Based on all frames in the first video frame sequence and their corresponding predicted frames, the score of each frame in the first video frame sequence is obtained. Based on the score and threshold of each frame in the first video frame sequence, the anomaly detection result of each frame in the first video frame sequence is obtained. Specifically, the second video frame sequence corresponding to the current frame is input into the trained video frame prediction model to obtain the predicted frame corresponding to the current frame, including: Preprocess the second video frame sequence corresponding to the current frame; The encoder is used to process the preprocessed second video frame sequence to obtain hierarchical features; The feature enhancement module is used to process the hierarchical features to obtain enhanced features, including: processing the hierarchical features using the frequency domain decoupling spatiotemporal correlation unit to obtain a first feature; processing the hierarchical features using the multi-scale spatiotemporal Mamba unit to obtain a second feature; and concatenating the first feature and the second feature in the channel dimension and projecting them back to the original dimension to obtain the enhanced features. The enhanced features are processed using the decoder to obtain the predicted frame. 2.The method of claim 1, wherein, The preprocessing includes size adjustment and normalization; the encoder includes an M-stage pyramid visual transformer, each stage of which outputs a hierarchical feature at a specific scale. 3.The method of claim 1, wherein, The hierarchical features are processed using the frequency domain decoupling spatiotemporal correlation unit to obtain the first feature, including: A one-dimensional fast Fourier transform is performed on the hierarchical features along the time dimension to obtain complex-valued frequency domain features and frequency amplitudes. Based on the frequency amplitude and the preset normalized frequency vector, the frequency-related weights are obtained; After weighting the complex-valued frequency domain features using the frequency-related weights, an inverse fast Fourier transform is performed to obtain the frequency-weighted features. The power spectral density is obtained by performing a two-dimensional fast Fourier transform on the frequency-weighted features along the spatiotemporal dimension. A two-dimensional inverse Fourier transform is performed on the power spectral density to obtain the spatiotemporal autocorrelation matrix; The spatiotemporal autocorrelation matrix is ​​used as attention weights to process the frequency-weighted features, resulting in the first feature. 4.The method of claim 1, wherein, The multi-scale spatiotemporal Mamba unit can be expressed by the following formula: , wherein, is an η times expanded recombination transform, is an η times expanded inverse recombination transform, η ∈ {1, 2, …, n}, n is the maximum expansion factor; f η-1 is an output of the η times expanded inverse recombination transform, f0is the hierarchical feature, STMamba is a spatio-temporal Mamba block, is a second feature corresponding to the hierarchical feature.

5. The method for detecting anomalies in UAV video under dynamic backgrounds according to claim 4, characterized in that, The processing procedure for the spatiotemporal Mamba block is as follows: The input features of the spatiotemporal Mamba block are sequentially normalized and linearly projected to obtain core features and gated features; Each pixel position of the core feature is flattened in row-major order and column-major order to obtain a flattened vector, and the flattened vectors are concatenated in chronological order to obtain a first sequence and a second sequence. Each frame of the core feature is divided into non-overlapping blocks. For each block position, the flattened vectors are obtained by flattening in row-major order and column-major order, and the flattened vectors are concatenated in chronological order to obtain the third sequence and the fourth sequence. The first sequence, the second sequence, the third sequence, and the fourth sequence are scanned forward and backward respectively to obtain 8 scan sequences; Each of the scan sequences is sequentially subjected to causal one-dimensional convolution, SSM state space model processing, and gating operation with the gated features; After summing the gating results of all scan sequences, linear mapping and residual connection are performed to obtain the output features of the spatiotemporal Mamba block.

6. The method for detecting anomalies in UAV video under dynamic backgrounds according to claim 1, characterized in that, The decoder can be expressed by the following formula: , , In the formula, f i For the i-th hierarchical feature, F i f i The corresponding enhancement features, δ represents the intermediate features of the decoder, P represents the predicted frame, Up2 represents the ConvTranspose2d operation, BN represents batch normalization, δ represents the ReLU activation function, and Up represents the upsampling operation.

7. The method for detecting anomalies in UAV video under dynamic backgrounds according to claim 1, characterized in that, Based on all frames in the first video frame sequence and their corresponding predicted frames, a score is obtained for each frame in the first video frame sequence, including: Based on all frames in the first video frame sequence and their corresponding predicted frames, the peak signal-to-noise ratio of each frame in the first video frame sequence is obtained. Based on the peak signal-to-noise ratio (PSNR) of all frames in the first video frame sequence, the maximum and minimum values ​​of the PSNR are obtained. The score for each frame in the first video frame sequence is obtained based on the peak signal-to-noise ratio (PSNR) of each frame and the maximum and minimum values ​​of the PSNR.

8. The method for detecting anomalies in UAV video under dynamic backgrounds according to claim 1, characterized in that, Based on the score and threshold of each frame in the first video frame sequence, the anomaly detection result of each frame in the first video frame sequence is obtained, including: Multiple selectable thresholds are generated from the [0,1] interval using a uniform distribution. For each optional threshold, predictive labels are generated based on the anomaly score, and the F1 score of these predictive labels is calculated compared with the true labels. Choose the threshold with the largest F1 score as the optimal threshold; Based on the score of each frame in the first video frame sequence and the optimal threshold, the anomaly detection result of each frame in the first video frame sequence is obtained.

Citation Information

Patent Citations

  • Self-supervision video anomaly detection method combined with self-attention module

    CN118262273A