Auto-regression score matching method for video anomaly detection

Through the autoregressive score matching method, combined with the conditional noise score transformer model of scene, motion and appearance information, the problem of insufficient likelihood in the prior art is solved, and more efficient video anomaly detection is achieved.

CN120339936APending Publication Date: 2025-07-18NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510291821.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing video anomaly detection method relies on distribution calculation likelihood to distinguish anomaly events in local patterns, resulting in insufficient generalization capabilities of the model.

Method used

The autoregressive score matching method is used to construct a conditional noise score transformer model, and the noise intensity is controlled using the diffusion time step, combined with scene, motion and appearance information, and the autoregressive noise denoising score matching mechanism performs abnormal score matching.

Benefits of technology

It improves the accuracy and efficiency of video abnormality detection, can effectively identify abnormal events in video, and is suitable for security monitoring scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339936A_ABST
    Figure CN120339936A_ABST
Patent Text Reader

Abstract

The invention specifically relates to an autoregressive score matching method for video anomaly detection, and the method comprises the steps: obtaining a preset number of video frames in a to-be-processed video, forming an input video clip, and inputting the input video clip into a conditional noise fraction transformer model which comprises a multi-layer perceptron, an embedded layer, a transformer block and a linear layer; the method comprises the following steps of: disturbing an input video clip by using noise with increasing intensity, and dividing disturbed noise-added data into a plurality of video image blocks; inputting the video image block into an embedding layer to obtain an embedding vector; inputting the diffusion time step length for controlling the noise intensity and the scene class label into a multi-layer perceptron to obtain an overall condition; inputting the embedded vector and the overall condition corresponding to the embedded vector into a transformer block for processing, and predicting noise added to the image block by using a linear layer; and obtaining an abnormal score based on likelihood through an autoregression denoising score matching mechanism. According to the method, the video anomaly detection task can be well processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly relates to an autoregressive score matching method for video anomaly detection. Background Art

[0002] Video surveillance systems are widely used in public and private domains, playing an important role in maintaining public order and ensuring the security of private assets. As one of the key applications of intelligent video surveillance systems, video anomaly detection aims to quickly and accurately identify abnormal events in videos, providing important support for real-time monitoring and event response.

[0003] Since abnormal events are very rare and scattered, it is impossible to collect abnormal event data fully and comprehensively. The semi-supervised setting of learning from relatively abundant normal event data is more practical. Therefore, related technologies set video anomaly detection as a semi-supervised task. In this semi-supervised setting, the training data only contains normal events without specific labels. The long-term goal of related technologies is to train a single classifier that faithfully learns the data distribution of normal event data while detecting and labeling events that deviate from the learned distribution as abnormal. Specifically, related technologies usually build deep autoencoders based on methods of reconstructing or predicting video frames to learn the features of normal events in videos. However, these methods often only consider the low-level pixel details in the videos and cannot ensure that the model does not produce unexpected generalization to abnormal events. Recently, the powerful pattern coverage ability of generative models has provided a new solution for the video anomaly detection task, that is, under the statistical model trained with normal data, abnormal events outside the distribution can be detected through the natural low likelihood of anomalies. These likelihood-based methods construct generative probability models to approximate the distribution of normal data and detect outliers that deviate from the learned pattern. However, although abnormal events are inexhaustible and very rare, they still have some commonalities. These relatively consistent abnormal events may exist in local patterns near the learned normal distribution. Relying solely on the distribution to calculate the likelihood is not sufficient to distinguish abnormal events in these local patterns. This inherent defect limits the further development of likelihood-based methods.

[0004] It should be noted that the information disclosed in the above background art section is only used to strengthen the understanding of the background of the present invention. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The present invention provides an autoregressive score matching method for video anomaly detection, which is used to solve the defect that in the existing methods, relying solely on the distribution to calculate the likelihood is not sufficient to distinguish abnormal events in these local patterns.

[0006] Other features and advantages of the present invention will become apparent from the following detailed description or, in part, be learned through the practice of the present invention.

[0007] According to a first aspect of the present invention, there is provided an autoregressive score matching method for video anomaly detection, the method comprising:

[0008] Obtain a preset number of video frames in the video to be processed to form an input video segment;

[0009] Construct a pre-trained conditional noise score Transformer model, the conditional noise score Transformer model including a multi-layer perceptron, an embedding layer, a Transformer block, and a linear layer, for predicting the noise added to the input video segment;

[0010] Add noises with different intensities to the input video segment, the intensity of the noise being controlled by the diffusion time step; at the same time, divide the perturbed noisy video segment into several video image blocks, and input the video image blocks into the conditional noise score Transformer model;

[0011] Convert the video image blocks into embedding vectors through the embedding layer, and input the diffusion time step and the video scene class label for controlling the noise intensity corresponding to the embedding vectors into the multi-layer perceptron to obtain the overall conditions corresponding to the embedding vectors;

[0012] Input the embedding vectors and the corresponding overall conditions into the Transformer block for processing, and use the linear layer to predict the noisy video image blocks;

[0013] Calculate the score of the input video segment at the corresponding noise intensity according to the predicted noisy video image blocks and the original video image blocks, calculate the motion weight of the score at the corresponding noise intensity according to the key frames of the video segment; set an objective function according to the score of the input video segment at the corresponding noise intensity and the motion weight, and train the conditional noise score Transformer model;

[0014] Calculate the anomaly score based on the norm of the score of the input video segment at the corresponding noise intensity in the inference stage, and predict the anomaly score corresponding to each intensity under increasing noise intensity in an autoregressive manner.

[0015] In some exemplary embodiments, the method further includes performing scaling, cropping, and mean normalization processing on a preset number of video frames.

[0016] In some exemplary embodiments, the Transformer block adopts an extensible adaptive layer normalization mechanism to embed and transmit the overall condition z:

[0017]

[0018] Among them, h and h' are the hidden outputs within the transformer module, both are hyperparameters, MHA is the multi-head self-attention mechanism, and FFN is the feed-forward neural network.

[0019] In some exemplary embodiments, calculating the motion weight of the score corresponding to the noise intensity according to the key frames of the video clip includes:

[0020] Taking the first frame and the last frame as key frames and calculating the absolute difference between the key frames;

[0021] Dividing the absolute difference into non-overlapping video image blocks;

[0022] Calculating the maximum difference along the c dimension and taking the cross-channel average of each video image block; where c is the number of input channels;

[0023] Normalizing the average value to obtain the motion weight.

[0024] In some exemplary embodiments, the objective function is:

[0025]

[0026] Among them, L represents how many noise levels, is the distribution expectation of the input under the noise addition condition, is the input video clip after noise addition is the video image block into which the input video clip is cut, P j is the video image block into which the original input video clip x is cut, represents the predicted score at the input noise level σ i and ω j is the motion weight.

[0027] In some exemplary embodiments, calculating the anomaly score based on the norm of the score corresponding to the input video clip at the corresponding noise intensity in the inference stage is specifically:

[0028] At each noise level σ i , dividing the video into η video clips each with T frames, and in each T-frame segment in the video, taking the maximum score as the score of the video clip; normalizing each score score i (t) to obtain the anomaly score S i (t) within the range of [0,1]:

[0029]

[0030] Wherein, is the L2 norm of the scores predicted by the network, and score i (t) represents the anomaly score of the input video frame t at the i-th noise level. The peak signal-to-noise ratio (PSNR) value between the input sequence x t and the denoised data is used as the denominator of the score norm.

[0031] In some exemplary embodiments, predicting the anomaly score corresponding to each intensity at increasing noise intensities in an autoregressive manner includes:

[0032] Adding noise of the first intensity, and denoising with the noise predicted by the conditional noise score transformer model. Replacing the original input video segment with the denoised reconstructed video segment and adding the second noise, and repeating this process to autoregressively obtain the anomaly scores corresponding to a series of intensities, wherein the first noise is less than the second noise.

[0033] According to a second aspect of the present invention, there is provided a storage medium having stored thereon a computer program, which when executed by a processor implements the autoregressive score matching method for video anomaly detection described in the first aspect above.

[0034] According to a third aspect of the present invention, there is provided a computer program product having stored thereon a computer program, which when executed by a processor implements the autoregressive score matching method for video anomaly detection described in the first aspect above.

[0035] According to a fourth aspect of the present invention, there is provided an electronic device, including:

[0036] A processor; and

[0037] A memory for storing executable instructions of the processor;

[0038] wherein the processor is configured to implement the autoregressive score matching method for video anomaly detection described in the first aspect above when executing the executable instructions.

[0039] The autoregressive score matching method for video anomaly detection provided by the embodiments of the present invention designs a novel conditional noise score Transformer based on the autoregressive denoising score matching mechanism, which combines the efficiency of score matching with the powerful performance of the diffusion probability model. To form a score function more suitable for video anomaly detection, the present invention considers the characteristics of video anomalies and modifies the score function from aspects of scene, motion, and appearance. Specifically, the scene condition is first embedded into the model to estimate the scene-related score function. Secondly, motion weights are assigned to the score function according to the differences between the key frames of the input sequence. Thirdly, the denoised data and the original data are compared to obtain the differences and aggregated with the score function to achieve enhanced appearance perception. Finally, the present invention proposes adding enhanced noise autoregressively to the denoised data during the inference process and estimating the corresponding score function for anomaly accumulation. The method of the present invention can handle video anomaly detection well.

[0040] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0042] Figure 1 It is a schematic diagram of the method flow of the present invention;

[0043] Figure 2 It is a schematic diagram of the overall model of the method of the present invention;

[0044] Figure 3 It is a schematic diagram of the conditional noise score Transformer block proposed by the present invention;

[0045] Figure 4 It is a schematic diagram of the autoregressive denoising score matching mechanism for video anomaly detection of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this invention will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments.

[0047] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0048] Aiming at the disadvantages and deficiencies of the prior art, in this exemplary embodiment, an autoregressive score matching method for video anomaly detection is provided. First, a preset number of video frames in the video to be processed are obtained to form an input video segment, which is input into the conditional noise score Transformer model. The conditional noise score Transformer model is constructed by a multi-layer perceptron, an embedding layer, a Transformer block, and a linear layer for predicting noise. In this model, the present invention perturbs the input video segment with noise of increasing intensity and divides the perturbed noisy data into several video image blocks. Next, the present invention inputs these video image blocks into the embedding layer to obtain embedding vectors. At the same time, the present invention inputs the diffusion time step for controlling the noise intensity and the scene class label into the multi-layer perceptron to obtain an overall condition. Then, the present invention inputs the embedding vectors and the corresponding overall conditions into the Transformer block for processing, and uses the linear layer to predict the noise added to the image blocks. Finally, the present invention obtains an anomaly score based on likelihood through an autoregressive denoising score matching mechanism. The method of the present invention can well handle the video anomaly detection task.

[0049] Reference Figure 1 As shown, it may specifically include the following steps:

[0050] Step 1: Obtain a preset number of video frames in the video to be processed to form an input video segment;

[0051] Exemplarily, a preset number of video frames in the video to be processed are obtained, and the preset number of video frames are scaled, cropped, and mean-normalized to obtain video data with a dimension of 3×T×160×160. Optionally, T represents the preset number of video frames selected from the video to be processed. In this embodiment, T = 8.

[0052] Step 2: Construct a pre-trained conditional noise score Transformer model, where the conditional noise score Transformer model includes a multi-layer perceptron, an embedding layer, a Transformer block, and a linear layer for predicting the noise added to the input video segment;

[0053] Step 3: Input the input video clip into the pre-trained conditional noise score Transformer model, add noises with different intensities to the input video clip, where the intensity of the noise is controlled by the diffusion time step, and at the same time, divide the perturbed noisy video data into several video image patches;

[0054] Exemplarily, perturb the video clip with noises of increasing intensity, and divide the perturbed noisy data into N = 8×100 video image patches P j ∈R d×d×c , where the size d of each video image patch is 16. Then, send the noisy video image patches into the pre-trained conditional noise score Transformer model.

[0055] Step 4: Convert the video image patches into embedding vectors through the embedding layer, and input the diffusion time step and the video scene class label that control the noise intensity corresponding to the embedding vectors into the multi-layer perceptron to obtain the overall conditions corresponding to the embedding vectors;

[0056] Step 5: Input the embedding vectors and the overall conditions corresponding to the embedding vectors into the Transformer block for processing, and use the linear layer to predict the noise added to the video image patches;

[0057] Exemplarily, input the embedding vectors and the overall conditions corresponding to the embedding vectors into the Transformer block for processing, and use the linear layer to predict the noisy video image patches;

[0058] Step 6: Calculate the error according to the predicted noise and the ground truth of the noise to train the model, regard this error as the score of the input video clip under the corresponding noise intensity, and at the same time calculate the motion weight according to the key frames of the video clip and assign it to the corresponding score;

[0059] Exemplarily, calculate the score of the input video clip under the corresponding noise intensity according to the predicted noisy video image patches and the original video image patches, calculate the motion weight of the score under the corresponding noise intensity according to the key frames of the video clip; set the objective function according to the score of the input video clip under the corresponding noise intensity and the motion weight, and train the conditional noise score Transformer model;

[0060] Step 7: In the inference stage, take the norm of the score of the input video clip under the corresponding noise intensity as the anomaly score; the network predicts the anomaly score corresponding to each intensity under increasing noise intensity in an autoregressive manner; the network predicts the anomaly score corresponding to each intensity under increasing noise intensity in an autoregressive manner. First, add noise of small intensity, and denoise with the noise predicted by the network. Replace the original input video clip with the denoised reconstructed video clip and add noise of greater intensity, and repeat this process to autoregressively obtain the anomaly scores corresponding to a series of intensities, so as to solve the anomaly detection task.

[0061] Specifically, the score matching mechanism is one of the common mechanisms for efficiently training non-linear probability models, where the score is defined as the gradient of the log-density of the data x The present invention can perturb the data x with independently and identically distributed Gaussian noise and use the score matching mechanism to train the score network to estimate the distribution of the perturbed data where σ is the noise level. The objective function is equivalent to matching the score of the non-parametric Parzen density estimator of the data x:

[0062]

[0063] This approximation can be further used to construct a generative model by perturbing the data with a series of Gaussian noises with increasing intensities and training the score network conditioned on the noise level σ i : If the present invention selects the noise distribution as then the score function is equivalent to which can be regarded as a vector field pointing in the denoising direction. Therefore, the objective function can be written as:

[0064]

[0065] Specifically, the conditional noise score Transformer model can not only coordinate the complex spatio-temporal information of the video, but also consider various conditions such as the noise level. The present invention modifies the denoising score matching formula to train a new type of conditional noise score Transformer. Let x t ={f t ,…,f t+n-1} be the n-frame input video segment at time step t, be a geometric sequence satisfying At the noise level σ i , the present invention perturbs the input video sequence with Gaussian noise ε to obtain whose distribution is The present invention samples these n perturbed frames to obtain the corresponding video image patches, and uses the uniform video image patch embedding method to obtain the embedding vectors of each video image patch. Let N denote the number of non-overlapping patches of size d×d×c from each sequence , where c is the number of input channels and d is a hyperparameter determined by n and N. For this purpose, the present invention changes the objective to make a per-patch denoising score matching function:

[0066]

[0067] Different from the Fisher score commonly used in statistics the score considered in the present invention is a function of the input sequence x t rather than the model parameter θ. Therefore, the present invention can improve this score according to the characteristics of the input data to obtain better performance. Especially in security-conscious monitoring scenarios, the dynamic foreground and static background of video frames are not balanced. Therefore, the score tends to reflect more static information because it is a vector field pointing in the direction where the probability density function has the maximum growth rate. For this reason, it is better to incorporate motion information into the score function additionally. The present invention uses the first and last frames of the input sequence as key frames during the operation of the denoising score matching mechanism and calculates the key frame difference therebetween to assign weights to the score. For each input sequence x t the present invention first calculates the absolute difference between the key frames to obtain g t = |f t+n-1 - f t |. Then, the present invention divides it into non-overlapping video image blocks Next, the present invention calculates the maximum difference along the c dimension and takes the cross-channel average of each block, which can be written as:

[0068]

[0069] Normalize this value to calculate the per-block weight and assign it to the scoring function of the present invention:

[0070]

[0071] where is the corresponding block after the noisy input video segment is cut, and P j is the corresponding block after the original input video segment is cut.

[0072] By implementing the motion-weighted denoising score matching process, the score output by the model of the present invention can contain more motion information.

[0073] Exemplarily, as shown in Figure 3 is a schematic diagram of a conditional noise score transformer block provided by an embodiment of the present invention. The transformer block includes layer normalization, scale shift, multi-head self-attention, scale, layer normalization, scale shift, feed-forward neural network, scale. Specifically, the present invention trains a model conditioned on the diffusion time step i to control the noise level σ i(Embedded via a multi-layer perceptron). With the attention mechanism-based architecture, continuous supervision of the multi-scale block denoising process can effectively generate a conservative vector field, making anomalies distinguishable. Considering the scene-dependent anomalies specific to the video anomaly detection task. For each video clip, the present invention assigns a scene class label y according to the scene it belongs to (i.e., the camera ID). Then, the present invention combines the scene class label y with the diffusion time step i to obtain the overall condition z. The present invention applies an extensible adaptive layer normalization mechanism to embed and transmit the overall condition z:

[0074]

[0075] where h and h' are the hidden outputs within the transformer module, both are hyperparameters, MHA is the multi-head self-attention mechanism, and FFN is the feed-forward neural network. In this more adaptive way, the present invention effectively transmits the conditional information to each transformer block, thus achieving excellent performance and faster model convergence.

[0076] Exemplarily, referring to Figure 4 shown, is a schematic diagram of the autoregressive denoising score matching process provided by an embodiment of the present invention. Specifically, in the inference stage, in order to obtain a more comprehensive and appropriate anomaly score, the present invention considers the L2 norm of the score:

[0077]

[0078] where, s θ (x) is the score function output by the network, is the gradient of the log density of the data x, is the log density, and p(x) is the data density.

[0079] Since the data density term appears in the denominator, normal events with high probability will correspond to low norms, while abnormal events with low probability will correspond to high norms. However, video anomalies may also cluster due to common features, such as being too fast or rarely occurring, resulting in meaningless scores because these anomalies are located in "flat" regions with small gradients. Utilizing multiple noise levels can alleviate this problem because multi-scale estimation can capture the gradient signal from the relevant context. The present invention proposes two improved methods for separating these "flat" regions: the autoregressive method and the appearance method during the inference process. For the former, each time the present invention perturbs the original data with enhanced noise and jointly estimates the score function, denoised data can be obtained The present invention modifies the denoising score matching mechanism, adds enhanced noise to the denoised data rather than the original data in an autoregressive manner, and estimates the corresponding score function. This method helps accumulate abnormal abnormal contexts, making the score function more effective for video anomaly detection. For the latter, the score function faithfully estimates the data distribution while ignoring the rare occurrence characteristics of the anomalies themselves. Using the peak signal-to-noise ratio (PSNR) value between x t and as the denominator of the score norm can semantically make up for this shortcoming without using more computing resources. Finally, for each input sequence x t , the present invention obtains L scores at time step t from various noise levels σ i :

[0080]

[0081] During the anomaly detection inference process, at each noise level σ i , the video is divided into η video segments with T frames. In each T-frame segment of the video, the maximum score is taken as the score of the video segment. Then, at each noise level σ i , the present invention normalizes each score score i (t) to obtain an anomaly score S i (t) in the range of [0,1] to quantify the probability of an anomaly occurring, which can be written as:

[0082]

[0083] Furthermore, the present invention performs a weighted sum and average on the obtained anomaly scores to obtain a final anomaly metric, completing the video anomaly detection.

[0084] It should be noted that, on the other hand, the present application also provides a storage medium, which may be included in an electronic device; or may exist separately without being assembled into the electronic device. The above storage medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the following embodiments.

[0085] In one embodiment, the present application provides a computer program product, including a computer program, which when executed by a processor implements the steps in the above method embodiments.

[0086] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, rather than for restrictive purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0087] Other embodiments of the present invention will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed by the present invention. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the present invention are pointed out by the claims.

[0088] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only defined by the appended claims.

Claims

1. An autoregressive score matching method for video anomaly detection, characterized in that The method includes: Obtaining a preset number of video frames in the video to be processed to form an input video segment; Constructing a pre-trained conditional noise score Transformer model, which includes a multi-layer perceptron, an embedding layer, a Transformer block, and a linear layer, for predicting the noise added to the input video segment; Adding noises with different intensities to the input video segment, where the intensity of the noise is controlled by the diffusion time step; at the same time, dividing the perturbed noisy video segment into several video image blocks and inputting the video image blocks into the conditional noise score Transformer model; Converting the video image blocks into embedding vectors through the embedding layer, and inputting the diffusion time step and the video scene class label that control the noise intensity corresponding to the embedding vectors into the multi-layer perceptron to obtain the overall conditions corresponding to the embedding vectors; Inputting the embedding vectors and the corresponding overall conditions into the Transformer block for processing, and using the linear layer to predict the noisy video image blocks; Calculating the score of the input video segment at the corresponding noise intensity according to the predicted noisy video image blocks and the original video image blocks, and calculating the motion weight of the score at the corresponding noise intensity according to the key frames of the video segment; setting an objective function according to the score of the input video segment at the corresponding noise intensity and the motion weight, and training the conditional noise score Transformer model; Calculating an anomaly score based on the norm of the score of the input video segment at the corresponding noise intensity in the inference stage, and predicting the anomaly score corresponding to each intensity under increasing noise intensities in an autoregressive manner.

2. The autoregressive score matching method for video anomaly detection according to claim 1, wherein The method further includes performing scaling, cropping, and mean normalization processing on a preset number of video frames.

3. The autoregressive score matching method for video anomaly detection according to claim 1, wherein The Transformer block adopts an extensible adaptive layer normalization mechanism to embed and transmit the overall condition z: Among them, h and h' are the hidden outputs within the transformer module, both are hyperparameters, MHA is the multi-head self-attention mechanism, and FFN is the feed-forward neural network.

4. The autoregressive score matching method for video anomaly detection according to claim 1, wherein The calculating the motion weight of the score at the corresponding noise intensity according to the key frames of the video segment includes: Taking the first frame and the last frame as key frames and calculating the absolute difference between the key frames; Dividing the absolute difference into non-overlapping video image blocks; Calculating the maximum difference along the c dimension and taking the cross-channel average value of each video image block; where c is the number of input channels; Normalizing the average value to obtain the motion weight.

5. The autoregressive score matching method for video anomaly detection according to claim 4, wherein The objective function is: Among them, L represents the number of noise levels, is the distribution expectation of the input under the noise addition condition, is the input video segment after noise addition is the video image block cut from it, P j is the video image block cut from the original input video segment x, represents the prediction score at the input noise level σ i and ω j is the motion weight.

6. The autoregressive score matching method for video anomaly detection according to claim 1, characterized in that The calculating the anomaly score based on the norm of the score of the input video segment at the corresponding noise intensity in the inference stage is specifically: At each noise level σ i the video is divided into η video segments each having T frames. In each T-frame segment of the video, the maximum score is taken as the score of the video segment; for each score score i (t) is normalized to obtain an anomaly score S i (t) in the range [0, 1]: Among them, is the L2 norm of the scores predicted by the network, score i (t) represents the anomaly score of the input video frame t at the i noise level. Taking the peak signal-to-noise ratio (PSNR) value between the input sequence x t and the denoised data as the denominator of the score norm.

7. The method according to claim 1, characterized in that, The predicting the anomaly score corresponding to each intensity under increasing noise intensities in an autoregressive manner includes: Adding the noise of the first intensity, denoising with the noise predicted by the conditional noise score Transformer model, replacing the original input video segment with the denoised reconstructed video segment and adding the second noise, and repeating this process to autoregressively obtain the anomaly scores corresponding to a series of intensities, where the first noise is less than the second noise.

8. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the autoregressive score matching method for video anomaly detection as described in any one of claims 1 to 7.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the autoregressive score matching method for video anomaly detection according to any one of claims 1 to 7.

10. An electronic device, characterized in that, It includes: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the autoregressive score matching method for video anomaly detection according to any one of claims 1 to 7 by executing the executable instructions.