Traffic video anomaly detection method and system based on AIGC
By performing unsupervised training on a generative model based on AIGC and utilizing adversarial learning between the generator and the discriminator, the problem of traditional traffic video anomaly detection methods relying on manual rules and high-cost labeling is solved, and efficient anomaly detection in complex scenarios is achieved, improving detection accuracy and adaptability.
Patent Information
- Application Number
- CN202510793593.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional traffic video anomaly detection methods rely on manually formulated rules and thresholds, which results in limited effectiveness in detecting complex and diverse abnormal events. In addition, manual labeling is costly and has poor versatility.
An unsupervised training method based on AIGC is adopted to build a generative model. The adversarial learning of the generator and the discriminator is used to generate and identify the normal features of traffic videos, thus realizing anomaly detection without the need to define rules in advance.
The accuracy and adaptability of traffic video anomaly detection are improved, manual intervention is reduced, the generalization and robustness of the model are enhanced, and it can adapt to traffic video data from different cities and scenarios.
Smart Images

Figure CN120673354A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning and artificial intelligence technology, and in particular to a traffic video anomaly detection method and system based on AIGC. Background Art
[0002] With the accelerating pace of urbanization, traffic issues are becoming increasingly prominent, and urban traffic management and safety have become critical issues. Traffic monitoring systems are becoming increasingly widely used in this context, but the challenges they face are also becoming increasingly severe. Traditional traffic monitoring systems often face the challenge of detecting traffic anomalies when processing large amounts of traffic video data. These anomalies include traffic accidents, traffic violations, and congestion, which directly impact traffic flow and safety. Therefore, effectively identifying anomalies in traffic video is crucial for improving urban traffic safety and has high application value.
[0003] Traditional traffic video anomaly detection relies on manually formulated rules, thresholds, or features, limiting its effectiveness in addressing complex and diverse anomalies. Furthermore, developing and adjusting these rules requires specialized knowledge and significant time investment. Generative models, on the other hand, are a type of machine learning method that aims to generate new data samples by learning the distributional characteristics of data. In traffic video anomaly detection, AIGC-based models can learn the distributional characteristics of normal traffic flow and then generate normal samples based on these characteristics. When the input traffic video data deviates significantly from the generated normal distribution, the generative model can identify the presence of an anomaly. AIGC-based traffic video anomaly detection methods combine the advantages of generative models and can detect various types of traffic anomalies without relying on predefined rules or thresholds. By learning data distribution and features, this method can adapt to complex anomaly scenarios and improve detection accuracy and robustness. Therefore, the present invention proposes a traffic video anomaly detection method and system based on AIGC. Summary of the Invention
[0004] In response to the shortcomings of current traffic video anomaly detection methods, the present invention provides a traffic video anomaly detection method and system based on AIGC. To solve the problem of high manual labeling costs, an unsupervised training method is adopted to effectively reduce manual intervention; to eliminate the correlation within the batch, the order of the initial features is randomized, which effectively improves the generalization ability of the model and improves the stability of the model; to solve the defect of poor versatility of traditional methods, by utilizing a generative model, it can adapt to different types of traffic video data, including different cities, different scenes, etc., so that the traffic video anomaly detection method based on AIGC can play a role in various application environments.
[0005] The present invention provides a traffic video anomaly detection method based on AIGC, the method comprising:
[0006] Build a traffic video dataset and perform preprocessing;
[0007] Build a feature extractor;
[0008] Extracting features from the preprocessed traffic video dataset using the feature extractor to generate a randomized feature vector;
[0009] Reconstructing the randomized feature vector using a generator;
[0010] The reconstructed features are input into the discriminator for classification and the probability value is output;
[0011] The output probability value of the discriminator is used to obtain the traffic video abnormal event score and realize traffic video anomaly detection.
[0012] Preferably, preprocessing the traffic video dataset includes:
[0013] Perform time-series segmentation on the video;
[0014] Use data enhancement technology to enhance the data of the segmented video;
[0015] Divide the data-augmented videos into training sets, validation sets, and test sets;
[0016] The time-series segmentation of the video includes:
[0017] V seg (i)={v i×L×f ,v (i+1)×L×f-1},i=1,2,…,N / (L×f)
[0018] Among them, V seg (i) represents the i-th time segment, v i represents the i-th frame of the video, L is the duration of the segment, f is the video frame rate, and N is the total number of frames in a video;
[0019] Data enhancement for the segmented video includes:
[0020] I bright (x,y)=I(x,y)+δ bright
[0021] I contrast (x,y)=α·(I(x,y)-μ)+μ
[0022] Among them, δ brightis the brightness adjustment amount, α is the contrast coefficient, μ is the average brightness of the image, I bright (x,y) and I contrast (x, y) is the enhanced video frame, and I(x, y) is the video frame image.
[0023] Preferably, extracting features from the pre-processed traffic video dataset using the feature extractor includes:
[0024] F i =Mobile-SFM(V seg (i))
[0025] Among them, F i Represents the feature extraction model network from the input data V seg (i) The extracted feature vectors.
[0026] Preferably, performing feature reconstruction on the randomized feature vector using a generator includes:
[0027] The default input random feature vector is X, and the output of the generator is X;
[0028] The generator consists of a series of convolutional layers, activation functions, and batch normalization. The output of each layer is:
[0029] x i =σ(BN(Conv3×3(x i-1 )))
[0030] Where x0 = X, the input vector; σ represents the activation function; Conv3×3(·) represents a 3×3 convolution operation; it consists of 5 layers in total, and the activation function of the last layer is tanh, while the activation functions of the remaining layers are all ReLU.
[0031] The output of the generator is judged by introducing a threshold control mechanism. Specifically, the samples reconstructed by the generator are divided, and samples larger than the threshold are input into the discriminator to train the generator. Specifically:
[0032]
[0033] When the reconstruction error Greater than the set threshold When , the generated feature vector is all 1, otherwise it is all 0; the threshold Set from 1 to 3.
[0034] Preferably, inputting the reconstructed features into a discriminator for classification includes:
[0035] The discriminator consists of a convolutional layer, an activation function, and a batch normalization layer. The output of each layer is expressed as:
[0036] x i =σ(BN(Conv3×3(x i-1 )))
[0037] x out =σ(Linear(x i ))
[0038] Where x0 = X, i.e., the input sample; σ represents the activation function; Conv3×3(·) represents a 3×3 convolution operation; it consists of 4 layers, including 3 convolutional layers and the last layer is a fully connected layer;
[0039] Discriminator loss function L D The cross entropy loss function is used to distinguish between generated samples and real samples. The expression is:
[0040] L D =logD(x)+log(1-D(x))
[0041] Among them, x represents the real sample, and x represents the sample generated by the generator.
[0042] The present invention also provides a traffic video anomaly detection system based on AIGC, the system is used to implement any one of the methods described above, the system comprising: a data set construction module, an extractor construction module, a feature extraction module, a feature reconstruction module, a feature classification module and an anomaly detection module;
[0043] The data set construction module is used to construct a traffic video data set and perform preprocessing;
[0044] The extractor construction module is used to construct a feature extractor;
[0045] The feature extraction module is used to extract features from the pre-processed traffic video data set using the feature extractor to generate a randomized feature vector;
[0046] The feature reconstruction module is used to perform feature reconstruction on the randomized feature vector using a generator;
[0047] The feature classification module is used to input the reconstructed features into the discriminator for classification and output a probability value;
[0048] The anomaly detection module is used to obtain the traffic video anomaly event score by using the output probability value of the discriminator, thereby realizing traffic video anomaly detection.
[0049] Preferably, preprocessing the traffic video dataset includes:
[0050] Perform time-series segmentation on the video;
[0051] Use data enhancement technology to enhance the data of the segmented video;
[0052] Divide the data-augmented videos into training sets, validation sets, and test sets;
[0053] The time-series segmentation of the video includes:
[0054] V seg (i)={v i×L×f ,v (i+1)×L×f-1},i=1,2,…,N / (L×f)
[0055] Among them, V seg (i) represents the i-th time segment, v i represents the i-th frame of the video, L is the duration of the segment, f is the video frame rate, and N is the total number of frames in a video;
[0056] Data enhancement for the segmented video includes:
[0057] I bright (x,y)=I(x,y)+δ bright
[0058] I contrast (x,y)=α(I(x,y)-μ)+μ
[0059] Among them, δ bright is the brightness adjustment amount, α is the contrast coefficient, μ is the average brightness of the image, I bright (x,y) and I contrast (x, y) is the enhanced video frame, and I(x, y) is the video frame image.
[0060] Preferably, extracting features from the pre-processed traffic video dataset using the feature extractor includes:
[0061] F i =Mobile-SFM(V seg (i))
[0062] Among them, F i Represents the feature extraction model network from the input data V seg (i) The extracted feature vectors.
[0063] Preferably, performing feature reconstruction on the randomized feature vector using a generator includes:
[0064] The default input random feature vector is X, and the output of the generator is X;
[0065] The generator consists of a series of convolutional layers, activation functions, and batch normalization. The output of each layer is:
[0066] x i =σ(BN(Conv3×3(x i-1 )))
[0067] Where x0 = X, the input vector; σ represents the activation function; Conv3×3(·) represents a 3×3 convolution operation; it consists of 5 layers in total, and the activation function of the last layer is tanh, while the activation functions of the remaining layers are all ReLU.
[0068] The output of the generator is judged by introducing a threshold control mechanism. Specifically, the samples reconstructed by the generator are divided, and samples larger than the threshold are input into the discriminator to train the generator. Specifically:
[0069]
[0070] When the reconstruction error Greater than the set threshold When , the generated feature vector is all 1, otherwise it is all 0; the threshold Set from 1 to 3.
[0071] Preferably, inputting the reconstructed features into a discriminator for classification includes:
[0072] The discriminator consists of a convolutional layer, an activation function, and a batch normalization layer. The output of each layer is expressed as:
[0073] x i =σ(BN(Conv3×3(x i-1 )))
[0074] x out =σ(Linear(x i ))
[0075] Where x0 = X, i.e., the input sample; σ represents the activation function; Conv3×3(·) represents a 3×3 convolution operation; it consists of 4 layers, including 3 convolutional layers and the last layer is a fully connected layer;
[0076] Discriminator loss function L D The cross entropy loss function is used to distinguish between generated samples and real samples. The expression is:
[0077] L D =logD(x)+log(1-D(x))
[0078] Among them, x represents the real sample, and x represents the sample generated by the generator.
[0079] Compared with the prior art, the present invention has the following beneficial effects:
[0080] This invention provides an effective method for detecting anomalies in traffic videos, leveraging the advantages of SFM fusion features to significantly improve the model's representational capabilities. This AIGC-based traffic video anomaly detection method overcomes the limitations of traditional methods, which are limited by manually formulated rules and features and struggle to handle complex and changing traffic scenarios. The invention of this AIGC-based traffic video anomaly detection method has significant benefits, improving the accuracy and adaptability of anomaly detection and providing strong support for security and management in the field of traffic monitoring. This technology has important practical application value for improving the urban traffic environment and ensuring driving safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0082] Figure 1 1 is a flow chart of a method for detecting anomalies in traffic video based on AIGC according to an embodiment of the present invention;
[0083] Figure 2 This is a schematic diagram of a synthesis and fusion module according to an embodiment of the present invention;
[0084] Figure 3 This is a schematic diagram of the generator structure of an embodiment of the present invention;
[0085] Figure 4 Schematic diagram of the discriminator structure of an embodiment of the present invention. DETAILED DESCRIPTION
[0086] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0087] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the described object changes, the relative position relationship may also change accordingly.
[0088] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0089] Example 1
[0090] The overall process of the present invention is as follows Figure 1 As shown, a traffic video anomaly detection method based on AIGC is provided. First, in order to solve the problem of high manual labeling costs, an unsupervised training method is adopted to effectively reduce manual intervention. At the same time, in order to eliminate the correlation within the batch, the order of the initial features is randomized, which effectively improves the generalization ability of the model and improves the stability of the model. Finally, in order to solve the defect of poor versatility of traditional methods, by utilizing a generative model, it can adapt to different types of traffic video data, including different cities, different scenes, etc., so that the traffic video anomaly detection method based on AIGC can play a role in various application environments. Finally, the discriminator is used to output the probability value of the abnormal event. The method specifically includes:
[0091] Build a traffic video dataset and perform preprocessing;
[0092] Build a feature extractor;
[0093] Extracting features from the preprocessed traffic video dataset using the feature extractor to generate a randomized feature vector;
[0094] Reconstructing the randomized feature vector using a generator;
[0095] The reconstructed features are input into the discriminator for classification and the probability value is output;
[0096] The output probability value of the discriminator is used to obtain the traffic video abnormal event score and realize traffic video anomaly detection.
[0097] In this embodiment, a traffic video dataset is constructed;
[0098] This step aims to construct a video dataset covering normal traffic scenes, ensuring that it does not contain any abnormal events such as traffic accidents, illegal parking, or congestion. The dataset will focus on normal traffic scenes in real-world scenarios, providing a reliable foundation for the development and evaluation of traffic video anomaly detection methods.
[0099] First, we need to clarify the definition of a normal traffic scene. A normal scene refers to a scene in which vehicles and pedestrians pass through traffic normally according to traffic rules, with smooth traffic flow and no obvious abnormal interference, such as no traffic accidents, no illegal parking, or abnormal congestion. With a clear definition, we can effectively filter out video clips that meet the requirements, thereby improving the diversity and representativeness of the dataset. During the data preprocessing stage, the video is first segmented in time series. Specifically, the goal of video segmentation is to cut a long video into multiple small time segments so that the model can effectively learn and analyze each sample while ensuring that the model can capture the spatiotemporal characteristics of the event. Assuming the total length of the video data is T and the video frame rate is f, the total number of frames in a video is N = T × f. To perform reasonable segmentation, we divide the video into several time segments of length L, each segment containing L × f frames. The segmented segments are shown in the following formula.
[0100] V seg (i)={v i×L×f ,v (i+1)×L×f-1},i=1,2,…,N / (L×f)
[0101] Among them, V seg (i) represents the i-th time segment, v i Denotes the i-th frame of the video, and L is the duration of the segment (s). This segmentation method helps the model capture the temporal information within each time segment and improves the ability to detect abnormal events.
[0102] Subsequently, data augmentation techniques are used to increase data diversity by adjusting the brightness, contrast, and rotation angle of the video frame, thereby improving the model's generalization ability. Assuming the video frame image is I(x, y), the goal of the data augmentation operation is to perform various transformations on the image to generate new image data, as shown in the following formula.
[0103] I bright (x,y)=I(x,y)+δ bright
[0104] I contrast (x,y)=α·(I(x,y)-μ)+μ
[0105] Among them, δ bright is the brightness adjustment amount, α is the contrast coefficient, μ is the average brightness of the image, I bright (x,y) and I contrast (x,y) is the enhanced video frame.
[0106] Finally, based on the characteristics of the dataset, the training set, validation set, and test set are divided into 70%, 10%, and 20% ratios to provide sufficient support for model training, optimization, and evaluation.
[0107] In this embodiment, a feature extractor is constructed;
[0108] In the feature extraction step of unlabeled videos, this paper proposes a feature extraction model called MobileV2-SFM, which combines the lightweight characteristics of MobileNetV2 and SFM to enhance the feature extraction capability of the model. First, the input video frame V seg (i),The feature extraction process is shown below.
[0109] F i =Mobile-SFM(V seg (i)) Where F i Represents the feature extraction model network from the input data V seg (i) The extracted feature vectors.
[0110] The feature extraction model consists of SFM and MobileV2, such as Figure 2 As shown in Figure 1, the SFM module contains three optional inputs: first, linearly scale the input, then add pixel by pixel, and finally fuse it using the MobileNet V2 network. This module can synthesize synthetic layers from original layers or simply use it for feature fusion. The SFM module is shown in the following formula.
[0111] F out =wiseAdd(F in +0.75F in +0.5F in )
[0112] Among them, F in Represents the input feature vector, represents the input feature vector of the SFM module.
[0113] The output feature vector is then input into the inverse residual module of MobileV2, which effectively enhances feature extraction capabilities. The inverse residual module employs an "expansion-contraction" approach, building a wide bottleneck layer between the input and output to increase the model's nonlinear expressiveness, thereby capturing a richer set of low-level and high-level features. This design not only better processes spatial features but also learns more meaningful features in the channel dimension, thereby improving model accuracy and robustness. The depthwise separable convolution and inverse residual architecture significantly reduce the model's parameter count and computational complexity while ensuring high performance. Depthwise separable convolution decomposes the conventional convolution operation into two independent steps: depthwise convolution and pointwise convolution, significantly reducing computational resources. Furthermore, the introduction of the inverse residual architecture further optimizes information flow, reduces redundant computation, and ensures efficient information transfer between network layers. The detailed process is shown in the equation below.
[0114] X1=ReLU6(Conv1×1(X))
[0115] X2=ReLU6(DepthwiseConv3×3(X1))
[0116] X out =Linear(Conv1×1(X2))+X
[0117] Where X represents the input feature vector, Conv1×1 represents the convolution with a convolution kernel of 1, ReLU6=min(max(0,x),6), X i Represents the intermediate process feature vector, DepthwiseConv3×3 represents the depth-separable convolution with a convolution kernel of 3, X out represents the output feature vector.
[0118] In this embodiment, the generator features reconstruction;
[0119] The generator uses a randomized feature vector as input to perform feature reconstruction operations. Its essence is a decoder structure that aims to generate feature representations consistent with the target distribution. In this process, the difference between the generated sample and the input sample is measured by calculating the reconstruction error, and the sample is classified based on a preset threshold. When the reconstruction error of a sample exceeds the set threshold, it is judged as an abnormal sample and input to the discriminator for further evaluation. Specifically, let the input random feature vector be X and the output of the decoder be X. The decoder consists of a series of convolutional layers, activation functions, and batch normalization, such as Figure 3 As shown, the output of each layer is shown in the following formula.
[0120] x i =σ(BN(Conv3×3(xi-1 )))
[0121] Where x0 = X, the input vector; σ represents the activation function (ReLU or tanh); and Conv3×3(·) represents a 3×3 convolution operation. The network consists of five layers, all of which use ReLU except the last one, which uses tanh.
[0122] Generally speaking, the generator is obtained by averaging the reconstruction errors L r Training:
[0123]
[0124] Where n is the number of samples, The reconstruction error of the nth sample represents the difference between the original feature and the reconstructed feature, calculated using the Euclidean norm; It is the feature vector value input to the generator, indicating the position of the sample in the two-dimensional feature space; is the value of the generator reconstruction vector; L r represents the average value of the reconstruction loss.
[0125] Specifically, the generator reconstructs the input feature vector and generates an output with the corresponding feature representation. Feature vectors with high reconstruction loss values usually represent abnormal samples that are significantly different from normal samples, while feature vectors with low reconstruction loss values indicate that the feature vector conforms to the normal pattern and is therefore considered a normal sample.
[0126] In the present invention, the training process of the generator is designed to optimize the performance of the generator by maximizing the error of the discriminator. The generator and the discriminator interact through adversarial learning to form a generative adversarial network framework. The idea of collaborative learning is adopted to train the discriminator through the output of the generator. In this process, the goal of the generator is to generate sufficiently realistic reconstructed features so that the discriminator cannot accurately distinguish between normal samples and abnormal samples, that is, to maximize the error rate of the discriminator. To achieve this, the present invention judges the output of the generator by introducing a threshold control mechanism. The samples reconstructed by the generator are divided, and the samples greater than the threshold are input into the discriminator to better train the generator, as shown in the following formula.
[0127]
[0128] From the above formula, we can see that when the reconstruction error is greater than Greater than the set threshold When , the generated feature vector is all 1, otherwise it is all 0; the threshold In this patent, it is set to 1 to 3, and the specific value can be adjusted according to the actual scenario and the distribution of training data.
[0129] In this embodiment, the discriminator improves the generator;
[0130] The discriminator is essentially a fully connected layer. This structure enables it to effectively integrate and process the global features of the input data without relying on local features or temporal information. Because the fully connected layer performs comprehensive calculations on each input feature through a weight matrix when processing data, it reduces sensitivity to noise and thus has greater robustness. In the task of anomaly detection in traffic videos, this robustness enables the discriminator to make stable judgments in the face of noise, ambiguous information, or incomplete data, ensuring the correct detection of abnormal events.
[0131] The discriminator improves the generator by providing accurate feedback, helping it optimize its output, enhance its ability to identify anomalous samples, and improve its robustness and generalization capabilities. Through collaborative training of the discriminator and generator, the generator can continuously improve its reconstruction process, leading to more effective detection of anomalies in traffic videos.
[0132] The discriminator is designed to play an adversarial role against the generator and separate the generated samples from the real samples. Specifically, let the real sample be X T , generate samples as X F . Not only should the adversarial relationship between the generator and the discriminator be considered, but also the balance between the generated samples and the real samples. Otherwise, as training progresses, the advantage or disadvantage of one will eventually lead to the inefficiency of the other. The discriminator consists of convolutional layers, activation functions, and batch normalization. The output of each layer can be expressed as:
[0133] x i =σ(BN(Conv3×3(x i-1 )))
[0134] x out =σ(Linear(x i ))
[0135] Where x0 = X, the input sample; σ represents the activation function (ReLU or tanh); Conv3×3(·) represents a 3×3 convolution operation. It consists of 4 layers, including 3 convolutional layers and the last layer is a fully connected layer. The samples generated by the discriminator share the same architecture as the real samples. Compared with the generator architecture, its setting is simpler, such as Figure 4 shown.
[0136] The discriminator continuously improves its discrimination ability by classifying generated samples and real samples, while also forcing the generator to generate more realistic samples.
[0137] Discriminator loss function LD The cross entropy loss function is used with the goal of enabling the discriminator to better distinguish between generated samples and real samples, as shown in the following formula.
[0138] L D =logD(x)+log(1-D(x))
[0139] Among them, x represents the real sample, and x represents the sample generated by the generator.
[0140] In this embodiment, traffic video anomaly detection is implemented;
[0141] This step uses the output probability value of the discriminator to obtain the abnormal event score of the traffic video.
[0142] This step uses the discriminator's output probability values to derive anomaly scores for traffic videos. Through adversarial training, the generator and discriminator continuously improve their capabilities in a competitive game. The generator's task is to generate samples that are as realistic as possible from random noise, aiming to make it difficult for the discriminator to distinguish between authentic and forged samples. As training progresses, the samples generated by the generator become increasingly realistic, gradually capturing the complex structure and distribution characteristics of the data.
[0143] The discriminator acts as a "referee," determining whether an input sample is genuine or forged by the generator. As the generator improves, the discriminator continuously learns to better distinguish between genuine and forged samples. Through continuous adversarial training, the discriminator continuously improves its accuracy in identifying forged samples, while also prompting the generator to produce more refined and realistic samples.
[0144] In the traffic video anomaly detection task, the goal is to determine whether the input traffic video conforms to the normal traffic video distribution during the training phase, and then detect whether there are any abnormal events (such as traffic jams, collisions, etc.). First, the input unlabeled traffic video clips are mapped back to the latent space used during training. By finding the closest feature points in the latent space, the difference between the samples generated by the generator and the input samples is minimized.
[0145] Specifically, the model implements this process through residual loss and discriminator loss. The reconstruction loss measures the difference between the generated sample and the input sample, reflecting the differences in details and structure of the sample; while the discriminator loss measures the distribution difference of the generated sample in the discriminator feature space, helping to determine the similarity between the generated sample and normal traffic video samples. Through these two loss functions, the model can find the generated sample that is closest to the input in the latent space and optimize its generation process.
[0146] After finding the closest latent space point, the model calculates the anomaly score. The anomaly score consists of two parts: first, the reconstruction score L r , that is, the pixel difference between the generated sample and the input sample, indicating the difference in the detail level of the video content; followed by the discriminator score L D , which refers to the difference between the generated samples and normal samples in the discriminator feature space, reflecting the deviation of the overall distribution. By comprehensively evaluating these two scores, the model can determine whether the input video conforms to the distribution of normal traffic scenes. A low anomaly score means that the input video conforms to the normal distribution and is likely to be normal. A high score indicates that there is a significant deviation between the input sample and the normal distribution, indicating the presence of an abnormal event.
[0147]
[0148] L D =logD(x)+log(1-D(x))
[0149] Based on the anomaly score, the model further determines whether the input video is anomalous. A high anomaly score indicates a significant difference between the input sample and the generated sample, and the model will flag the video as anomalous. Furthermore, by analyzing the reconstructed sample (i.e., the difference between the input and generated samples), the model can precisely locate the area where the anomaly occurred, providing more detailed anomaly detection information and further supporting the location and analysis of abnormal events.
[0150] It should be noted that the method of the embodiments of the present disclosure can be performed by a single device, such as a computer or server. The method of the embodiments of the present disclosure can also be applied in a distributed scenario, where multiple devices cooperate to perform the method. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiments of the present disclosure, and the multiple devices will interact with each other to complete the method.
[0151] It should be noted that the above describes some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, it should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention. The actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-tasking and parallel processing are also possible or may be advantageous.
[0152] Example 2
[0153] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present invention further provides a traffic video anomaly detection system based on AIGC, the system being used to implement any of the methods described above, the system comprising: a data set construction module, an extractor construction module, a feature extraction module, a feature reconstruction module, a feature classification module, and an anomaly detection module;
[0154] The dataset construction module is used to construct a traffic video dataset and perform preprocessing;
[0155] The extractor building module is used to build the feature extractor;
[0156] The feature extraction module is used to extract features from the pre-processed traffic video data set using the feature extractor to generate a randomized feature vector;
[0157] The feature reconstruction module is used to perform feature reconstruction on the randomized feature vector using a generator;
[0158] The feature classification module is used to input the reconstructed features into the discriminator for classification and output the probability value;
[0159] The anomaly detection module is used to obtain the traffic video anomaly event score using the output probability value of the discriminator and realize traffic video anomaly detection.
[0160] In this embodiment, preprocessing the traffic video dataset includes:
[0161] Perform time-series segmentation on the video;
[0162] Use data enhancement technology to enhance the data of the segmented video;
[0163] Divide the data-augmented videos into training sets, validation sets, and test sets;
[0164] The time-series segmentation of the video includes:
[0165] V seg (i)={v i×L×f ,v (i+1)×L×f-1},i=1,2,…,N / (L×f)
[0166] Among them, V seg (i) represents the i-th time segment, v i represents the i-th frame of the video, L is the duration of the segment, f is the video frame rate, and N is the total number of frames in a video;
[0167] Data enhancement for the segmented video includes:
[0168] I bright(x,y)=I(x,y)+δ bright
[0169] I contrast (x,y)=α·(I(x,y)-μ)+μ
[0170] Among them, δ bright is the brightness adjustment amount, α is the contrast coefficient, μ is the average brightness of the image, I bright (x,y) and I contrast (x, y) is the enhanced video frame, and I(x, y) is the video frame image.
[0171] In this embodiment, extracting features from the pre-processed traffic video dataset using a feature extractor includes:
[0172] F i =Mobile-SFM(V seg (i))
[0173] Among them, F i Represents the feature extraction model network from the input data V seg (i) The extracted feature vectors.
[0174] In this embodiment, using a generator to perform feature reconstruction on a randomized feature vector includes:
[0175] The default input random feature vector is X, and the output of the generator is X;
[0176] The generator consists of a series of convolutional layers, activation functions, and batch normalization. The output of each layer is:
[0177] x i =σ(BN(Conv3×3(x i-1 )))
[0178] Where x0 = X, the input vector; σ represents the activation function; Conv3×3(·) represents a 3×3 convolution operation; it consists of 5 layers in total, and the activation function of the last layer is tanh, while the activation functions of the remaining layers are all ReLU.
[0179] The output of the generator is judged by introducing a threshold control mechanism. Specifically, the samples reconstructed by the generator are divided, and samples larger than the threshold are input into the discriminator to train the generator. Specifically:
[0180]
[0181] When the reconstruction error Greater than the set threshold When , the generated feature vector is all 1, otherwise it is all 0; the threshold Set from 1 to 3.
[0182] In this embodiment, inputting the reconstructed features into the discriminator for classification includes:
[0183] The discriminator consists of a convolutional layer, an activation function, and a batch normalization layer. The output of each layer is expressed as:
[0184] x i =σ(BN(Conv3×3(x i-1 )))
[0185] x out =σ(Linear(x i ))
[0186] Where x0 = X, i.e., the input sample; σ represents the activation function; Conv3×3(·) represents a 3×3 convolution operation; it consists of 4 layers, including 3 convolutional layers and the last layer is a fully connected layer;
[0187] Discriminator loss function L D The cross entropy loss function is used to distinguish between generated samples and real samples. The expression is:
[0188] L D =logD(x)+log(1-D(x))
[0189] Among them, x represents the real sample, and x represents the sample generated by the generator.
[0190] The system of the above embodiment is used to implement a corresponding AIGC-based traffic video anomaly detection method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0191] It should be noted that the above-mentioned AIGC-based traffic video anomaly detection system is embodied in the form of functional units. The term "module" here can be implemented in the form of software and / or hardware, and is not specifically limited to this.
[0192] For example, a "module" may be a software program, a hardware circuit, or a combination of the two that implements the aforementioned functionality. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (e.g., a shared processor, a dedicated processor, or a group of processors) and memory for executing one or more software or firmware programs, combined logic circuits, and / or other suitable components that support the described functionality.
[0193] The embodiments of the present invention are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention should be included in the scope of protection of this disclosure.
Claims
1. A traffic video anomaly detection method based on AIGC, characterized in that: The method comprises: Build a traffic video dataset and perform preprocessing; Build a feature extractor; Extracting features from the preprocessed traffic video dataset using the feature extractor to generate a randomized feature vector; Reconstructing the randomized feature vector using a generator; The reconstructed features are input into the discriminator for classification and the probability value is output; The output probability value of the discriminator is used to obtain the traffic video abnormal event score and realize traffic video anomaly detection.
2. The method according to claim 1, characterized in that Preprocessing the traffic video dataset includes: Perform time-series segmentation on the video; Use data enhancement technology to enhance the data of the segmented video; Divide the data-augmented videos into training sets, validation sets, and test sets; The time-series segmentation of the video includes: V seg (i)={v i×L×f ,v (i+1)×L×f-1 },i=1,2,…,N / (L×f) Among them, V seg (i) represents the i-th time segment, v i represents the i-th frame of the video, L is the duration of the segment, f is the video frame rate, and N is the total number of frames in a video; Data enhancement for the segmented video includes: I bright (x,y)=I(x,y)+δ bright I contrast (x,y)=α(I(x,y)-μ)+μ Among them, δ bright is the brightness adjustment amount, α is the contrast coefficient, μ is the average brightness of the image, I bright (x,y) and I contrast (x, y) is the enhanced video frame, and I(x, y) is the video frame image.
3. The method according to claim 2, characterized in that Extracting features from the preprocessed traffic video dataset using the feature extractor includes: F i =Mobile-SFM(V seg (i)) Among them, F i Represents the feature extraction model network from the input data V seg (i) The extracted feature vectors.
4. The method according to claim 1, wherein Performing feature reconstruction on the randomized feature vector using a generator includes: The default input random feature vector is X, and the output of the generator is X; The generator consists of a series of convolutional layers, activation functions, and batch normalization. The output of each layer is: x i =σ(BN(Conv3×3(x i-1 ))) Where x0 = X, the input vector; σ represents the activation function; Conv3×3(·) represents a 3×3 convolution operation; it consists of 5 layers in total, and the activation function of the last layer is tanh, while the activation functions of the remaining layers are all ReLU. The output of the generator is judged by introducing a threshold control mechanism. Specifically, the samples reconstructed by the generator are divided, and samples larger than the threshold are input into the discriminator to train the generator. Specifically: When the reconstruction error Greater than the set threshold When , the generated feature vector is all 1, otherwise it is all 0; the threshold Set from 1 to 3.
5. The method according to claim 1, wherein Inputting the reconstructed features into the discriminator for classification includes: The discriminator consists of a convolutional layer, an activation function, and a batch normalization layer. The output of each layer is expressed as: x i =σ(BN(Conv3×3(x i-1 ))) x out =σ(Linear(x i )) Where x0 = X, i.e., the input sample; σ represents the activation function; Conv3×3(·) represents a 3×3 convolution operation; it consists of 4 layers, including 3 convolutional layers and the last layer is a fully connected layer; Discriminator loss function L D The cross entropy loss function is used to distinguish between generated samples and real samples. The expression is: L D =logD(x)+log(1-D(x)) Among them, x represents the real sample, and x represents the sample generated by the generator.
6. A traffic video anomaly detection system based on AIGC, the system being used to implement the method according to any one of claims 1 to 5, characterized in that: The system includes: a data set construction module, an extractor construction module, a feature extraction module, a feature reconstruction module, a feature classification module and an anomaly detection module; The data set construction module is used to construct a traffic video data set and perform preprocessing; The extractor construction module is used to construct a feature extractor; The feature extraction module is used to extract features from the pre-processed traffic video data set using the feature extractor to generate a randomized feature vector; The feature reconstruction module is used to perform feature reconstruction on the randomized feature vector using a generator; The feature classification module is used to input the reconstructed features into the discriminator for classification and output a probability value; The anomaly detection module is used to obtain the traffic video anomaly event score by using the output probability value of the discriminator, thereby realizing traffic video anomaly detection.
7. The system according to claim 6, characterized in that Preprocessing the traffic video dataset includes: Perform time-series segmentation on the video; Use data enhancement technology to enhance the data of the segmented video; Divide the data-augmented videos into training sets, validation sets, and test sets; The time-series segmentation of the video includes: V seg (i)={v i×L×f ,v (i+1)×L×f-1 },i=1,2,…,N / (L×f) Among them, V seg (i) represents the i-th time segment, v i represents the i-th frame of the video, L is the duration of the segment, f is the video frame rate, and N is the total number of frames in a video; Data enhancement for the segmented video includes: I bright (x,y)=I(x,y)+δ bright I contrast (x,y)=α(I(x,y)-μ)+μ Among them, δ bright is the brightness adjustment amount, α is the contrast coefficient, μ is the average brightness of the image, I bright (x,y) and I contrast (x, y) is the enhanced video frame, and I(x, y) is the video frame image.
8. The system according to claim 7, characterized in that Extracting features from the preprocessed traffic video dataset using the feature extractor includes: F i =Mobile-SFM(V seg (i)) Among them, F i Represents the feature extraction model network from the input data V seg (i) The extracted feature vectors.
9. The system according to claim 6, wherein: Performing feature reconstruction on the randomized feature vector using a generator includes: The default input random feature vector is X, and the output of the generator is X; The generator consists of a series of convolutional layers, activation functions, and batch normalization. The output of each layer is: x i =σ(BN(Conv3×3(x i-1 ))) Where x0 = X, the input vector; σ represents the activation function; Conv3×3(·) represents a 3×3 convolution operation; it consists of 5 layers in total, and the activation function of the last layer is tanh, while the activation functions of the remaining layers are all ReLU. The output of the generator is judged by introducing a threshold control mechanism. Specifically, the samples reconstructed by the generator are divided, and samples larger than the threshold are input into the discriminator to train the generator. Specifically: When the reconstruction error Greater than the set threshold When , the generated feature vector is all 1, otherwise it is all 0; the threshold Set from 1 to 3.
10. The system according to claim 6, wherein: Inputting the reconstructed features into the discriminator for classification includes: The discriminator consists of a convolutional layer, an activation function, and a batch normalization layer. The output of each layer is expressed as: x i =σ(BN(Conv3×3(x i-1 ))) x out =σ(Linear(x i )) Where x0 = X, i.e., the input sample; σ represents the activation function; Conv3×3(·) represents a 3×3 convolution operation; it consists of 4 layers, including 3 convolutional layers and the last layer is a fully connected layer; Discriminator loss function L D The cross entropy loss function is used to distinguish between generated samples and real samples. The expression is: L D =logD(x)+log(1-D(x)) Among them, x represents the real sample, and x represents the sample generated by the generator.