Video anomaly detection method, pre-training method and system of feature extractor
The method enhances video anomaly detection in fixed-view angle scenarios by using frame difference and depth estimation to guide selective masking, improving feature extraction and detection performance.
Patent Information
- Application Number
- CN202411764772.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-12-03
AI Technical Summary
The prior art performs poorly in fixed-view videos, making it difficult to adapt to specific scenarios such as surveillance videos, and there are domain differences problems when training existing network models, resulting in low abnormal detection accuracy and efficiency.
A visual mask modeling method based on space-time cues guidance is adopted to generate a mask probability map through inter-frame difference and depth estimation, selectively mask the video block, and optimize the feature extractor with perspective reconstruction loss function to improve the model's learning ability of key areas.
The abnormal detection accuracy and efficiency of fixed-view videos are improved, and the abnormal detection performance in surveillance videos is significantly improved.
Smart Images

Figure CN119600516B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video anomaly detection, and in particular to a video anomaly detection method, a pre-training method for a feature extractor, and a system. Background Art
[0002] Video Anomaly Detection (VAD) plays a crucial role in large-scale monitoring systems, and excellent visual features are the key to ensuring high accuracy of the anomaly detection network. Currently, most studies still rely on networks based on 3D convolution such as I3D and C3D to extract features on general behavior recognition datasets. The domain difference between the general behavior recognition datasets (such as Kinetics400) used for training I3D and C3D networks and the datasets of fixed-view videos makes these networks perform well in general videos, but their performance in special videos such as fixed-view videos, for example, surveillance videos, is not satisfactory.
[0003] Since the attention-based backbone network Transformer was proposed, research results on Masked Visual Modeling (MVM) have emerged continuously. Masked Visual Modeling adopts a simple masking strategy and a pixel reconstruction task, and can learn general visual representations in large-scale unlabeled datasets, with strong visual feature extraction capabilities. When similar to the above-mentioned networks based on 3D convolution, although it performs better in general videos, it is difficult to show good adaptability in fixed-view videos. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a video anomaly detection method, a pre-training method for a feature extractor, and a system, so as to solve the problems existing in the prior art.
[0005] To achieve the foregoing invention purpose, the technical solutions adopted by the present invention include:
[0006] In a first aspect, the present invention provides a pre-training method for a feature extractor for video anomaly detection, where the feature extractor is used for anomaly detection of fixed-view videos, including:
[0007] Dividing the training video into blocks and performing mask covering, covering a part of the video blocks and retaining another part of the video blocks as input video blocks;
[0008] Using an encoder to learn the feature representation of the input video blocks;
[0009] The decoder is used to reconstruct the masked video blocks based on the feature representation, and update the parameters of the encoder based on the similarity between the reconstructed video blocks and the masked video blocks, and finally serve as the feature extractor;
[0010] The process of mask covering specifically includes:
[0011] Obtain the frame difference map of the training video using inter-frame difference, and obtain the depth map of the training video using depth estimation;
[0012] Form a temporal probability map using the frame difference map, form a spatial probability map using the depth map, and fuse the spatial probability map and the temporal probability map to form a mask probability map;
[0013] Select and form a mask map based on the mask probability map, and selectively mask the video blocks based on the mask map;
[0014] Wherein, when forming the mask map, the probability that any video block is selected for masking is positively correlated with the degree of motion in the corresponding temporal probability map and negatively correlated with the depth value in the spatial probability map.
[0015] In a second aspect, the present invention also provides an abnormal detection method for videos with a fixed perspective, which includes:
[0016] Use the feature extractor obtained by pre-training using the above pre-training method to extract features from the video to be detected to obtain video features;
[0017] Use a scorer to evaluate the abnormality of the video features to obtain an abnormal detection result.
[0018] In a third aspect, the present invention also provides a pre-training system for a feature extractor for video anomaly detection, where the feature extractor is used for anomaly detection of videos with a fixed perspective, and includes:
[0019] A mask module, configured to divide the training video into blocks and perform mask covering, covering a part of the video blocks and retaining another part of the video blocks as input video blocks;
[0020] An encoder module, configured to use an encoder to learn the feature representation of the input video blocks;
[0021] A decoder module, configured to use a decoder to reconstruct the masked video blocks based on the feature representation, and update the parameters of the encoder based on the similarity between the reconstructed video blocks and the masked video blocks, and finally serve as the feature extractor;
[0022] The mask module specifically includes:
[0023] A differential depth unit is used to obtain a frame difference map of the training video by using inter-frame difference and obtain a depth map of the training video by using depth estimation;
[0024] A probability map unit is used to form a temporal probability map by using the frame difference map, form a spatial probability map by using the depth map, and fuse the spatial probability map and the temporal probability map to form a mask probability map;
[0025] A selection masking unit is used to select and form a mask map based on the mask probability map and perform selective masking on the video block based on the mask map;
[0026] Wherein, when forming the mask map, the probability that any of the video blocks is selected for masking is positively correlated with the degree of motion in the corresponding temporal probability map and negatively correlated with the depth value in the spatial probability map.
[0027] In a fourth aspect, the present invention further provides a readable storage medium in which a computer program is stored, and when the computer program is run, the steps of the above-mentioned pre-training method or video anomaly detection method are executed.
[0028] Based on the above technical solutions, compared with the prior art, the beneficial effects of the present invention at least include:
[0029] The pre-training method provided by the present invention can provide a specific visual feature extractor for video anomaly detection of a fixed perspective, solve the domain difference problem between the general behavior recognition data set used in the training of the current network model and the fixed perspective video, and improve the accuracy and efficiency of anomaly detection of fixed perspective videos by extracting feature representations customized for video anomaly detection.
[0030] The above description is only an overview of the technical solutions of the present invention. In order to enable those skilled in the art to understand the technical means of the present application more clearly and implement it in accordance with the content of the specification, the following will describe the preferred embodiments of the present invention in conjunction with detailed drawings as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a schematic diagram of the process of video anomaly detection provided by a typical embodiment of the present invention. DETAILED DESCRIPTION
[0032] In view of the deficiencies in the prior art, the inventors of this case have proposed the technical solutions of the present invention through long-term research and a large number of practices. The following will further explain the technical solutions, their implementation processes and principles.
[0033] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways than those described herein, and thus, the scope of the present invention is not limited by the specific embodiments disclosed below.
[0034] Moreover, relational terms such as "first" and "second" are only used to distinguish one component or method step with the same name from another, and do not necessarily require or imply any actual relationship or order between these components or method steps.
[0035] See Figure 1 As shown, an embodiment of the present invention provides a pre-training method for a feature extractor for video anomaly detection. The feature extractor is used for anomaly detection of videos with a fixed perspective, and the general process is the same as that of the existing self-attention-based backbone network Transformer, including the following main steps:
[0036] Divide the training video into chunks and perform mask covering, covering a part of the video chunks and retaining another part of the video chunks as input video chunks;
[0037] Use the encoder to learn the feature representation of the input video chunks;
[0038] Use the decoder to reconstruct the masked video chunks based on the feature representation, and update the parameters of the encoder based on the similarity between the reconstructed video chunks and the masked video chunks, and finally use it as the feature extractor.
[0039] In the specific embodiments provided by the present invention, VideoMAE is used as the basic architecture of the visual feature extractor. VideoMAE uses self-supervised learning to obtain effective visual feature representations. One of its core designs is the mask mechanism, which includes two modules: an encoder and a decoder. The structures of both modules are composed of traditional Transformer blocks. The encoder is responsible for learning the feature representation from the input video chunks, and the decoder is used to reconstruct the masked video chunks.
[0040] During the training process of this model, the random mask strategy is a commonly used mask method. It randomly masks the input video chunks, and each video chunk has the same mask probability. However, the inventors of the present invention have found that this strategy may be simple and effective in general video recognition datasets, but it is less applicable in videos with a fixed perspective, such as video surveillance datasets.
[0041] This is because compared with general action recognition datasets, moving objects in surveillance scenarios are usually smaller, and there is a significant distinction between the foreground and the background. The movement range of the objects often lies within a specific area. Therefore, the random mask strategy may obscure the key objects in the surveillance scenario, resulting in the loss of important information and preventing the model from learning relevant context information. To address this issue, the present invention proposes a visual mask modeling based on spatio-temporal cue guidance for video surveillance scenarios.
[0042] Thus, there are key differences between the present invention and the above-mentioned existing network models. The process of mask covering specifically includes:
[0043] Obtain the frame difference map of the training video using frame difference, and obtain the depth map of the training video using depth estimation;
[0044] Form a temporal probability map using the frame difference map, form a spatial probability map using the depth map, and fuse the spatial probability map and the temporal probability map to form a mask probability map;
[0045] Selectively form a mask map based on the mask probability map, and selectively cover the video blocks based on the mask map;
[0046] Among them, when forming the mask map, the probability of any video block being selected for covering is positively correlated with the degree of movement in the corresponding temporal probability map and negatively correlated with the depth value in the spatial probability map.
[0047] Based on the observation and analysis of the fixed-view dataset, the inventors obtained two important characteristics of fixed-view videos: (1) objects in fixed-view videos have the characteristic of being larger when closer and smaller when farther; (2) usually only some areas in the fixed-view scene have movement changes. Based on these characteristics, the spatio-temporal semantic information (frame difference and depth information) of the video data was respectively extracted to guide the selection of the mask area.
[0048] The target task in the pre-training stage of the visual feature extractor is pixel reconstruction. To emphasize the difference in the contribution degree of different regions in the video scene to the reconstruction task, depth information is used to assign corresponding weights to the reconstruction losses of different regions, so as to enhance the expression ability of the visual extractor and enable it to better capture the visual features of key regions.
[0049] In the final application, in the abnormal detection feature extraction stage, the pre-trained visual feature extractor is applied to the video abnormal detection dataset, and the extracted features are input into the video abnormal detection network for abnormal scoring and determination, and thus better detection performance can be achieved compared with the general random mask method.
[0050] Specifically, in some implementation schemes, the formation method of the mask probability map can be expressed as:
[0051] P m = αP s + βP d ;;
[0052] Wherein, P m represents the mask probability map; P s represents the temporal probability map; P d represents the spatial probability map; α and β represent weight coefficients, and the sum of the two is 1.
[0053] In some embodiments, the generation processes of the temporal probability map and the spatial probability map can be expressed as:
[0054]
[0055] Wherein, represents the probability value corresponding to video block i in the temporal probability map; represents the probability value corresponding to the video block in the spatial probability map; s (·) represents the frame difference change amplitude of the video block; d (·) represents the depth value of the video block; i represents the serial number of the video block; HW / N 2 represents the number of video blocks, where H and W are the number of pixels of the video frame height and width respectively. In some typical examples, H = W = 224 and N = 14, but this parameter setting is not limited to this.
[0056] Regarding some specific technical details, in some embodiments, the process of mask covering may further include:
[0057] Based on the frame difference map, distinguish the moving area and the static area; Usually, since the background area remains unchanged and the change amplitude of the inter-frame difference is small, the moving area and the static area can be effectively distinguished by setting a threshold. Of course, other technical means with the same function are not excluded to distinguish the moving area and the static area.
[0058] In the moving area, randomly select a preset proportion of video blocks, and specify the probability that the selected video block is masked as 0 in the mask probability map, and force the retention of the selected video block.
[0059] As some typical examples, for the input original video V ∈ R T×H×W×C , the present invention uses the inter-frame difference method and the monocular depth estimation network (MiDas) to obtain the frame difference map S ∈ R T×H×W and the depth map D ∈ R T×H×W , which respectively represent the temporal information and the spatial information of the video. Next, perform the softmax operation on the spatio-temporal information respectively to obtain the temporal probability map and spatial probability maps The value range of the probability is [0, 1]. These probability maps will be used to guide the selection process of video mask blocks, which represent the mask probability of each input video block. The higher the probability value, the greater the possibility that the video block will be masked; on the contrary, the smaller the probability value, the more likely the video block will be retained.
[0060] In the spatial probability map, the foreground region has a higher mask probability. The same-sized video blocks of the same target in the foreground and background regions show different density information; for the same-sized video blocks, the foreground region usually contains higher redundant information, while the background region contains more complex and high-density information.
[0061] In the temporal probability map, the moving region has a higher mask probability. The reconstruction of the moving region by the decoder enables the model to better learn the feature representation of the moving target. However, in order to prevent the model from being unable to obtain the context information of the moving target due to the complete masking of the moving region, in the preferred embodiment of the present invention, about 25% of the moving regions are randomly selected, and the probability value of their being masked is set to 0 to retain some moving regions for the model to learn and prevent a significant decrease in the model training efficiency.
[0062] Specifically, in some embodiments, the preset ratio can generally be 10 - 40%. In the embodiments provided by the present invention, this value can be set to 25%, but it does not mean that adaptive adjustment cannot be performed. For different data sets and different base models, the optimal value range may vary.
[0063] The final mask probability map will be obtained by the weighted sum of the temporal probability map and the spatial probability map. Through the comparative analysis of the mask visualization effect, in the general method settings, it is usually selected that α = 0.7, β = 0.3 as a better setting, and this setting is more in line with the expected effect of model training in this scenario.
[0064] Specifically, in some embodiments, the value range of α is usually 0.2 - 0.4. For example, in a typical embodiment, it is set to about 0.3.
[0065] The above presents the generation details of spatio-temporal probability information. Regarding the process of forming a mask map based on probability information, in some embodiments, the mask map formation process uses the roulette selection algorithm for multiple rounds of selection. For any one of the video blocks, the probability of being selected in the current round is related to the probability value in the corresponding mask probability map until the proportion of the number of selected video blocks reaches the preset mask rate.
[0066] Specifically, in some embodiments, the calculation method of the probability of the video block being selected is expressed as:
[0067]
[0068] Among them, p(xi) represents the selection probability of video block i; q(x i ) represents the cumulative probability; P(x.) represents the probability value corresponding to the corresponding video block in the mask probability map; j represents the serial number of the video block; the rest of the symbols have been clarified in the above formula and will not be elaborated here.
[0069] The formation process of the mask graph specifically includes being carried out in multiple repeated rounds:
[0070] Generate a random number r between [0, 1];
[0071] If r < q(x i ), then select video block i to be selected;
[0072] Otherwise, jump to video block i + 1 for the next round of selection until a video block k is selected such that: q[Xk - 1] < r ≤ q[x k holds; where k - 1 represents a video block adjacent to video block k and with a cumulative probability value less than that of video block k after sorting according to the cumulative probability, and k represents that after selecting the kth video block, its cumulative probability value is greater than or equal to the random number r.
[0073] As a typical example, in the above process, according to the mask probability map representing all video blocks, the roulette wheel selection algorithm is used to select the mask blocks. The roulette wheel algorithm is a selection strategy in genetic algorithms, and its core idea is that the probability of an individual being selected is proportional to its fitness value. In the process of selecting mask blocks, the present invention regards the probability value in the mask probability map as the fitness of each video block, calculates the selection probability of each video block accordingly, and then calculates the cumulative probability until the number of selected video blocks reaches the number of mask blocks required by the preset mask rate (such as 90%, and of course, it can also be appropriately adjusted up and down according to training requirements, such as 60 - 95%). Then, the encoder learns the feature representation from the remaining video blocks, and the decoder reconstructs and predicts the masked video blocks. Compared with the traditional random masking strategy, the probability masking strategy based on spatio - temporal clue guidance proposed by the present invention can enable the model to pay more attention to the motion regions and long - view regions rich in semantic information, achieving significantly different technical effects.
[0074] Regarding the loss function, in some embodiments, the loss function of the training method includes a perspective reconstruction loss, and the perspective reconstruction loss uses the depth value in the depth map as the weight of the corresponding video block, calculates the difference between the reconstruction result of the video block and the ground truth by weighting and sums it up, where the weight of the video block is positively correlated with the depth value.
[0075] More specifically, in some embodiments, the perspective loss function is expressed as:
[0076]
[0077] Among them, L rec represents the perspective loss function; ψ(·) represents the mapping function of the depth value, and the value size is positively correlated with the depth value. In a typical representative case, it corresponds to Figure 1 the "loss window" in represents the reconstructed video block; I k,t represents the masked video block; k represents the serial number of the video block; t represents the serial number of the video frame. The rest of the symbols have been clearly defined in the above formula and will not be elaborated here.
[0078] Among them, the above mapping function can be a positively correlated continuous function or a piecewise function that is positively correlated and takes fixed values in segments. Specifically, for example, Figure 1 as shown in
[0079] The above solution proposes a Perspective Reconstruction Loss based on the depth information of the video. This loss uses the mean square error to measure the reconstruction quality of the model for the masked area, and at the same time introduces depth information to emphasize the contribution degree difference of different regions in the surveillance video scene to the reconstruction task, thereby promoting the model to learn a better feature representation.
[0080] This weight assigns a higher weight value to the far-view area, thereby increasing the learning difficulty of the model for the far-view area target and enhancing the learning ability of the model for the far-view area target containing rich semantic information. Compared with the reconstruction loss that has the same attention to each area, the perspective reconstruction loss can guide the model to pay more attention to specific areas or targets, thereby improving the reconstruction quality and feature representation ability.
[0081] As a further application of the above technical solution, the second aspect of the embodiments of the present invention also provides an abnormal detection method for videos with a fixed perspective, which includes the following steps:
[0082] Use the feature extractor obtained by pre-training with the pre-training method provided in any of the above embodiments to perform feature extraction on the video to be detected to obtain video features;
[0083] Use a scorer to perform abnormal evaluation on the video features to obtain an abnormal detection result.
[0084] In the feature extraction process of anomaly detection, typical implementation cases of the present invention use the above-mentioned pre-trained model as a visual feature extractor to replace traditional feature extractors such as I3D and C3D, and extract features from the anomaly detection data set. The extracted features can then be used in various anomaly detection networks (such as TEVAD) for anomaly prediction and scoring.
[0085] Correspondingly, an embodiment of the present invention further provides a pre-training system for a feature extractor of video anomaly detection. The feature extractor is used for anomaly detection of videos with a fixed perspective and includes:
[0086] A mask module for dividing the training video into blocks and performing mask covering, covering a part of the video blocks and retaining another part of the video blocks as input video blocks;
[0087] An encoder module for using an encoder to learn the feature representation of the input video blocks;
[0088] A decoder module for using a decoder to reconstruct the masked video blocks based on the feature representation, and updating the parameters of the encoder based on the similarity between the reconstructed video blocks and the masked video blocks, and finally serving as the feature extractor;
[0089] Similarly corresponding to the training system of the existing network structure, the specific content of the mask module includes:
[0090] A differential depth unit for obtaining a frame difference map of the training video using inter-frame difference and obtaining a depth map of the training video using depth estimation;
[0091] A probability map unit for forming a temporal probability map using the frame difference map, forming a spatial probability map using the depth map, and fusing the spatial probability map and the temporal probability map to form a mask probability map;
[0092] A selective masking unit for selecting a mask map based on the mask probability map and selectively masking the video blocks based on the mask map;
[0093] Wherein, when forming the mask map, the probability of any video block being selectively masked is positively correlated with the degree of motion in the corresponding temporal probability map and negatively correlated with the depth value in the spatial probability map.
[0094] In addition, an embodiment of the present invention further provides a readable storage medium storing a computer program, and when the computer program is run, it executes the steps of the above-mentioned pre-training method or video anomaly detection method.
[0095] The technical solutions of the present invention are further described in detail below through several embodiments in conjunction with the accompanying drawings. However, the selected embodiments are only used to illustrate the present invention and do not limit the scope of the present invention.
[0096] Example 1
[0097] In this example, the process of the pre-training method shown above is applied to the pre-training of the feature extractor and the anomaly detection of the surveillance video, as specifically shown below.
[0098] Experimental settings: In the pre-training stage of the model, ViT-B is used as the default architecture of the encoder, and the decoder has the same architecture. Train for 800 rounds on the surveillance dataset, with a batch size of 512, and use 128 Ascend 910N PUs. The AdamW optimizer and the cosine learning rate scheduler are used, with an initial learning rate of 1.5e-4, a weight decay of 0.05, β1 of 0.9, β2 of 0.999, and a warm-up period of 40 rounds. During the mask generation process, a mask ratio of 90% is applied. In the anomaly detection stage, the pre-trained model is used as the video encoder.
[0099] The training data uses a self-built surveillance dataset, and actual detection is performed on the anomaly detection datasets UCF-Crime and TAD. The area under the curve (AUC) of the frame-level receiver operating characteristic curve (ROC) is used to evaluate the performance of the method adopted in this example on the UCF-Crime and TAD benchmarks.
[0100] Table 1 Comparison results of method performance on UCF-Crime and TAD datasets
[0101]
[0102] Table 1 shows the performance comparison between this method and other methods. Among them, an AUC performance of 86.0% is achieved on the UCF-Crime dataset, which is 1.7% to 3.9% higher than other methods; a performance of 90.3% is achieved on the TAD dataset.
[0103] Comparative Example 1
[0104] This comparative example is generally the same as Example 1, and the main differences are as follows:
[0105] Instead of performing statistical analysis of spatio-temporal information, a random masking strategy is used to mask video blocks, and the loss function does not involve the weight term related to the depth value (loss window).
[0106] To fully prove the effectiveness of each module proposed in this method, we conducted ablation experiments on the UCF-Crime and TAD datasets. The specific settings are to adopt a random masking strategy, only utilize temporal information, and a masking strategy guided by the fusion of temporal and spatial information.
[0107] Table 2 Results of ablation experiments on the masking strategy
[0108]
[0109] Table 2 shows the ablation experiment comparison on the masking strategy. When adopting the random masking strategy, the AUC performances of UCF-Crime and TAD are 82.6% and 88.1% respectively. When only using the masking strategy guided by time information, the performances of both datasets are improved by 1.0%; while when fusing the time and space information to guide the masking strategy, the performance of UCF-Crime is further improved, and the performance of TAD slightly decreases compared with the random masking strategy.
[0110] Comparative Example 2
[0111] This comparative example is generally the same as Example 1, and the main differences are as follows:
[0112] Delete the weight term (loss window) related to the depth value in the loss function.
[0113] Ablation experiment of the loss function in Table 3
[0114]
[0115] Table 3 shows the ablation experiment on the proposed perspective reconstruction loss. The results show that after adopting the improved perspective reconstruction loss, the performances of UCF-Crime and TAD are both significantly improved, fully indicating that the control of the loss weight is very effective for the video encoder to learn better feature representations.
[0116] Based on the above embodiments and comparative examples, it can be clearly seen that the pre-training method provided by the embodiments of the present invention can provide a specific visual feature extractor for this type of video anomaly detection with a fixed perspective, solve the domain difference problem between the general behavior recognition dataset used in the current network model training and the fixed-perspective video, and improve the accuracy and efficiency of the anomaly detection of the fixed-perspective video by extracting the feature representation customized for video anomaly detection.
[0117] It should be understood that the above embodiments are only used to illustrate the technical concept and characteristics of the present invention, and their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A pre-training method for a feature extractor for video anomaly detection, where the feature extractor is used for anomaly detection of videos with a fixed perspective, including: Dividing the training video into blocks and performing mask covering, covering a part of the video blocks and retaining another part of the video blocks as input video blocks; Using an encoder to learn the feature representation of the input video blocks; Using a decoder to reconstruct the masked video blocks based on the feature representation, and updating the parameters of the encoder based on the similarity between the reconstructed video blocks and the masked video blocks, and finally serving as the feature extractor; It is characterized in that the process of mask covering specifically includes: Using inter-frame difference to obtain the frame difference map of the training video, and using depth estimation to obtain the depth map of the training video; Using the frame difference map to form a temporal probability map, using the depth map to form a spatial probability map, and fusing the spatial probability map and the temporal probability map to form a mask probability map; Based on the mask probability map, a mask map is selected to be formed, and the video blocks are selectively masked based on the mask map; Among them, when forming the mask map, the probability that any video block is selected to be masked is positively correlated with the degree of motion in the corresponding temporal probability map and negatively correlated with the depth value in the spatial probability map.
2. The pre-training method according to claim 1, wherein The formation method of the mask probability map is expressed as: P m = αP s + βP d ; Among them, P m represents the mask probability map; P s represents the temporal probability map; P d represents the spatial probability map; α and β represent weight coefficients, and the sum of the two is 1.
3. The pre-training method according to claim 2, characterized in that, The generation processes of the temporal probability map and the spatial probability map are expressed as: Among them, represents the probability value corresponding to video block i in the time probability map; represents the probability value corresponding to video block i in the spatial probability map; s (·) represents the frame difference change amplitude of the video block; d (·) represents the depth value of the video block; i represents the serial number of the video block; HW / N 2 represents the number of video blocks, where H and W are the number of pixels of the video frame height and width respectively.
4. The pre-training method according to any one of claims 1-3, characterized in that The process of mask covering further includes: Based on the frame difference map, differentiating the moving regions and the stationary regions; In the moving regions, a preset proportion of video blocks are randomly selected, and the probability that the selected video blocks are masked is specified as 0 in the mask probability map, and the selected video blocks are forcibly retained.
5. The pre-training method according to claim 1, wherein The mask map formation process uses the roulette wheel selection algorithm for multiple rounds of selection. For any video block, the probability of being selected in the current round is related to the probability value in the corresponding mask probability map until the proportion of the number of selected video blocks reaches the preset mask rate.
6. The pre-training method according to claim 5, wherein The calculation method of the probability that the video block is selected is expressed as: Among them, p(xi) represents the selection probability of video block i; q(x i ) represents the cumulative probability; P(x.) represents the probability value corresponding to the corresponding video block in the mask probability map; j represents the serial number of the video block; The formation process of the mask map specifically includes being carried out in multiple rounds of repetition: Generating a random number r between [0, 1]; If r < q(x i ), then the video block i is selected; Otherwise, jump to the (i + 1)-th video block for the next round of selection until a video block k is selected such that: q[xk - 1] < r ≤ q[x k holds; where k - 1 represents a video block adjacent to video block k with a cumulative probability value less than that of video block k after sorting by cumulative probability, and k represents that after selecting the k-th video block, its cumulative probability value is greater than or equal to the random number r.
7. The pre-training method according to claim 1, wherein The loss function of the training method includes a perspective reconstruction loss. The perspective reconstruction loss uses the depth value in the depth map as the weight of the corresponding video block, calculates the difference between the reconstruction result of the video block and the ground truth in a weighted manner and summarizes it, where the weight of the video block is positively correlated with the depth value.
8. The pre-training method according to claim 7, wherein The perspective loss function is expressed as: Among them, L rec represents the perspective loss function; ψ(·) represents the mapping function of the depth value, and the magnitude of the value is positively correlated with the depth value; represents the reconstructed video block; I k,t represents the masked video block; k represents the serial number of the video block; t represents the serial number of the video frame.
9. An abnormal detection method for videos with a fixed viewing angle, characterized in that, Including: Performing feature extraction on the video to be detected using the feature extractor pre-trained by the training method described in any one of claims 1-8 to obtain video features; Using a scorer to perform anomaly evaluation on the video features to obtain an anomaly detection result.
10. A pre-training system for a feature extractor for video anomaly detection, where the feature extractor is used for anomaly detection of videos with a fixed perspective, including: A mask module for dividing the training video into blocks and performing mask covering, covering a part of the video blocks and retaining another part of the video blocks as input video blocks; An encoder module for using an encoder to learn the feature representation of the input video blocks; A decoder module, which is used to reconstruct the masked video block based on the feature representation by using a decoder, and update the parameters of the encoder based on the similarity between the reconstructed video block and the masked video block, and finally serve as the feature extractor; It is characterized in that the mask module specifically includes: A differential depth unit, which is used to obtain the frame difference map of the training video by using inter-frame difference, and obtain the depth map of the training video by using depth estimation; A probability map unit, which is used to form a temporal probability map by using the frame difference map, form a spatial probability map by using the depth map, and fuse the spatial probability map and the temporal probability map to form a mask probability map; A selective masking unit, which is used to select and form a mask map based on the mask probability map, and selectively mask the video block based on the mask map; Among them, when forming the mask map, the probability that any of the video blocks is selected for masking is positively correlated with the degree of motion in the corresponding temporal probability map and negatively correlated with the depth value in the spatial probability map.
Citation Information
Patent Citations
Pre-training-based MOOC self-adaptive learning system construction method and pre-training-based MOOC self-adaptive learning system construction device
CN114567815A
Visual target tracking method based on mask contrast learning pre-training
CN118674749A