Video anomaly detection model training method and device based on contrast learning and video anomaly detection method and device

By adopting a training method based on comparison learning in the video anomaly detection model, negative samples and positive samples feature groups are generated and model parameters are adjusted, the problem of insufficient modeling ability of abnormal samples in the existing technology is solved, and the model's distinction ability and recognition ability of normal videos are improved.

CN120198841AActive Publication Date: 2025-06-24ARTIFICIAL INTELLIGENCE RES INST OF HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ARTIFICIAL INTELLIGENCE LAB)
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510690582.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-24
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

In the prior art, memory-enhanced autoencoder lacks the ability to model abnormal samples in video anomaly detection, and cannot effectively learn the boundary between normal samples and abnormal samples, affecting the model's distinction ability.

Method used

The training method of the video anomaly detection model based on contrast learning is adopted. By inputting the sample video features into the feature extraction and reconstruction module of the initial video anomaly detection model, negative sample video feature groups and positive sample video feature groups are generated, and model parameters are adjusted using contrast learning loss function and reconstruction loss function to enhance the learning ability of the memory module.

Benefits of technology

The video anomaly detection model's ability to identify normal videos is improved, the boundaries between positive and negative samples are clarified, the model's distinction ability is enhanced, and the difference between normal and abnormal samples is learned without supervision, reducing the dependence on labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198841A_ABST
    Figure CN120198841A_ABST
Patent Text Reader

Abstract

The invention provides a video anomaly detection model training method and device based on comparative learning and a video anomaly detection method and device, which can be applied to the technical field of deep reinforcement learning. The training method comprises the following steps: inputting a sample video into a feature extraction module to obtain a sample video feature; inputting the sample video features into a reconstruction module to obtain a reconstructed sample video, and based on a reconstruction loss function, obtaining reconstruction loss according to the sample video and the reconstructed sample video; generating a negative sample video feature group based on the sample video features and the reconstruction loss; performing data enhancement processing on the sample video features to obtain a positive sample video feature group; based on a comparative learning loss function, obtaining comparative learning loss according to the negative sample video feature group and the positive sample video feature group; the parameters of the initial video anomaly detection model are adjusted based on the comparison learning loss and the reconstruction loss, the target video detection model is obtained, the recognition capability of normal videos is improved, and video anomaly detection is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep reinforcement learning, and more particularly to a training method and device for a video anomaly detection model based on contrast learning, and a video anomaly detection method and device. Background Art

[0002] Video anomaly detection is an important computer vision task, which is widely used in fields such as security monitoring, behavior recognition, and industrial inspection. Its goal is to automatically detect anomaly events that are significantly different from normal patterns from a video stream.

[0003] In the prior art, video anomalies are generally detected by a memory-augmented autoencoder. The memory-augmented autoencoder stores the features of the normal pattern through a memory module, thereby improving the reconstruction ability of the target video anomaly detection model for the normal pattern. However, the feature update of the memory module only depends on normal samples, lacks the ability to model anomaly samples, and cannot effectively learn the boundary between normal samples and anomaly samples, thus affecting the discrimination ability of the model. Summary of the Invention

[0004] In view of the above problems, the present invention provides a training method for a video anomaly detection model based on contrast learning and a video anomaly detection method.

[0005] According to a first aspect of the present invention, there is provided a training method for a video anomaly detection model based on contrast learning, including: inputting a sample video into a feature extraction module of an initial video anomaly detection model to obtain sample video features; inputting the sample video features into a reconstruction module of the initial video anomaly detection model to obtain a reconstructed sample video, and based on a reconstruction loss function, obtaining a reconstruction loss according to the sample video and the reconstructed sample video; generating a negative sample video feature group based on the sample video features and the reconstruction loss; performing data augmentation processing on the sample video features to obtain a positive sample video feature group; obtaining a contrast learning loss based on a contrast learning loss function according to the negative sample video feature group and the positive sample video feature group; and adjusting parameters of the initial video anomaly detection model based on the contrast learning loss and the reconstruction loss to obtain a target video detection model.

[0006] According to an embodiment of the present invention, the above-mentioned contrast learning loss function obtains a contrast learning loss according to the above-mentioned negative sample video feature group and positive sample video feature group, including: inputting the above-mentioned negative sample video feature group and the above-mentioned positive sample video feature group into the memory module of the above-mentioned initial video anomaly detection model respectively to obtain a negative sample similarity and a positive sample similarity, where the negative sample similarity represents the degree of similarity between the above-mentioned negative sample video feature group and the memory vector feature group stored in the above-mentioned memory module, and the positive sample similarity represents the degree of similarity between the above-mentioned positive sample video feature group and the memory vector feature group; based on the contrast learning loss function, the above-mentioned contrast learning loss is obtained according to the above-mentioned negative sample similarity and the above-mentioned positive sample similarity.

[0007] According to an embodiment of the present invention, the above-mentioned generation of the negative sample video feature group based on the above-mentioned sample video feature and the above-mentioned reconstruction loss includes: calculating a gradient vector of the above-mentioned reconstruction loss with respect to the above-mentioned sample video feature based on a gradient extraction function, where the gradient vector represents the sensitivity of the above-mentioned sample video feature to the reconstruction error; generating the above-mentioned negative sample video feature group based on the gradient direction information in the above-mentioned gradient vector and the above-mentioned sample video feature.

[0008] According to an embodiment of the present invention, the above-mentioned data augmentation processing of the above-mentioned sample video feature to obtain a positive sample video feature group includes: generating a first noise based on a first preset noise range, where the first noise represents the change of environmental factors in the sample video; generating a second noise based on a second preset noise range, where the second noise represents the visual interference factors in the sample video; performing proportional scaling on the above-mentioned sample video feature according to the above-mentioned first noise to obtain a scaled positive sample video feature group; performing random offset on the above-mentioned scaled positive sample video feature group according to the above-mentioned second noise to obtain the above-mentioned positive sample video feature group.

[0009] According to an embodiment of the present invention, the above-mentioned memory vector feature group includes P memory vector features, the above-mentioned negative sample video feature group includes M negative sample video features, and the above-mentioned positive sample video feature group includes N positive sample video features. P, M, and N are all positive integers greater than 0. p is a positive integer less than or equal to P and greater than 0, m is a positive integer less than or equal to M and greater than 0, and n is a positive integer less than or equal to N and greater than 0. The above-mentioned steps of inputting the above-mentioned negative sample video feature group and the above-mentioned positive sample video feature group into the memory module of the above-mentioned initial video anomaly detection model to obtain negative sample similarity and positive sample similarity include: for the m-th negative sample video feature in the above-mentioned negative sample video feature group, using the cosine similarity function, based on the m-th negative sample video feature and the p-th above-mentioned memory vector feature, obtain the p-th sub-negative sample cosine similarity to obtain a sub-negative sample cosine similarity group, and the sub-negative sample cosine similarity group includes P sub-negative sample cosine similarities; for the n-th positive sample video feature in the above-mentioned positive sample video feature group, using the above-mentioned cosine similarity function, based on the n-th positive sample video feature and the p-th memory vector feature, obtain the p-th sub-positive sample cosine similarity to obtain a sub-positive sample cosine similarity group, and the sub-positive sample cosine similarity group includes P sub-positive sample cosine similarities; using a preset normalization function, normalize and sum the above-mentioned P sub-negative sample cosine similarities respectively to obtain the m-th negative sample similarity to obtain M negative sample similarities; using a preset normalization function, normalize and sum the P sub-positive sample cosine similarities respectively to obtain the n-th positive sample similarity to obtain N positive sample similarities.

[0010] According to an embodiment of the present invention, the above-mentioned steps of obtaining the above-mentioned contrast learning loss based on the above-mentioned negative sample similarity and the above-mentioned positive sample similarity include: summing the above-mentioned N positive sample similarity indices to obtain a positive sample index sum; summing the above-mentioned M negative sample similarity indices to obtain a negative sample index sum; based on the above-mentioned positive sample index sum and the above-mentioned negative sample index sum, obtain a sample index sum; based on the above-mentioned positive sample index sum and the above-mentioned sample index sum, obtain the above-mentioned contrast learning loss.

[0011] According to an embodiment of the present invention, the above-mentioned steps of obtaining the above-mentioned contrast learning loss based on the above-mentioned negative sample similarity and the above-mentioned positive sample similarity include: for the n-th positive sample similarity, calculate the difference between the n-th positive sample similarity and the above-mentioned M negative sample similarities respectively to obtain the n-th sub-positive and negative sample similarity differences, and obtain M×N sub-positive and negative sample similarity differences; based on a preset boundary value, sum the above-mentioned M×N sub-positive and negative sample similarity differences to obtain M×N sub-contrast learning losses; accumulate the non-negative values in the above-mentioned M×N sub-contrast learning losses to obtain the above-mentioned contrast learning loss.

[0012] The second aspect of the present invention provides a video anomaly detection method, including: obtaining a target video; inputting the target video into a target video anomaly detection model to obtain a detection result, where the detection result indicates whether there is an anomaly in the target video, and the target video anomaly detection model is trained by using the method of the video anomaly detection model based on contrast learning.

[0013] The third aspect of the present invention provides a training device for a video anomaly detection model based on contrast learning, including: an input extraction module, configured to input a sample video into a feature extraction module of an initial video anomaly detection model to obtain sample video features; an input reconstruction module, configured to input the sample video features into a reconstruction module of the initial video anomaly detection model to obtain a reconstructed sample video, and based on a reconstruction loss function, obtain a reconstruction loss according to the sample video and the reconstructed sample video; a negative sample generation module, configured to generate a negative sample video feature group based on the sample video features and the reconstruction loss; a positive sample enhancement module, configured to perform data enhancement processing on the sample video features to obtain a positive sample video feature group; a contrast learning loss calculation module, configured to obtain a contrast learning loss based on a contrast learning loss function according to the negative sample video feature group and the positive sample video feature group; an adjustment module, configured to adjust parameters of the initial video anomaly detection model based on the contrast learning loss and the reconstruction loss to obtain a target video detection model.

[0014] The fourth aspect of the present invention provides a video anomaly detection device, including: an acquisition module, configured to acquire a target video; an input module, configured to input the target video into a target video anomaly detection model to obtain a detection result, where the detection result indicates whether there is an anomaly in the target video, and the target video anomaly detection model is trained by using the training device of the video anomaly detection model based on contrast learning.

[0015] The fifth aspect of the present invention provides an electronic device, including: one or more processors; a memory, configured to store one or more computer programs, where the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0016] The sixth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0017] The seventh aspect of the present invention further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0018] According to an embodiment of the present invention, by inputting the extracted sample video features into a reconstruction module, a reconstructed sample video is obtained, and a reconstruction loss is obtained based on the sample video and the reconstructed sample video. Under the guidance of the reconstruction loss, a negative sample video feature group is generated based on the sample video features, and at the same time, noise is added to the sample video features to generate a positive sample video feature group. A contrastive learning loss is calculated based on the negative sample video feature group and the positive sample video feature group, and the initial video anomaly detection model is jointly optimized based on the contrastive learning loss and the reconstruction loss, so that the memory module can learn a more complex positive sample video feature representation, thereby improving the recognition ability of normal videos, and clarifying the boundary between positive samples and negative samples based on the negative sample video feature group, making video anomaly detection more accurate. At the same time, the target video detection model can learn the differences between normal samples and abnormal samples without supervision, reducing the dependence on labeled data. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0020] Figure 1 The application scenario diagrams of the training method of the video anomaly detection model based on contrastive learning, the video anomaly detection method, the training device of the video anomaly detection model based on contrastive learning, and the video anomaly detection device according to the embodiments of the present invention are shown;

[0021] Figure 2 The flowchart of the training method of the video anomaly detection model based on contrastive learning according to the embodiments of the present invention is shown;

[0022] Figure 3 The flowchart of the video anomaly detection method according to the embodiments of the present invention is shown;

[0023] Figure 4 The flowchart of another training method of the video anomaly detection model according to the embodiments of the present invention is shown;

[0024] Figure 5 The structural block diagram of the training device of the video anomaly detection model based on contrastive learning according to the embodiments of the present invention is shown;

[0025] Figure 6 The structural block diagram of the video anomaly detection device according to the embodiments of the present invention is shown;

[0026] Figure 7 The block diagram of the electronic device suitable for implementing the training method of the video anomaly detection model based on contrastive learning according to the embodiments of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, numerous specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.

[0028] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0029] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0030] In cases where expressions similar to "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0031] Video anomaly detection is an important computer vision task and is widely applied in fields such as security monitoring, behavior recognition, and industrial inspection. Its goal is to automatically detect anomaly events in a video stream that are significantly different from the normal pattern. The memory-augmented autoencoder is a classic video anomaly detection method that enhances the model's ability to reconstruct the normal pattern by introducing a memory module to store the features of the normal pattern. The memory module stores the feature representations of normal samples, and the features of the input samples are matched and enhanced with the memory module to generate more accurate reconstruction results. The reconstruction error of normal samples is low, while the reconstruction error of anomaly samples is high because they cannot be effectively matched by the memory module, thus realizing anomaly detection. In the prior art, the feature update in the memory module usually depends on normal samples and lacks the ability to model anomaly samples, making the model unable to learn stronger discrimination ability in the feature space.

[0032] In view of this, an embodiment of the present invention provides a training method for a video anomaly detection model based on contrast learning, including: inputting a sample video into a feature extraction module of an initial video anomaly detection model to obtain sample video features; inputting the sample video features into a reconstruction module of the initial video anomaly detection model to obtain a reconstructed sample video, and based on a reconstruction loss function, obtaining a reconstruction loss according to the sample video and the reconstructed sample video; generating a negative sample video feature group based on the sample video features and the reconstruction loss; performing data augmentation processing on the sample video features to obtain a positive sample video feature group; obtaining a contrast learning loss based on the negative sample video feature group and the positive sample video feature group according to a contrast learning loss function; and adjusting parameters of the initial video anomaly detection model based on the contrast learning loss and the reconstruction loss to obtain a target video detection model.

[0033] Figure 1 FIG. shows an application scenario diagram of a training method for a video anomaly detection model based on contrast learning, a video anomaly detection method, a training device for a video anomaly detection model based on contrast learning, and a video anomaly detection device according to an embodiment of the present invention.

[0034] As Figure 1 shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0035] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0036] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0037] The server 105 may be a server that provides various services. For example, it may be a background management server (only for example) that supports the websites browsed by the user using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0038] It should be noted that the training method and video anomaly detection method of the video anomaly detection model based on contrastive learning provided in the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the training device and video anomaly detection device of the video anomaly detection model based on contrastive learning provided in the embodiments of the present invention can generally be set in the server 105. The training method and video anomaly detection method of the video anomaly detection model based on contrastive learning provided in the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the training device and video anomaly detection device of the video anomaly detection model based on contrastive learning provided in the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0039] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0040] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 The following will be based on Figure 2 the described scenario, and will describe in detail the training method of the video anomaly detection model based on contrastive learning in the embodiments of the invention through

[0041] Figure 2 shows a flowchart of the training method of the video anomaly detection model based on contrastive learning according to an embodiment of the present invention.

[0042] As Figure 2 shown, the training method of the video anomaly detection model based on contrastive learning in this embodiment includes operations S210 to S260.

[0043] In operation S210, the sample video is input into the feature extraction module of the initial video anomaly detection model to obtain the sample video features.

[0044] According to an embodiment of the present invention, the above sample video is a normal video segment, and the size is where B is the batch size, C is the number of channels, S is the length of the frame sequence, H and W are the height and width of a frame image in the sample video, F represents the number of feature channels after extraction, and the feature extraction module extracts video features. The feature extraction module includes an initial convolutional layer and three downsampling layers, and the feature size is gradually reduced to , , .

[0045] In operation S220, the sample video features are input into the reconstruction module of the initial video anomaly detection model to obtain a reconstructed sample video, and based on the reconstruction loss function, the reconstruction loss is obtained according to the sample video and the reconstructed sample video.

[0046] According to the embodiments of the present invention, the reconstruction module gradually restores the spatial dimension of the sample video features, and the output sizes are successively , , . Finally, a reconstructed sample video is generated through the output convolutional layer, and the size of the reconstructed sample video is . The reconstruction loss can be calculated by the following formula (1).

[0047] (1)

[0048] where represents the reconstruction loss, represents the i-th frame image of the sample video, represents the corresponding i-th frame image of the reconstructed sample video, represents calculating the Euclidean distance, which is used to measure the reconstruction error, represents the total number of frames of the sample video.

[0049] In operation S230, a negative sample video feature group is generated based on the sample video features and the reconstruction loss.

[0050] According to the embodiments of the present invention, perturbation noise can be added to the sample video features under the guidance of the reconstruction loss to obtain a negative sample video feature group, or random noise can be added to the sample video features to generate a negative sample video feature group. The feature dimension of the above sample video features is , so the feature dimension of each negative sample video feature in the obtained negative sample video feature group after adding noise is also .

[0051] In operation S240, data augmentation processing is performed on the sample video features to obtain a positive sample video feature group.

[0052] According to an embodiment of the present invention, the above data augmentation processing may be, for example, adding random noise to the sample video features to obtain a set of positive sample video features. Since the feature dimension of the sample video features is , the feature dimension of each positive sample video feature in the obtained set of positive sample video features is also .

[0053] According to an embodiment of the present invention, the generated random noise may be added to the sample video features to obtain a set of positive sample video features.

[0054] In operation S250, based on the contrastive learning loss function, a contrastive learning loss is obtained according to the set of negative sample video features and the set of positive sample video features.

[0055] In operation S260, the parameters of the initial video anomaly detection model are adjusted based on the contrastive learning loss and the reconstruction loss to obtain the target video detection model.

[0056] According to an embodiment of the present invention, the contrastive learning loss and the reconstruction loss can be expressed as a total loss, and the total loss can be calculated by the following formula (2):

[0057] (2)

[0058] Wherein, represents the total loss, represents the contrastive learning loss, represents the reconstruction loss, , is the weight parameter.

[0059] According to an embodiment of the present invention, by inputting the extracted sample video features into the reconstruction module, a reconstructed sample video is obtained, and a reconstruction loss is obtained based on the sample video and the reconstructed sample video. Under the guidance of the reconstruction loss, a set of negative sample video features is generated based on the sample video features, and at the same time, noise is added to the sample video features to generate a set of positive sample video features. A contrastive learning loss is calculated according to the set of negative sample video features and the set of positive sample video features, and the initial video anomaly detection model is jointly optimized based on the contrastive learning loss and the reconstruction loss, so that the memory module can learn more complex positive feature representations, thereby improving the recognition ability of normal videos, and clarifying the boundary between positive samples and negative samples according to the set of negative sample video features, making video anomaly detection more accurate. At the same time, the target video detection model can learn the differences between normal samples and abnormal samples without supervision, reducing the dependence on labeled data.

[0060] According to an embodiment of the present invention, generating a negative sample video feature group based on sample video features and a reconstruction loss includes: calculating a gradient vector of the reconstruction loss with respect to the sample video features based on a gradient extraction function, where the gradient vector characterizes the sensitivity of the sample video features to the reconstruction error; generating the negative sample video feature group based on the gradient direction information in the gradient vector and the sample video features.

[0061] According to an embodiment of the present invention, the negative sample video feature group can be calculated through the following formula (3).

[0062] (3)

[0063] Where Z represents the sample video features, is a hyperparameter of the perturbation intensity, used to control the degree to which the generated negative sample video features N deviate from the sample video features Z, represents the reconstruction loss, used to measure the error between the sample video and the sample reconstructed video, represents the reconstruction loss with respect to the sample video features gradient vector, represents the gradient direction information in the gradient vector.

[0064] According to an embodiment of the present invention, by adding gradient-based perturbations to the sample video features, the generated negative sample video features will deliberately deviate from the distribution of normal samples, making them more challenging. The gradient direction points in the direction where the reconstruction error increases, which means that the generated negative sample video features will be more difficult to be reconstructed by the target video detection model, and thus will be closer to the distribution of abnormal samples in the feature space.

[0065] According to an embodiment of the present invention, the value of will affect the generated negative sample video features, being too small may cause the negative sample video features to be too close to the normal samples and difficult to achieve a distinguishing effect. Being too large may cause the negative sample video features to completely deviate from the normal pattern, lose their structure, and affect the training effect. It is recommended to adjust through experiments, usually select to set within [0.01, 0.1].

[0066] According to an embodiment of the present invention, the negative sample video features generated based on confrontation are completely generated by gradient vectors without manually annotating abnormal samples, greatly reducing the data requirements and annotation costs. Further, by generating negative sample video features that are difficult to reconstruct, the model learns stronger discrimination ability in the feature space, and thus the normal samples and abnormal samples are more separated in the feature space. In addition, the distribution of the negative sample video features deviates from the sample video features but still retains a certain structure, which can effectively fill the lack of negative samples in unsupervised learning. And the perturbation intensity is controllable, facilitating the adjustment of the difficulty of the negative sample video features, so that the target video detection model gradually adapts to more complex abnormal scenarios; more importantly, the gradient direction is closely related to the sample video features and the current state of the target video detection model, so the generated negative sample video features will change dynamically with training and have high diversity. This dynamically generated negative sample video features avoids the overfitting problem that may be caused by static negative sample video features and improves the discrimination ability and robustness of the target video detection model.

[0067] According to an embodiment of the present invention, data augmentation processing is performed on the sample video features to obtain a positive sample video feature group, including: generating a first noise based on a first preset noise range, where the first noise represents the change of environmental factors in the sample video; generating a second noise based on a second preset noise range, where the second noise represents the visual interference factors in the sample video; performing proportional scaling on the sample video features according to the first noise to obtain a scaled positive sample video feature group; and performing random offset on the scaled positive sample video feature group according to the second noise to obtain the positive sample video feature group.

[0068] According to an embodiment of the present invention, the positive sample video feature group is to enhance the learning ability of the target video detection model for normal modes and at the same time improve the generalization performance of the target video detection model in complex scenarios. By adding noise perturbation or feature transformation, the generated positive sample video feature group is semantically consistent with the original sample video features but has a more complex background or feature distribution, making the target video detection model face greater challenges during training.

[0069] The positive sample video feature group can be calculated by the following formula (4).

[0070] (4)

[0071] Wherein, represents the positive sample video feature, Z represents the sample video feature, represents the first noise, represents the second noise.

[0072] According to an embodiment of the present invention, the first preset noise range is the N(1, 1) range determined from the noise sampled from a normal distribution. The first noise is used to scale the sample video features. The sampled noise values fluctuate around the mean value of 1, ensuring that the overall amplitude change of the sample video features is small and the semantics remain consistent, but the feature distribution is perturbed. This step simulates the changes of the sample at different scales or amplitudes, such as changes in light intensity or scene details. The second preset noise range is the N(0, 1) range determined from the noise sampled from a normal distribution. The second noise is used to add random offsets to the sample features. The sampled noise values fluctuate around the mean value of 0, simulating the random changes of the sample in the background or environment. This step simulates the random perturbations of the sample in the background or environment, such as noise interference or minor occlusion. The positive sample video features have noise perturbations but maintain semantic consistency with the sample video features.

[0073] According to an embodiment of the present invention, the first noise is generated through the first preset noise range, and the second noise is generated through the second preset noise range. Further, the sample video features are scaled according to the first noise, and the scaled positive sample video feature group is randomly offset according to the second noise to obtain the positive sample video feature group, so that the positive sample video feature group simulates the changes of complex scenes in the sample video on the premise of ensuring semantic consistency, enabling the target video detection model to adapt to more diverse normal modes, showing stronger robustness in the case of noise or environmental changes. And because the feature distribution of the positive sample video feature group is more complex, the target video detection model needs stronger capabilities to correctly identify these samples as normal samples during training, thereby improving the modeling ability of the target video detection model for normal modes.

[0074] According to an embodiment of the present invention, based on the contrastive learning loss function, the contrastive learning loss is obtained according to the negative sample video feature group and the positive sample video feature group, including: respectively inputting the negative sample video feature group and the positive sample video feature group into the memory module of the initial video anomaly detection model to obtain the negative sample similarity and the positive sample similarity. The negative sample similarity represents the similarity degree between the negative sample video feature group and the memory vector feature group stored in the memory module, and the positive sample similarity represents the similarity degree between the positive sample video feature group and the memory vector feature group; based on the contrastive learning loss function, the contrastive learning loss is obtained according to the negative sample similarity and the positive sample similarity.

[0075] According to an embodiment of the present invention, the memory vector feature group includes P memory vector features, the negative sample video feature group includes M negative sample video features, and the positive sample video feature group includes N positive sample video features. P, M, and N are all positive integers greater than 0. p is a positive integer less than or equal to P and greater than 0, m is a positive integer less than or equal to M and greater than 0, and n is a positive integer less than or equal to N and greater than 0. The negative sample video feature group and the positive sample video feature group are respectively input into the memory module of the initial video anomaly detection model to obtain the negative sample similarity and the positive sample similarity, including: for the m-th negative sample video feature in the negative sample video feature group, using the cosine similarity function, based on the m-th negative sample video feature and the p-th memory vector feature to obtain the p-th sub-negative sample cosine similarity, so as to obtain a sub-negative sample cosine similarity group, and the sub-negative sample cosine similarity group includes P sub-negative sample cosine similarities; for the n-th positive sample video feature in the positive sample video feature group, using the cosine similarity function, based on the n-th positive sample video feature and the p-th memory vector feature to obtain the p-th sub-positive sample cosine similarity, so as to obtain a sub-positive sample cosine similarity group, and the sub-positive sample cosine similarity group includes P sub-positive sample cosine similarities; using a preset normalization function, respectively normalizing and summing the P sub-negative sample cosine similarities to obtain the m-th negative sample similarity, and obtaining M negative sample similarities; using a preset normalization function, respectively normalizing and summing the P sub-positive sample cosine similarities to obtain the n-th positive sample similarity, and obtaining N positive sample similarities.

[0076] According to an embodiment of the present invention, the input of the memory module includes input features and the memory vector feature group stored in the memory module , the memory vector feature group includes P memory vectors, d represents the dimension of each memory vector, represents the first memory vector feature in the memory vector feature group, and so on, represents the p-th memory vector feature in the memory vector feature group.

[0077] According to an embodiment of the present invention, the sub-negative sample cosine similarity and the sub-positive sample sine similarity can be calculated by the following formula (5):

[0078] (5)

[0079] where k represents the k-th negative sample video feature in the negative sample video feature group or the k-th positive sample video feature in the positive sample video feature group, represents the p-th memory vector feature in the memory vector feature group, Denotes the sub-negative sample cosine similarity between the k-th negative sample video feature in the negative sample video feature group and the p-th memory vector feature in the memory vector feature group, or the sub-positive sample cosine similarity between the k-th positive sample video feature in the positive sample video feature group and the p-th memory vector feature in the memory vector feature group. Denotes and the dot product of Denotes of the norm of Denotes of the norm of

[0080] According to an embodiment of the present invention, the above-mentioned preset normalization function can be, for example, the Softmax function. The Softmax function combines the advantages of the maximum similarity and the average similarity, is more flexible and robust, and can calculate the negative sample similarity and the positive sample similarity through the following formula (6):

[0081] (6)

[0082] Wherein, Denotes the sub-negative sample similarity between the k-th negative sample video feature in the negative sample video feature group and the memory vector feature group, or the sub-positive sample similarity between the k-th positive sample video feature in the positive sample video feature group and the memory vector feature group. Denotes the sub-negative sample cosine similarity between the k-th negative sample video feature in the negative sample video feature group and the p-th memory vector feature in the memory vector feature group, or the sub-positive sample cosine similarity between the k-th positive sample video feature in the positive sample video feature group and the p-th memory vector feature in the memory vector feature group. Denotes the memory vector feature group, k denotes the input feature, P is the number of memory vector features, Denotes the p-th memory vector feature in the memory vector feature group, Denotes the j-th memory vector feature in the memory vector feature group, Denotes the temperature parameter.

[0083] According to an embodiment of the present invention, the temperature parameter usually needs to be adjusted according to the specific task and data distribution, and is generally set to 0.1 or 0.07. A smaller will make the distribution sharper and emphasize the difference between positive and negative samples; a larger will make the distribution smoother and reduce the gradient fluctuation. In practical applications, the most suitable temperature parameter can be selected through experimental optimization.

[0084] According to an embodiment of the present invention, the memory module of the initial video anomaly detection model is trained through positive and negative sample video feature groups. The target video detection model can learn different similarity features between positive samples and negative samples and the memory vector feature group respectively, so as to better understand different types of video features, enhance the ability to distinguish normal and abnormal video features, contribute to improving the detection accuracy of the model for video anomaly situations in practical applications, and through the contrastive learning loss function, convert the similarity difference between positive and negative samples into an optimizable numerical index, which can enable the target video detection model to learn in a direction more conducive to distinguishing normal and abnormal video features, thereby improving the overall performance of the target video detection model.

[0085] According to an embodiment of the present invention, obtaining the contrastive learning loss according to the negative sample similarity and the positive sample similarity includes: summing the N positive sample similarity indices to obtain the positive sample index sum; summing the M negative sample similarity indices to obtain the negative sample index sum; obtaining the sample index sum based on the positive sample index sum and the negative sample index sum; and obtaining the contrastive learning loss based on the positive sample index sum and the sample index sum.

[0086] According to an embodiment of the present invention, the contrastive learning loss can be calculated through the following formula (7).

[0087] (7)

[0088] Wherein, represents the contrastive learning loss, represents the positive sample video feature group, represents the negative sample video feature group, represents the temperature parameter, which is used to control the smoothness of the distribution.

[0089] According to an embodiment of the present invention, the numerator part of the above formula (7) calculates the similarity between the positive sample video feature group and the memory feature vector group in the memory module, and the denominator part calculates the total similarity between the positive sample video feature group and the negative sample video feature group and the memory feature vector group in the memory module.

[0090] According to an embodiment of the present invention, obtaining the contrastive learning loss according to the negative sample similarity and the positive sample similarity includes: obtaining the contrastive learning loss according to the negative sample similarity and the positive sample similarity includes: for the nth positive sample similarity, calculating the difference between the nth positive sample similarity and each of the M negative sample similarities respectively to obtain the nth sub positive-negative sample similarity difference, and obtaining M×N sub positive-negative sample similarity differences; summing the preset boundary value and each of the M×N sub positive-negative sample similarity differences respectively to obtain M×N sub contrastive learning losses; and accumulating the non-negative values among the M×N sub contrastive learning losses to obtain the contrastive learning loss.

[0091] According to an embodiment of the present invention, the above contrast learning loss function can also be a triplet loss function, and the contrast learning loss can be calculated by the following formula (8):

[0092] (8)

[0093] Wherein, represents the contrast learning loss obtained according to the triplet loss function, represents a preset boundary value, which is used to control the similarity difference between the positive and negative sample video features, represents taking a non-negative value, which is used to ensure that the loss is not less than 0, represents the positive sample video feature group, represents the negative sample video feature group, represents the negative sample similarity between the negative sample video feature group and the memory vector feature group, represents the positive sample similarity between the positive sample video feature group and the memory vector feature group.

[0094] According to an embodiment of the present invention, the ratio of the positive sample video feature group and the negative sample video feature group is calculated through a normalized probability distribution to optimize the overall distribution of the feature space. Further, the contrast learning loss is calculated through the triplet loss function, which will constrain the similarity difference between the positive sample video feature group and the negative sample video feature group, thereby optimizing the local structure of the feature space.

[0095] Figure 3 Fig. shows a flowchart of a video anomaly detection method according to an embodiment of the present invention.

[0096] As Figure 3 shown, the video anomaly detection method of this embodiment includes operation S310 to operation S320.

[0097] In operation S310, a target video is acquired.

[0098] In operation S320, the target video is input into the target video anomaly detection model to obtain a detection result.

[0099] Wherein, the detection result characterizes whether there is an anomaly in the target video, and the target video anomaly detection model is trained by using the training method of the above video anomaly detection model based on contrast learning.

[0100] According to an embodiment of the present invention, the target video anomaly detection model includes a target memory module and a target reconstruction module. The target video features extracted from the target video first pass through the target memory module to calculate the corresponding weighted target video features, and then are reconstructed by the target reconstruction module to obtain the reconstructed target video. The corresponding weights can be calculated by the following formula (9).

[0101] The weighted target video features can be calculated through the following formula (9).

[0102] (9)

[0103] Wherein, represents the weighted target video features, represents the i-th target memory vector feature in the target memory vector feature group stored in the target memory module, represents the target video features, is the temperature parameter.

[0104] According to the embodiments of the present invention, when identifying a video after training is completed, there is no need to generate a positive sample video feature group and a negative sample video feature group anymore. Instead, it is directly determined whether the target video is abnormal through the reconstruction error between the target video and the target reconstructed video.

[0105] Figure 4 FIG. shows a flowchart of another method for training a video anomaly detection model according to an embodiment of the present invention.

[0106] In operation S410, the initial video anomaly detection model extracts features from the sample video to obtain sample video features.

[0107] In operation S420, the reconstruction module of the initial video anomaly detection model reconstructs based on the sample video features to obtain a reconstructed sample video.

[0108] In operation S430, a reconstruction loss is calculated based on the reconstructed sample video and the sample video.

[0109] In operation S440, noise is added to the sample video features to obtain a positive sample video feature group.

[0110] In operation S450, noise is added to the sample video based on the reconstruction loss to obtain a negative sample video feature group.

[0111] In operation S460, a contrast learning loss is obtained based on the positive sample video feature group and the negative sample video feature group.

[0112] In operation S470, the contrast learning loss and the reconstruction loss are jointly optimized.

[0113] In operation S480, when the contrast learning loss and the reconstruction loss reach preset conditions, a target video detection model is obtained.

[0114] Based on the above method for training a video anomaly detection model based on contrast learning, the present invention also provides a device for training a video anomaly detection model based on contrast learning. The following will be combined with Figure 5 to describe this device in detail.

[0115] Figure 5 The block diagram of the training device of the video anomaly detection model based on contrast learning according to an embodiment of the present invention is shown.

[0116] As Figure 5 shown, the training device 500 of the video anomaly detection model based on contrast learning in this embodiment includes an input extraction module 510, an input reconstruction module 520, a negative sample generation module 530, a positive sample enhancement module 540, a contrast learning loss calculation module 550, and an adjustment module 560.

[0117] The input extraction module 510 is used to input the sample video into the feature extraction module of the initial video anomaly detection model to obtain the sample video features. In one embodiment, the input extraction module 510 can be used to perform the operation S210 described above, which will not be elaborated here.

[0118] The input reconstruction module 520 is used to input the sample video features into the reconstruction module of the initial video anomaly detection model to obtain the reconstructed sample video, and based on the reconstruction loss function, obtain the reconstruction loss according to the sample video and the reconstructed sample video. In one embodiment, the input reconstruction module 520 can be used to perform the operation S220 described above, which will not be elaborated here.

[0119] The negative sample generation module 530 is used to generate a negative sample video feature group based on the sample video features and the reconstruction loss. In one embodiment, the negative sample generation module 530 can be used to perform the operation S230 described above, which will not be elaborated here.

[0120] The positive sample enhancement module 540 is used to perform data enhancement processing on the sample video features to obtain a positive sample video feature group. In one embodiment, the positive sample enhancement module 540 can be used to perform the operation S240 described above, which will not be elaborated here.

[0121] The contrast learning loss calculation module 550 is used to obtain the contrast learning loss based on the contrast learning loss function according to the negative sample video feature group and the positive sample video feature group. In one embodiment, the contrast learning loss calculation module 550 can be used to perform the operation S250 described above, which will not be elaborated here.

[0122] The adjustment module 560 is used to adjust the parameters of the initial video anomaly detection model based on the contrast learning loss and the reconstruction loss to obtain the target video detection model. In one embodiment, the adjustment module 560 can be used to perform the operation S260 described above, which will not be elaborated here.

[0123] According to an embodiment of the present invention, any multiple of the input extraction module 510, the input reconstruction module 520, the negative sample generation module 530, the positive sample enhancement module 540, the contrastive learning loss calculation module 550, and the adjustment module 560 can be combined and implemented in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the input extraction module 510, the input reconstruction module 520, the negative sample generation module 530, the positive sample enhancement module 540, the contrastive learning loss calculation module 550, and the adjustment module 560 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as hardware or firmware through circuit integration or packaging, or can be implemented in any one of the three implementation manners of software, hardware, and firmware or in any appropriate combination of several of them. Alternatively, at least one of the input extraction module 510, the input reconstruction module 520, the negative sample generation module 530, the positive sample enhancement module 540, the contrastive learning loss calculation module 550, and the adjustment module 560 can be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.

[0124] Based on the above video anomaly detection method, the present invention also provides a video anomaly detection device. The following will be combined with Figure 6 to describe this device in detail.

[0125] Figure 6 The structural block diagram of the video anomaly detection device according to an embodiment of the present invention is shown.

[0126] As Figure 6 shown, the video anomaly detection device 600 of this embodiment includes an acquisition module 610 and an input module 620.

[0127] The acquisition module 610 is used to acquire a target video. In one embodiment, the acquisition module 610 can be used to perform the operation S310 described above, which will not be elaborated here.

[0128] The input module 620 is used to input the target video into the target video anomaly detection model to obtain a detection result, and the detection result indicates whether there is an anomaly in the target video. The target video anomaly detection model is trained by using the training device of the above video anomaly detection model based on contrastive learning. In one embodiment, the input module 620 can be used to perform the operation S320 described above, which will not be elaborated here.

[0129] According to an embodiment of the present invention, any number of modules among the acquisition module 610 and the input module 620 can be combined and implemented in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the acquisition module 610 and the input module 620 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable manner of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the acquisition module 610 and the input module 620 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.

[0130] Figure 7 A block diagram of an electronic device suitable for implementing a training method of a video anomaly detection model based on contrastive learning according to an embodiment of the present invention is shown.

[0131] As Figure 7 shown, the electronic device 700 according to an embodiment of the present invention includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM 702) or a program loaded from a storage section 708 into a random access memory (RAM 703). The processor 701 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 701 can also include on-board memory for caching purposes. The processor 701 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0132] In the RAM 703, various programs and data required for the operation of the electronic device 700 are stored. The processor 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The processor 701 performs various operations of the method flow according to an embodiment of the present invention by executing the programs in the ROM 702 and / or the RAM 703. It should be noted that the programs can also be stored in one or more memories other than the ROM 702 and the RAM 703. The processor 701 can also perform various operations of the method flow according to an embodiment of the present invention by executing the programs stored in the one or more memories.

[0133] According to an embodiment of the present invention, the electronic device 700 may further include an input / output (I / O) interface 705, and the input / output (I / O) interface 705 is also connected to the bus 704. The electronic device 700 may further include one or more of the following components connected to the input / output (I / O) interface 705: an input portion 706 including a keyboard, a mouse, etc.; an output portion 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 708 including a hard disk, etc.; and a communication portion 709 including a network interface card such as a LAN card, a modem, etc. The communication portion 709 performs communication processing via a network such as the Internet. The drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read from it can be installed into the storage portion 708 as needed.

[0134] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiments of the present invention is implemented.

[0135] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM 703), a read-only memory (ROM 702), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 702 and / or RAM 703 and / or one or more memories other than ROM 702 and RAM 703.

[0136] An embodiment of the present invention further includes a computer program product, which includes a computer program, and the computer program includes program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the training method of the video anomaly detection model based on contrast learning provided by the embodiments of the present invention.

[0137] When the computer program is executed by the processor 701, the above functions defined in the system / apparatus of the embodiments of the present invention are executed. According to the embodiments of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0138] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program can also be transmitted and distributed in the form of signals on a network medium, and be downloaded and installed through the communication part 709, and / or be installed from the removable medium 711. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0139] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or be installed from the removable medium 711. When the computer program is executed by the processor 701, the above functions defined in the system of the embodiments of the present invention are executed. According to the embodiments of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0140] According to the embodiments of the present invention, the program code for executing the computer program provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to, such as Java, C++, python, the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0142] Those skilled in the art will appreciate that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0143] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in the respective embodiments cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A training method for a video anomaly detection model based on contrastive learning, characterized in that, The method includes: Inputting the sample video into the feature extraction module of the initial video anomaly detection model to obtain the sample video features; Inputting the sample video features into the reconstruction module of the initial video anomaly detection model to obtain the reconstructed sample video, and based on the reconstruction loss function, obtaining the reconstruction loss according to the sample video and the reconstructed sample video; Generating a negative sample video feature group based on the sample video features and the reconstruction loss; Performing data augmentation processing on the sample video features to obtain a positive sample video feature group; Based on the contrastive learning loss function, obtaining the contrastive learning loss according to the negative sample video feature group and the positive sample video feature group; Adjusting the parameters of the initial video anomaly detection model based on the contrastive learning loss and the reconstruction loss to obtain the target video detection model.

2. The method according to claim 1, characterized in that, The obtaining the contrastive learning loss according to the negative sample video feature group and the positive sample video feature group based on the contrastive learning loss function includes: Inputting the negative sample video feature group and the positive sample video feature group into the memory module of the initial video anomaly detection model respectively to obtain the negative sample similarity and the positive sample similarity. The negative sample similarity represents the similarity degree between the negative sample video feature group and the memory vector feature group stored in the memory module, and the positive sample similarity represents the similarity degree between the positive sample video feature group and the memory vector feature group; Based on the contrastive learning loss function, obtaining the contrastive learning loss according to the negative sample similarity and the positive sample similarity.

3. The method according to claim 1, characterized in that The generating the negative sample video feature group based on the sample video features and the reconstruction loss includes: Calculating the gradient vector of the reconstruction loss with respect to the sample video features based on the gradient extraction function. The gradient vector represents the sensitivity degree of the sample video features to the reconstruction error; Generating the negative sample video feature group based on the gradient direction information in the gradient vector and the sample video features.

4. The method according to claim 1, wherein The performing data augmentation processing on the sample video features to obtain a positive sample video feature group includes: Generating the first noise based on the first preset noise range. The first noise represents the environmental factor change in the sample video; Generating the second noise based on the second preset noise range. The second noise represents the visual interference factor in the sample video; Performing proportional scaling on the sample video features according to the first noise to obtain the scaled positive sample video feature group; Performing random offset on the scaled positive sample video feature group according to the second noise to obtain the positive sample video feature group.

5. The method according to claim 2, wherein The memory vector feature group includes P memory vector features, the negative sample video feature group includes M negative sample video features, and the positive sample video feature group includes N positive sample video features. P, M, and N are all positive integers greater than 0. p is a positive integer less than or equal to P and greater than 0, m is a positive integer less than or equal to M and greater than 0, and n is a positive integer less than or equal to N and greater than 0; Inputting the negative sample video feature group and the positive sample video feature group into the memory module of the initial video anomaly detection model respectively to obtain a negative sample similarity and a positive sample similarity, includes: For the m-th negative sample video feature in the negative sample video feature group, using the cosine similarity function, obtaining the p-th sub-negative sample cosine similarity based on the m-th negative sample video feature and the p-th memory vector feature, so as to obtain a sub-negative sample cosine similarity group, where the sub-negative sample cosine similarity group includes P sub-negative sample cosine similarities; For the n-th positive sample video feature in the positive sample video feature group, using the cosine similarity function, obtaining the p-th sub-positive sample cosine similarity based on the n-th positive sample video feature and the p-th memory vector feature, so as to obtain a sub-positive sample cosine similarity group, where the sub-positive sample cosine similarity group includes P sub-positive sample cosine similarities; Using a preset normalization function, normalizing and summing the P sub-negative sample cosine similarities respectively to obtain the m-th negative sample similarity, so as to obtain M negative sample similarities; Using a preset normalization function, normalizing and summing the P sub-positive sample cosine similarities respectively to obtain the n-th positive sample similarity, so as to obtain N positive sample similarities.

6. The method according to claim 5, wherein Obtaining the contrastive learning loss according to the negative sample similarity and the positive sample similarity, includes: Summing the N positive sample similarity exponents to obtain a positive sample exponent sum; Summing the M negative sample similarity exponents to obtain a negative sample exponent sum; Based on the positive sample exponent sum and the negative sample exponent sum, obtaining a sample exponent sum; Based on the positive sample exponent sum and the sample exponent sum, obtaining the contrastive learning loss.

7. The method according to claim 5, wherein Obtaining the contrastive learning loss according to the negative sample similarity and the positive sample similarity, includes: For the n-th positive sample similarity, subtracting the n-th positive sample similarity from each of the M negative sample similarities respectively to obtain the n-th sub-positive and negative sample similarity difference, and obtaining M×N sub-positive and negative sample similarity differences; Based on a preset boundary value, summing the M×N sub-positive and negative sample similarity differences to obtain M×N sub-contrastive learning losses; Accumulating the non-negative values in the M×N sub-contrastive learning losses to obtain the contrastive learning loss.

8. A video anomaly detection method, characterized in that, The method includes: Obtaining a target video; Inputting the target video into a target video anomaly detection model to obtain a detection result, where the detection result indicates whether there is an anomaly in the target video, and the target video anomaly detection model is trained by using the method described in any one of claims 1 to 7.

9. A training device for a video anomaly detection model based on contrastive learning, characterized in that, The device includes: An input extraction module, configured to input a sample video into the feature extraction module of an initial video anomaly detection model to obtain sample video features; An input reconstruction module, configured to input the sample video features into the reconstruction module of the initial video anomaly detection model to obtain a reconstructed sample video, and based on a reconstruction loss function, obtaining a reconstruction loss according to the sample video and the reconstructed sample video; A negative sample generation module, configured to generate a negative sample video feature group based on the sample video features and the reconstruction loss; A positive sample enhancement module, configured to perform data enhancement processing on the sample video features to obtain a positive sample video feature group; A contrastive learning loss calculation module, configured to obtain a contrastive learning loss based on a contrastive learning loss function according to the negative sample video feature group and the positive sample video feature group; An adjustment module, configured to adjust parameters of the initial video anomaly detection model based on the contrastive learning loss and the reconstruction loss to obtain a target video detection model.

10. A video anomaly detection device, characterized in that, The apparatus includes: An acquisition module, configured to acquire a target video; An input module, configured to input the target video into a target video anomaly detection model to obtain a detection result, where the detection result indicates whether there is an anomaly in the target video, and the target video anomaly detection model is trained by using the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Surveillance video anomaly detection method and system based on adversarial learning

    CN115909144A

  • Weak supervision video anomaly detection-oriented dual dynamic memory network construction method

    CN116563744A

  • Data anomaly detection method and device, equipment and storage medium

    CN117009903A

  • Electric energy meter anomaly detection method and device based on confrontation contrast auto-encoder

    CN117092582A

  • Image anomaly detection method and system based on comparative learning

    CN118115450A