Video anomaly detection, model training method, device, model and storage medium

By using a semi-supervised learning method in video anomaly detection, the detection model is trained using normal video samples to predict and constrain the time-domain fluctuation state of visual feature errors, the video anomaly detection model is solved in the lack of abnormal data, and more efficient abnormal detection is achieved.

CN113515993BActive Publication Date: 2025-08-05ALIBABA GROUP HOLDING LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011321517.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-23
Publication Date
2025-08-05
Estimated Expiration
2040-11-23

AI Technical Summary

Technical Problem

In the absence of abnormal event data in the prior art, the performance of the video anomaly detection model is poor, making it difficult to effectively train and detect abnormal events.

Method used

The semi-supervised learning method is used to train the detection model using normal video samples in the target scenario, and optimize the model parameters by predicting the visual feature errors of video frames and future frames, and using the fluctuation state of the visual feature error in the time domain as a constraint.

Benefits of technology

The detection model's prediction performance for future frames is improved, and abnormal videos in the target scenario can be more accurately identified, and misjudgment is reduced. It is suitable for scenarios such as safety supervision, emergency management and traffic supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113515993B_ABST
    Figure CN113515993B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a video anomaly detection, model training method, device, model and storage medium. In the embodiment of the present application, a normal video sample in a target scene is used as a training sample, and a detection model is used to predict the visual features of a sample future frame of at least one sample video frame in the normal video sample, and the visual feature error between the sample video frame and the corresponding sample future frame is calculated; the detection model is trained with the fluctuation state of the visual feature error in the time domain reaching the target state as a constraint condition. Accordingly, the trained detection model can ensure that the fluctuation state of the visual feature error in the time domain of the normal video in the target scene tends to the target state, while the fluctuation state of the visual feature error in the time domain of the abnormal video will deviate from the target state. Based on this, the time domain fluctuation of the visual feature error can be used as an effective basis for judging whether an anomaly occurs in the video, thereby quickly and accurately distinguishing abnormal videos in the target scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video processing technology, and in particular to a video anomaly detection, model training method, device, model and storage medium. Background Art

[0002] Video anomaly detection aims to automatically perceive abnormal events in video sequences.

[0003] In natural scenarios, abnormal events, such as fires and explosions, can be dangerous and even cause serious damage. Therefore, in practice, it is difficult to obtain a large number of training samples containing abnormal events. Consequently, in the absence of abnormal data, detection models trained using traditional fully supervised methods perform poorly and may even fail to complete training.

[0004] Therefore, there is an urgent need for a solution that can improve the dilemma of video anomaly detection. Summary of the Invention

[0005] Various aspects of the present application provide a video anomaly detection, model training method, device, model and storage medium to improve the accuracy of video anomaly detection.

[0006] The present invention provides a model training method, including:

[0007] Input normal video samples under the target scene into the detection model for video anomaly detection;

[0008] Using the detection model, respectively predicting visual features of sample future frames of at least one sample video frame in the normal video sample; and calculating a visual feature error between the at least one sample future frame and its corresponding sample video frame in the normal video sample;

[0009] The detection model is trained with the first constraint condition that the fluctuation state of the visual feature error in the time domain reaches a target state, so that the detection model obtains model parameters under the target scene.

[0010] The present application also provides a video anomaly detection method applicable to a detection model, including:

[0011] Get the video to be detected;

[0012] respectively predicting visual features of future frames of at least one video frame in the video to be detected;

[0013] Calculating a visual feature error between the at least one future frame and its corresponding video frame in the to-be-detected video frame;

[0014] If the fluctuation state of the visual feature error in the time domain is in an abnormal state that does not meet the target state, it is determined that the video to be detected has an abnormality;

[0015] The target state is a constraint condition set for the fluctuation state of the visual feature error of a normal video in the time domain during the training process of the detection model.

[0016] The present invention also provides a method for detecting video anomalies, including:

[0017] In response to a request to call a target service, determining a processing resource corresponding to the target service, and performing the following steps using the processing resource corresponding to the target service:

[0018] Get the video to be tested;

[0019] respectively predicting visual features of future frames of at least one video frame in the video to be detected;

[0020] Calculating a visual feature error between the at least one future frame and its corresponding video frame in the to-be-detected video frame;

[0021] If the fluctuation state of the visual feature error in the time domain is in an abnormal state that does not meet the target state, it is determined that the video to be detected has an abnormality;

[0022] The target state is a constraint condition set for the fluctuation state of the visual feature error of a normal video in the time domain during the training process of the detection model.

[0023] The embodiment of the present application also provides a detection model, including an encoder, a memory unit and a processor;

[0024] The encoder is used to obtain a video to be detected; and extract visual features of at least one video frame in the video to be detected;

[0025] The memory unit is configured to respectively predict visual features of future frames of the at least one video frame;

[0026] The processor is configured to calculate a visual feature error between the at least one future frame and its corresponding video frame in the video frame to be detected; if a fluctuation state of the visual feature error in the time domain is abnormal and does not meet a target state, determining that an abnormality exists in the video to be detected;

[0027] The target state is a constraint condition set for the fluctuation state of the visual feature error of a normal video in the time domain during the training process of the detection model.

[0028] An embodiment of the present application further provides a computing device, including a memory and a processor;

[0029] The memory is used to store one or more computer instructions;

[0030] The processor is coupled to the memory and configured to execute the one or more computer instructions for:

[0031] Input normal video samples under the target scene into the detection model for video anomaly detection;

[0032] Using the detection model, respectively predicting visual features of sample future frames of at least one sample video frame in the normal video sample; and calculating a visual feature error between the at least one sample future frame and its corresponding sample video frame in the normal video sample;

[0033] The detection model is trained with the first constraint condition that the fluctuation state of the visual feature error in the time domain reaches a target state, so that the detection model obtains model parameters under the target scene.

[0034] An embodiment of the present application also provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the aforementioned model training method or the aforementioned video anomaly detection method.

[0035] In an embodiment of the present application, a normal video sample under a target scene is used as a training sample, and a detection model is used to predict the visual features of a sample future frame of at least one sample video frame in the normal video sample, and the visual feature error between the sample video frame and the corresponding sample future frame is calculated; the detection model is trained with the constraint that the fluctuation state of the visual feature error in the time domain reaches the target state. This semi-supervised training method can effectively improve the prediction performance of the detection model for future frames. Based on this, in the process of video anomaly detection, the detection model can be used to ensure that the fluctuation state of the visual feature error in the time domain in the normal video under the target scene tends to the target state, while the fluctuation state of the visual feature error in the time domain in the abnormal video will deviate from the target state, thereby more accurately detecting abnormal situations in the video under the target scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0037] Figure 1 A schematic structural diagram of a detection model provided by an exemplary embodiment of the present application;

[0038] Figure 2An internal logic diagram of a detection model provided by an exemplary embodiment of the present application;

[0039] Figure 3 A flowchart of a model training method provided by an exemplary embodiment of the present application;

[0040] Figure 4 A flowchart of a video anomaly detection method provided by an exemplary embodiment of the present application;

[0041] Figure 5 A flowchart of another video anomaly detection method provided by an exemplary embodiment of the present application;

[0042] Figure 6 A schematic diagram of the structure of a computing device provided as an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0043] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0044] In order to solve the technical problem that the detection model trained by the existing fully supervised method has poor performance and may even be unable to complete the training, in some embodiments of the present application: normal video samples in the target scene are used as training samples, and the detection model is used to predict the visual features of the sample future frame of at least one sample video frame in the normal video sample, and the visual feature error between the sample video frame and the corresponding sample future frame is calculated; the detection model is trained with the constraint that the fluctuation state of the visual feature error in the time domain reaches the target state. This semi-supervised training method can effectively improve the prediction performance of the detection model for future frames. Based on this, in the process of video anomaly detection, the detection model can be used to ensure that the fluctuation state of the visual feature error in the time domain in the normal video in the target scene tends to the target state, while the fluctuation state of the visual feature error in the time domain in the abnormal video will deviate from the target state, thereby more accurately detecting the abnormal situation in the video under the target scene.

[0045] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0046] Figure 1 A structural diagram of a detection model provided as an exemplary embodiment of the present application. Figure 2This is a schematic diagram of the internal logic of a detection model provided by an exemplary embodiment of the present application. Figure 1 and Figure 2 As shown, the detection model at least includes an encoder 10 , a memory unit 20 and a processor 40 .

[0047] refer to Figure 1 and Figure 2 , encoder 10, can be used to extract the visual features of each video frame in the video sequence, which can be referred to Figure 2 The feature map extraction part in the encoder 10 can use a deep convolutional deep network to extract visual features from the video frame. Specifically, according to the preset step size, feature extraction can be performed on the video frame through the convolution kernel to obtain the visual features of the video frame (such as Figure 2 f in 2,1 、f T,1 For example, visual features can be represented in the form of feature maps. Visual features can be used to represent semantic information in video frames.

[0048] The memory unit 20 can be used to predict the visual features of future frames of the video frame (such as Figure 2 in etc.), please refer to Figure 2 Here, the memory unit 20 may predict the visual features of a future frame of the current video frame in the video sequence based on the visual features of the current video frame or the visual features of the current video frame and its previous historical video frames. The future frame may be the next frame of the current video frame.

[0049] The encoder 10 and the memory unit 20 have the same functions in the training phase and the application phase of the detection model, while the processor 40 has slightly different functions in the training phase and the application phase of the detection model. Figure 1 and 2 , respectively, the training scheme and application scheme of the detection model are explained.

[0050] During the training phase of the detection model, the main focus is on optimizing the performance of the encoder 10 and the memory unit 20 . Figure 3 A flow chart of a model training method provided by an exemplary embodiment of the present application, refer to Figure 3 , the method comprising:

[0051] Step 300: Input a normal video sample under the target scene into a detection model for video anomaly detection;

[0052] Step 301: Using the detection model, respectively predict visual features of sample future frames of at least one sample video frame in the normal video sample; and calculate the visual feature error between the at least one sample future frame and its corresponding sample video frame in the normal video sample;

[0053] Step 302: The detection model is trained with the first constraint condition that the fluctuation state of the visual feature error in the time domain reaches the target state, so that the detection model obtains model parameters under the target scene.

[0054] The detection model and model training scheme provided in this embodiment can be applied to various scenarios requiring video anomaly detection, such as security supervision, emergency management, and traffic supervision. It can accurately monitor abnormal events in various application scenarios, such as fires, explosions, and traffic violations. This embodiment does not limit the application scenarios.

[0055] In this embodiment, detection models can be trained separately for different scenarios to improve their performance. Of course, this embodiment is not limited to this; similar scenarios can also share the same detection model. The background content in the video varies across different scenarios. This article will use the target scenario as an example to illustrate the detection model training process. It should be understood that the target scenario can be any desired scenario.

[0056] In step 300, normal videos of the target scene can be used as training samples, referred to as normal video samples. Normal video samples can be used as the video sequences mentioned above. It is understood that normal video samples refer to videos without abnormal events. In this embodiment, only normal videos are used as training samples so that the detection model can fully learn the characteristics of normal videos.

[0057] In step 301, the encoder 10 in the detection model can be used to extract visual features of at least one sample video frame in the normal video sample, and the memory unit 20 can be used to respectively predict visual features of a sample future frame of the at least one sample video frame. In this way, visual features of at least one sample video frame and its corresponding sample future frame can be obtained. The sample future frame predicted by the memory unit 20 can be time-aligned with at least one video frame in the normal video sample to determine the corresponding relationship between the sample video frame and the sample future frame. Figure 2 , sample video frame f2 and sample future frame Alignment, that is, the sample future frame corresponding to the sample video frame f2 is

[0058] Based on this, the processor 40 may calculate a visual feature error between at least one sample future frame and its corresponding sample video frame in the normal video sample. The visual feature error may be a mean square error (MSE), although this embodiment is not limited thereto. The visual feature error is used to characterize the difference between the sample video frame and its corresponding sample future frame.

[0059] After obtaining the visual feature error between at least one sample future frame and its corresponding sample video frame in the normal video sample, in step 302, the detection model can be trained with the fluctuation state of the visual feature error in the time domain reaching the target state as the first constraint condition, so that the detection model can obtain model parameters under the target scene.

[0060] The target state may be that the fluctuation of the visual feature error in the time domain is minimal. Of course, this embodiment is not limited thereto, and the target state may also be other effect states that can reflect the normality of the video.

[0061] In practical applications, refer to Figure 2 , the standard deviation of the visual feature error in the time domain can be calculated as the time domain fluctuation loss function (such as Figure 2 L in σ2 ); The detection model is trained with the time domain fluctuation loss function converging to the minimum as the first constraint. In practical applications, Figure 2 The L2 distance in is used to characterize the visual feature error. Of course, this embodiment is not limited to this.

[0062] Furthermore, in this embodiment, the temporal fluctuation state of visual feature errors can be constrained by region. As mentioned above, encoder 10 can employ a convolutional neural network for visual feature extraction. Based on this, in this embodiment, the coverage position of the convolution kernel described above, that is, the field of view during a single convolution, can be considered as a spatial position. Thus, a normal video sample will have at least one spatial position, and accordingly, a single sample video frame will also have at least one spatial position.

[0063] Taking the first sample video frame from at least one sample video frame as an example, in this embodiment, the encoder 10 may extract visual features of the first sample video frame at a first spatial position. When predicting the visual features of a sample future frame, the memory unit 20 may also use a single spatial position within a single sample video frame as a prediction unit, thereby predicting the visual features of the sample future frame corresponding to the first sample video frame at the first spatial position. The first sample video frame may be any one of the at least one sample video frame, and the first spatial position may be any one of at least one spatial position within a normal sample video frame. This ensures semantic alignment of the sample video frame and the sample future frame in a local area, thereby improving the performance of the memory unit 20. Based on this, the processor 40 may calculate the visual feature error between the first sample video frame and its corresponding sample future frame at the first spatial position. Consequently, the processor 40 may calculate the visual feature error at each spatial position for the at least one sample video frame and its corresponding sample future frame. In practical applications, a fluctuation curve of the visual feature error along the time axis may be constructed for a single spatial position to characterize the fluctuation of the visual feature error in the temporal domain. Of course, this embodiment is not limited to this.

[0064] Based on this, in this embodiment, the first constraint condition can be that the fluctuation state of the visual feature error corresponding to at least one spatial position of at least one sample video frame reaches a target state. Thus, the fluctuation state of the visual feature error in the time domain can be constrained at at least one spatial position.

[0065] Under the constraints of the first constraint, the detection model can continuously perform backpropagation, thereby continuously optimizing the model parameters in the encoder 10 and the memory unit 20, until the fluctuation state of the visual feature error in the time domain determined according to the detection model can reach the target state. In the trained detection model, the feature extraction performance of the encoder 10 and the prediction performance of the visual features of future frames by the memory unit 20 have been optimized. With the cooperation of the encoder 10 and the memory unit 20, the fluctuation of the visual feature error in the time domain corresponding to the normal video in the target scene will tend to the target state and be relatively stable.

[0066] In addition to the first constraint mentioned above, refer to Figure 2 In this embodiment, the detection model can be trained by setting the second constraint condition that the video feature error corresponding to at least one sample video frame meets the first preset requirement.

[0067] In practical applications, such as Figure 2 As shown, the video feature error corresponding to at least one sample video frame can be used as the feature prediction loss function (such as Figure 2 L in f); the detection model is trained with the feature prediction loss function converging to a minimum as the second constraint. Similarly, in this embodiment, the visual feature error corresponding to at least one sample video frame at at least one spatial location can be used as the feature prediction loss function. The semantic alignment of the sample video frame and the sample future frame based on spatial location can effectively reduce the difficulty of predicting future frames, thereby optimizing future frame prediction performance in complex videos.

[0068] In this embodiment, the detection model can be trained using both the first and second constraints described above. The second constraint constrains the semantic information at each spatial location within the sample video frame, enabling the detection model to more effectively predict complex content within the sample video frame. This ensures that, even when a normal video is complex, the temporal fluctuations of the visual feature error remain stable and approach the target state. This prevents complex videos from being misclassified as abnormal.

[0069] refer to Figure 2 In this embodiment, the detection model may further include a decoder 30, which may restore the future frame based on the visual features of the future frame. Based on this, the processor 40 may further calculate the pixel error between at least one sample video frame and its corresponding sample future frame; the detection model is trained with the pixel error corresponding to at least one sample video frame meeting the second preset requirement as the third constraint condition. In practical applications, the pixel error corresponding to at least one sample video frame may be used as a pixel prediction loss function (e.g., Figure 2 L in I ); The detection model is trained with the pixel prediction loss function converging to a minimum as the third constraint. The third constraint can further optimize the performance of the encoder 10 and the memory unit 20.

[0070] The pixel error is in units of pixels, that is, in this embodiment, the pixel error between at least one sample video frame and its corresponding sample future frame in at least one pixel can be calculated, and the convergence of these pixel errors to a minimum is used as the third constraint condition.

[0071] In this embodiment, Figure 2 As shown, the detection model is trained simultaneously with the first, second, and third constraints to optimize the detection model to the desired performance. In this embodiment, during the detection model training process, the temporal and spatial states of visual features can be taken into account simultaneously. Effectively applying temporal and spatial constraints to model training can effectively reduce visual feature errors caused by normal content changes in the video and allow temporal fluctuations to serve as an effective basis for determining whether anomalies occur in the video. Furthermore, effectively decoupling temporal and spatial visual feature analysis can help simplify the model training task.

[0072] The following will continue to explain the application of the detection model in the target scenario.

[0073] In this embodiment, the video to be detected in the target scene can be input into the trained detection model; the trained detection model is used to predict the visual features of future frames of at least one video frame in the video to be detected; the visual feature error between at least one future frame and its corresponding video frame is calculated; if the fluctuation state of the visual feature error in the time domain is abnormal and does not meet the target state, it is determined that the video to be detected is abnormal.

[0074] As mentioned above, the trained detection model, through the cooperation of encoder 10 and memory unit 20, ensures that the temporal fluctuation state of the visual feature error corresponding to normal videos approaches the target state. Therefore, during the application phase of the detection model, abnormal videos can be identified by determining whether the temporal fluctuation state of the visual feature error meets the target state.

[0075] The processing logic of encoder 10 and memory unit 20 during the application phase is identical to that during the training phase and will not be further described here. However, encoder 10 and memory unit 20 now execute the processing logic based on the model parameters obtained after training. Therefore, the visual features they output ensure that the temporal fluctuations of the visual feature errors corresponding to normal video tend to the target state.

[0076] However, there is a difference between the processing logic of the processor 40 in the application stage and the processing logic in the training stage. In the application stage, the processor 40 can be used to determine whether the fluctuation state of the visual feature error in the time domain meets the target state to distinguish abnormal videos. In actual applications, the processor 40 can calculate the reference error value corresponding to the video to be detected based on the visual feature error corresponding to at least one video frame; if there is a target video feature error in the video feature error corresponding to at least one video frame, and the difference from the reference error value exceeds a preset range, it is determined that the video frame to which the target video feature error belongs is abnormal. From the aforementioned spatial position level, the processor 40 can calculate the reference error value of the video to be detected at the first spatial position based on the visual feature error of at least one video frame at the first spatial position; wherein the first spatial position is any one of at least one spatial position in the video to be detected.

[0077] The processor 40 may calculate the mean of the visual feature errors of at least one video frame at the first spatial position as the reference error value for the video to be detected at the first spatial position. Of course, this embodiment is not limited to this. The processor 40 may also calculate the median of the visual feature errors of at least one video frame at the first spatial position as the reference error value for the video to be detected at the first spatial position.

[0078] In practice, for normal videos with simpler content, the reference error value will be slightly smaller than that for more complex content. In either case, the visual feature error corresponding to a normal video will fluctuate slightly around the reference error value. However, anomalous videos have not been trained by the detection model, so the visual feature errors corresponding to anomalous videos will fluctuate more significantly, making them easier to distinguish.

[0079] In this way, the processor 40 can determine whether any of the video feature errors corresponding to at least one video frame contains a target video feature error whose difference from the reference error value exceeds a preset range. If so, the processor 40 can determine that an abnormality exists in the video being detected. Furthermore, the processor can also locate the target video frame containing the abnormality, or even the target spatial location within the target video frame, based on the target video feature error. Using image recognition and other technologies, the processor can then identify the type of abnormal event, such as a fire, explosion, or running a red light.

[0080] In actual applications, an alarm component can be added after the detection model. The alarm component can issue an alarm based on the abnormal detection results, abnormal event types and other information output by the detection model to prompt staff to deal with abnormal events in a timely manner.

[0081] Accordingly, in this embodiment, normal video samples in the target scene can be used as training samples, and the detection model can be used to predict the visual features of the sample future frames of at least one sample video frame in the normal video sample, and the visual feature error between the sample video frame and the corresponding sample future frame can be calculated; the detection model is trained with the constraint that the fluctuation state of the visual feature error in the time domain reaches the target state. This semi-supervised training method can effectively improve the prediction performance of the detection model for future frames. Based on this, in the process of video anomaly detection, the detection model can be used to ensure that the fluctuation state of the visual feature error in the time domain in the normal video in the target scene tends to the target state, while the fluctuation state of the visual feature error in the time domain in the abnormal video will deviate from the target state, thereby more accurately detecting abnormal situations in the video under the target scene.

[0082] Figure 4 A flowchart of a video anomaly detection method provided by an exemplary embodiment of the present application is provided. Figure 4 , the method comprising:

[0083] Step 400: Obtain the video to be detected;

[0084] Step 401: predict visual features of future frames of at least one video frame in the video to be detected;

[0085] Step 402: Calculate the visual feature error between at least one future frame and its corresponding video frame in the video frame to be detected;

[0086] Step 403: If the fluctuation state of the visual feature error in the time domain is abnormal and does not meet the target state, it is determined that the video to be detected has an abnormality;

[0087] The target state is a constraint condition set for the fluctuation state of the visual feature error of normal videos in the time domain during the training of the detection model.

[0088] The video anomaly detection method provided in this embodiment can be implemented based on the detection model in the aforementioned embodiment. The scene corresponding to the video to be detected and the detection model should be as consistent as possible to ensure the accuracy of the detection.

[0089] In an optional embodiment, step 403 may include:

[0090] Calculating a reference error value corresponding to the video to be detected based on the visual feature error corresponding to at least one video frame;

[0091] If there is a target video feature error in the video feature errors corresponding to at least one video frame, and the difference from the reference error value exceeds a preset range, it is determined that the video frame to which the target video feature error belongs has an abnormality.

[0092] In an optional embodiment, the step of calculating a reference error value corresponding to the video to be detected based on the visual feature error corresponding to at least one video frame may include:

[0093] Calculating a reference error value of the video to be detected at the first spatial position based on a visual feature error of each of the at least one video frame at the first spatial position;

[0094] The first spatial position is any one of at least one spatial position in the video to be detected.

[0095] In an optional embodiment, the step of calculating a reference error value of the video to be detected at the first spatial position based on the visual feature error of each of at least one video frame at the first spatial position may include:

[0096] An average of visual feature errors of at least one video frame at the first spatial position is calculated as a reference error value of the video to be detected at the first spatial position.

[0097] The training scheme and application scheme of the detection model have been described in detail in the aforementioned embodiments. Based on this, the technical details in each step of this embodiment can refer to the description of the application scheme of the detection model in the aforementioned embodiments and will not be repeated here, but this should not cause any loss to the scope of protection of this application.

[0098] In one possible design, the above-mentioned video anomaly detection solution can be implemented by a cloud server or cloud platform. Figure 5 A flowchart of another video anomaly detection method provided by an exemplary embodiment of the present application. Figure 5 , the method comprising:

[0099] Step 500: In response to a request to call a target service, determine the processing resources corresponding to the target service, and use the processing resources corresponding to the target service to perform the following steps:

[0100] Step 501: Obtain the video to be detected;

[0101] Step 502: predict visual features of future frames of at least one video frame in the video to be detected;

[0102] Step 503: Calculate the visual feature error between at least one future frame and its corresponding video frame in the video frame to be detected;

[0103] Step 504: If the fluctuation state of the visual feature error in the time domain is abnormal and does not meet the target state, it is determined that the video to be detected is abnormal; wherein the target state is the constraint condition set for the fluctuation state of the visual feature error in the time domain of the normal video during the training process of the detection model.

[0104] In this embodiment, the target service can be deployed on a cloud server or cloud platform, which can receive requests to invoke the target service to provide video anomaly detection capabilities. The target service can include a detection model for video anomaly detection, and the cloud server or cloud platform can also obtain several normal videos of the target scene as training samples to train the detection model included in the target service. This allows the target service to support anomaly detection in the target scene's videos.

[0105] Regarding the video anomaly detection solution and the training solution for the detection model that the target service can provide, please refer to the relevant description in the aforementioned embodiments and will not be repeated here, but this should not cause any loss to the scope of protection of this application.

[0106] It should be noted that the execution entity of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution entity of steps 401 to 403 can be device A; for another example, the execution entity of steps 401 and 402 can be device A, and the execution entity of step 403 can be device B; and so on.

[0107] In addition, in some of the processes described in the above embodiments and the accompanying drawings, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The sequence numbers of the operations, such as 401, 402, etc., are merely used to distinguish between different operations, and the sequence numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that descriptions such as "first" and "second" in this article are used to distinguish different video frames, positions, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to being different types.

[0108] Figure 6 This is a schematic diagram of a computing device provided by an exemplary embodiment of the present application. Figure 6 As shown, the device includes: a memory 60, a processor 61 and a communication component 62.

[0109] The memory 60 is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc.

[0110] The memory 60 may be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0111] The processor 61 is coupled to the memory 60 and is configured to execute the computer program in the memory 60 to:

[0112] Input normal video samples under the target scene into the detection model for video anomaly detection;

[0113] Using the detection model, respectively predicting visual features of a sample future frame of at least one sample video frame in the normal video sample; and calculating a visual feature error between the at least one sample future frame and its corresponding sample video frame in the normal video sample;

[0114] The detection model is trained with the first constraint that the fluctuation state of the visual feature error in the time domain reaches the target state, so that the detection model can obtain the model parameters under the target scene.

[0115] In an optional embodiment, the target state is that the fluctuation of the visual feature error in the time domain is minimized.

[0116] In an optional embodiment, the processor 61 is configured to: when training the detection model with the first constraint that the fluctuation state of the visual feature error in the time domain reaches the target state,

[0117] Calculate the standard deviation of the visual feature error in the time domain as the time domain fluctuation loss function;

[0118] The detection model is trained with the time domain fluctuation loss function converging to the minimum as the first constraint.

[0119] In an optional embodiment, when calculating the visual feature error between at least one sample future frame and its corresponding sample video frame in the normal video sample, the processor 61 is configured to:

[0120] extracting visual features of a first sample video frame at a first spatial position;

[0121] Obtaining visual features of a sample future frame corresponding to the first sample video frame at a first spatial position;

[0122] Calculating a visual feature error between a first sample video frame and its corresponding sample future frame at a first spatial position;

[0123] The first sample video frame is any one of at least one sample video frame, and the first spatial position is any one of at least one spatial position included in a normal sample video frame.

[0124] In an optional embodiment, when the first constraint condition is that the fluctuation state of the visual feature error in the time domain reaches the target state, the processor 61 is configured to:

[0125] The first constraint condition is that the fluctuation state of the visual feature error corresponding to at least one sample video frame at at least one spatial position reaches a target state.

[0126] In an optional embodiment, the processor 61 is further configured to:

[0127] The detection model is trained with the second constraint condition that the video feature error corresponding to at least one sample video frame meets the first preset requirement.

[0128] In an optional embodiment, when training the detection model with the video feature error corresponding to at least one sample video frame meeting the first preset requirement as the second constraint condition, the processor 61 is configured to:

[0129] Using the video feature error corresponding to at least one sample video frame as the feature prediction loss function;

[0130] The detection model is trained with the feature prediction loss function converging to the minimum as the second constraint.

[0131] In an optional embodiment, the processor 61 is further configured to:

[0132] calculating a pixel error between at least one sample video frame and its corresponding sample future frame;

[0133] The detection model is trained with the third constraint condition that the pixel error corresponding to at least one sample video frame meets the second preset requirement.

[0134] In an optional embodiment, when training the detection model with the pixel error corresponding to at least one sample video frame meeting the second preset requirement as the third constraint condition, the processor 61 is configured to:

[0135] Using the pixel error corresponding to each of at least one sample video frame as a pixel prediction loss function;

[0136] The detection model is trained with the pixel prediction loss function converging to the minimum as the third constraint.

[0137] In an optional embodiment, after training the detection model, the processor 61 is further configured to:

[0138] Input the video to be detected in the target scene into the trained detection model;

[0139] Using the trained detection model, respectively predict visual features of future frames of at least one video frame in the video to be detected; and calculate the visual feature error between the at least one future frame and its corresponding video frame;

[0140] If the fluctuation state of the visual feature error in the time domain is abnormal and does not meet the target state, it is determined that the video to be detected has an abnormality.

[0141] In an optional embodiment, if the fluctuation state of the visual feature error in the time domain is abnormal and does not meet the target state, the processor 61 is configured to:

[0142] Calculating a reference error value corresponding to the video to be detected based on the visual feature error corresponding to at least one video frame;

[0143] If there is a target video feature error in the video feature errors corresponding to at least one video frame, and the difference from the reference error value exceeds a preset range, it is determined that the video frame to which the target video feature error belongs has an abnormality.

[0144] In an optional embodiment, when calculating the reference error value corresponding to the video to be detected based on the visual feature error, the processor 61 is configured to:

[0145] Calculating a reference error value of the video to be detected at the first spatial position based on a visual feature error of each of the at least one video frame at the first spatial position;

[0146] The first spatial position is any one of at least one spatial position in the video to be detected.

[0147] In an optional embodiment, when calculating the reference error value of the video to be detected at the first spatial position based on the visual feature error of each of at least one video frame at the first spatial position, the processor 61 is configured to:

[0148] An average of visual feature errors of at least one video frame at the first spatial position is calculated as a reference error value of the video to be detected at the first spatial position.

[0149] Further, if Figure 6 As shown, the computing device also includes: a communication component 62, a power supply component 63 and other components. Figure 6 Only some components are shown schematically, and it does not mean that the computing device only includes Figure 6 Components shown.

[0150] It is worth noting that the technical details in the above-mentioned embodiments of the computing device can be referred to the relevant description in the aforementioned method embodiment. In order to save space, they will not be repeated here, but this should not cause any loss of the scope of protection of this application.

[0151] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed, can implement the steps that can be executed by a computing device in the above method embodiment.

[0152] above Figure 6 The communication component in is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0153] above Figure 6The power supply component in a device provides power to various components of the device in which the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component is located.

[0154] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0155] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0156] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0157] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0158] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0159] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0160] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0161] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0162] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included in the protection scope of the present application.

Claims

1. A model training method, characterized in that: include: Inputting a normal video sample in the target scene into a detection model for video anomaly detection, wherein the normal video sample refers to a video without abnormal events; Using the detection model, respectively predict visual features of a sample future frame of at least one sample video frame in the normal video sample; and calculate a visual feature error at at least one spatial position between the at least one sample future frame and its corresponding sample video frame in the normal video sample; and calculate a pixel error between the at least one sample video frame and its corresponding sample future frame; The first constraint condition is that the fluctuation state of the visual feature error corresponding to at least one sample video frame at at least one spatial position in the time domain reaches the target state, and the first constraint condition is used to constrain the fluctuation state of the visual feature error in the time domain. The second constraint condition is that the video feature error corresponding to at least one sample video frame at at least one spatial position reaches a first preset requirement, and the second constraint condition is used to constrain the semantic information of the sample video frame at at least one spatial position. The third constraint condition is that the pixel error corresponding to each of the at least one sample video frames reaches a second preset requirement. The detection model is trained so that the detection model obtains the model parameters under the target scene.

2. The method according to claim 1, characterized in that The target state is that the fluctuation of the visual feature error in the time domain is minimized.

3. The method according to claim 2, characterized in that The detection model is trained based on the first constraint condition that the fluctuation state of the visual feature error in the time domain reaches the target state, including: Calculating the standard deviation of the visual feature error in the time domain as a time domain fluctuation loss function; The detection model is trained with the time domain fluctuation loss function converging to a minimum as the first constraint condition.

4. The method according to claim 1, wherein The calculating of the visual feature error between the at least one sample future frame and its corresponding sample video frame in the normal video sample includes: extracting visual features of a first sample video frame at a first spatial position; Obtaining visual features of a sample future frame corresponding to the first sample video frame at the first spatial position; Calculating a visual feature error between the first sample video frame and its corresponding sample future frame at the first spatial position; The first sample video frame is any one of the at least one sample video frame, and the first spatial position is any one of the at least one spatial position included in the normal sample video frame.

5. The method according to claim 4, characterized in that The first constraint condition of taking the fluctuation state of the visual feature error in the time domain to reach the target state includes: The first constraint condition is that the fluctuation state of the visual feature error corresponding to the at least one sample video frame at the at least one spatial position reaches the target state.

6. The method according to claim 1, characterized in that The training of the detection model based on the second constraint condition that the video feature error corresponding to the at least one sample video frame meets the first preset requirement includes: Using the video feature error corresponding to the at least one sample video frame as a feature prediction loss function; The detection model is trained with the feature prediction loss function converging to a minimum as the second constraint condition.

7. The method according to claim 1, characterized in that The training of the detection model based on the third constraint condition that the pixel error corresponding to each of the at least one sample video frame meets the second preset requirement includes: Using pixel errors corresponding to each of the at least one sample video frame as a pixel prediction loss function; The detection model is trained with the pixel prediction loss function converging to a minimum as the third constraint condition.

8. The method according to claim 1, characterized in that After the detection model is trained, the method further includes: Inputting the video to be detected in the target scene into the trained detection model; Using the trained detection model, respectively predicting visual features of future frames of at least one video frame in the video to be detected; calculating a visual feature error between the at least one future frame and its corresponding video frame; If the fluctuation state of the visual feature error in the time domain is in an abnormal state that does not meet the target state, it is determined that the video to be detected has an abnormality.

9. The method according to claim 8, characterized in that If the fluctuation state of the visual feature error in the time domain is in an abnormal state that does not meet the target state, determining that the video to be detected has an abnormality includes: Calculating a reference error value corresponding to the video to be detected based on the visual feature errors corresponding to each of the at least one video frame; If, among the video feature errors corresponding to the at least one video frame, there is a target video feature error whose difference from the reference error value exceeds a preset range, it is determined that the video frame to which the target video feature error belongs is abnormal.

10. The method according to claim 9, characterized in that The step of calculating a reference error value corresponding to the video to be detected based on the visual feature error includes: Calculating a reference error value of the video to be detected at the first spatial position based on the visual feature error of each of the at least one video frame at the first spatial position; The first spatial position is any one of at least one spatial position in the video to be detected.

11. The method according to claim 10, characterized in that Calculating a reference error value of the video to be detected at the first spatial position based on the visual feature error of each of the at least one video frame at the first spatial position includes: An average of the visual feature errors of each of the at least one video frame at the first spatial position is calculated as a reference error value of the video to be detected at the first spatial position.

12. A video anomaly detection method, suitable for use in a detection model, characterized in that: include: Get the video to be detected; respectively predicting visual features of future frames of at least one video frame in the video to be detected; Calculating a visual feature error between the at least one future frame and its corresponding video frame in the to-be-detected video frame; If the fluctuation state of the visual feature error in the time domain is in an abnormal state that does not meet the target state, it is determined that the video to be detected has an abnormality; The target state is a constraint condition set for the fluctuation state of the visual feature error of a normal video in the time domain during the training of the detection model, and the detection model is trained according to the method according to any one of claims 1 to 11.

13. The method according to claim 12, characterized in that If the fluctuation state of the visual feature error in the time domain is in an abnormal state that does not meet the target state, determining that the video to be detected has an abnormality includes: Calculating a reference error value corresponding to the video to be detected based on the visual feature errors corresponding to each of the at least one video frame; If, among the video feature errors corresponding to the at least one video frame, there is a target video feature error whose difference from the reference error value exceeds a preset range, it is determined that the video frame to which the target video feature error belongs is abnormal.

14. The method according to claim 13, characterized in that The step of calculating a reference error value corresponding to the video to be detected based on the visual feature error includes: Calculating a reference error value of the video to be detected at the first spatial position based on the visual feature error of each of the at least one video frame at the first spatial position; The first spatial position is any one of at least one spatial position in the video to be detected.

15. The method according to claim 14, characterized in that Calculating a reference error value of the video to be detected at the first spatial position based on the visual feature error of each of the at least one video frame at the first spatial position includes: An average of the visual feature errors of each of the at least one video frame at the first spatial position is calculated as a reference error value of the video to be detected at the first spatial position.

16. A video anomaly detection method, characterized in that: include: In response to a request to call a target service, determining a processing resource corresponding to the target service, and performing the following steps using the processing resource corresponding to the target service: Obtain the video to be detected through the detection model; Predicting visual features of future frames of at least one video frame in the video to be detected by using a detection model; Calculating, by means of a detection model, a visual feature error between the at least one future frame and its corresponding video frame in the video frame to be detected; and determining that an abnormality exists in the video to be detected if a fluctuation state of the visual feature error in the time domain does not satisfy a target state. The target state is a constraint condition set for the fluctuation state of the visual feature error of a normal video in the time domain during the training of the detection model, and the detection model is trained according to the method according to any one of claims 1 to 11.

17. A computing device, characterized in that The computing device deploys a detection model, wherein the detection model includes an encoder, a memory unit, and a processor; The encoder is used to obtain a video to be detected; and extract visual features of at least one video frame in the video to be detected; The memory unit is configured to respectively predict visual features of future frames of the at least one video frame; The processor is configured to calculate a visual feature error between the at least one future frame and its corresponding video frame in the video frame to be detected; if a fluctuation state of the visual feature error in the time domain is abnormal and does not meet a target state, determining that an abnormality exists in the video to be detected; The target state is a constraint condition set for the fluctuation state of the visual feature error of a normal video in the time domain during the training of the detection model, and the detection model is trained according to the method according to any one of claims 1 to 11.

18. The computing device according to claim 17, wherein: The memory unit is configured to, when predicting visual features of a future frame of the at least one video frame: predicting visual features of a future frame of the first video frame at a first spatial position; The processor, when calculating the visual feature error between the at least one future frame and its corresponding video frame in the to-be-detected video frame, is configured to: Calculate the visual feature error between the first video frame and the corresponding future frame at the first spatial position.

19. The computing device according to claim 17, wherein: Also included is a decoder configured to restore the at least one future frame based on visual features of the at least one future frame; The processor is further configured to optimize the future frame prediction performance of the memory unit based on a constraint condition that a pixel error between the at least one video frame and its corresponding future frame meets a preset requirement.

20. A computing device, characterized in that including memory and processor; The memory is used to store one or more computer instructions; The processor is coupled to the memory and configured to execute the one or more computer instructions for: Inputting a normal video sample in the target scene into a detection model for video anomaly detection, wherein the normal video sample refers to a video without abnormal events; Using the detection model, respectively predict visual features of a sample future frame of at least one sample video frame in the normal video sample; and calculate a visual feature error at at least one spatial position between the at least one sample future frame and its corresponding sample video frame in the normal video sample; and calculate a pixel error between the at least one sample video frame and its corresponding sample future frame; The first constraint condition is that the fluctuation state of the visual feature error corresponding to the at least one sample video frame at at least one spatial position in the time domain reaches the target state, and the second constraint condition is that the video feature error corresponding to the at least one sample video frame reaches a first preset requirement. The second constraint condition is used to constrain the semantic information of the sample video frame at at least one spatial position, and the third constraint condition is that the pixel error corresponding to each of the at least one sample video frames reaches a second preset requirement. The detection model is trained so that the detection model obtains the model parameters under the target scene.

21. A computer-readable storage medium storing computer instructions, characterized in that: When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the model training method described in any one of claims 1-11 or the video anomaly detection method described in any one of claims 12-16.

Citation Information

Patent Citations

  • Abnormal behavior detection method based on generative adversarial network

    CN110705376A