Detection Method, Device, Equipment and Storage Medium for Abnormal Events in Surveillance Videos
By training the event detection model, combining the self-coding network and direction gradient histogram features, the label frequency and biased positive examples and label-free sample learning methods are used to solve the problem of manual label consumption in monitoring video abnormal event detection, and efficient and fast abnormal event detection is achieved.
Patent Information
- Application Number
- CN202110396821.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-04-13
AI Technical Summary
The existing methods for detecting abnormal events in surveillance video require a large amount of precise manual labeling, resulting in serious human and material consumption and affecting practical applications.
By training the event detection model, combining the central object features of labeled and labelless image data, using the self-encoding network and directional gradient histogram features, combining label frequency and biased positive examples and labelless sample learning methods to quickly and accurately detect abnormal events.
It realizes that no large amount of accurate manual labeling is required, and can efficiently and quickly process large-scale monitoring video data, accurately detect abnormal events, reduce manpower and material consumption, and improve detection efficiency.
Smart Images

Figure CN115205773B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video surveillance, and in particular, to a method, device, equipment and storage medium for detecting abnormal events in surveillance videos. Background Art
[0002] With the continuous development of video processing technology, the detection technology of abnormal events in surveillance videos has gradually become a research hotspot. The detection technology of abnormal events in surveillance videos can use a computer to achieve real-time detection of surveillance videos. When a suspicious abnormal event appears, it reminds the surveillance personnel to conduct a secondary manual confirmation. In this way, the vast majority of normal events can be filtered out, and there is no need for the surveillance personnel to continuously monitor the screen content, overcoming the physiological characteristics of the surveillance personnel such as inherent mental fatigue and inattention, greatly improving the work efficiency of the surveillance personnel, and also shortening the reaction time when an abnormal event occurs, and reducing the harm of abnormal events to a certain extent.
[0003] Surveillance video data has the characteristics of high dimension and redundancy. Therefore, the main difficulty of the method for detecting abnormal events in surveillance videos lies in how to efficiently and quickly process large-scale data. For the detection of abnormal events in surveillance videos, a large amount of accurate manual annotation is required, which causes a large amount of manpower and material resources consumption and seriously affects the practical application of the method for detecting abnormal events. Summary of the Invention
[0004] The present invention provides a method, device, equipment and storage medium for detecting abnormal events in surveillance videos, which is used to solve the defect that the method for detecting abnormal events in surveillance videos in the prior art requires a large amount of accurate manual annotation, resulting in a large amount of manpower and material resources consumption, and realizes that there is no need for a large amount of accurate manual annotation, can efficiently and quickly process large-scale data, and accurately detect abnormal events in surveillance videos.
[0005] The present invention provides a method for detecting abnormal events in surveillance videos, including: obtaining a target surveillance video image; obtaining the posterior probability of the central object of the target surveillance video image through an event detection model, where the event detection model is obtained by training based on a historical surveillance video including normal event labels; correcting the posterior probability of the central object of the target surveillance video image through label frequency, where the label frequency is obtained by processing the central object features extracted from the historical surveillance video using a method of learning with biased positive examples and unlabeled samples, and the central object is the image area where the moving target is located; detecting whether the target surveillance video image includes an abnormal event according to the correction result.
[0006] According to the method for detecting abnormal events in surveillance videos provided by the present invention, before correcting the posterior probability of the central object of the target surveillance video image through label frequency, it further includes: obtaining the historical surveillance video, extracting labeled image data and unlabeled image data by performing moving object extraction on the historical surveillance video; extracting the central object features of the labeled image data, and extracting the central object features of the unlabeled image data; processing the central object features of the labeled image data and the central object features of the unlabeled image data through a method of learning with biased positive examples and unlabeled samples to obtain the label frequency.
[0007] According to the method for detecting abnormal events in surveillance videos provided by the present invention, before obtaining the posterior probability of the central object of the target surveillance video image through the event detection model, and after extracting the central object features of the labeled image data and extracting the central object features of the unlabeled image data, it further includes: training an event detection model according to the central object features of the labeled image data and the central object features of the unlabeled image data, where the event detection model is a binary classification model.
[0008] According to the method for detecting abnormal events in surveillance videos provided by the present invention, performing moving object extraction on the historical surveillance video to obtain labeled image data and including unlabeled image data, includes: extracting the moving object data of the video frames with normal event labels in the historical surveillance video to obtain the labeled image data; extracting the moving object data of the video frames without labels in the historical surveillance video to obtain the unlabeled image data.
[0009] According to the method for detecting abnormal events in surveillance videos provided by the present invention, extracting the central object features of the labeled image data includes: extracting the autoencoder network features of the central object of the labeled image data; obtaining the previous video frame and the subsequent video frame of the video frame corresponding to the labeled image data in the historical surveillance video, and the histogram of oriented gradients features at the corresponding position of the central object of the labeled image data according to the previous video frame and the subsequent video frame; splicing and combining the autoencoder network features and the histogram of oriented gradients features in a chronological relationship to obtain the central object features of the labeled image data.
[0010] According to the method for detecting abnormal events in surveillance videos provided by the present invention, the label frequencies are obtained by processing the central object features of the labeled image data and the central object features of the unlabeled image data through a method of learning with biased positive examples and unlabeled samples, including: using a hierarchical partitioning method to obtain the attribute sub-domains containing normal event labels among the central object features of the labeled image data and the central object features of the unlabeled image data, where the attribute sub-domains are sample regions formed by partitioning the labeled image data and the unlabeled image data according to feature attributes; counting the total number of all samples and the total number of labeled samples in the attribute sub-domains; and obtaining the label frequencies based on a preset inequality, the total number of all samples, and the total number of labeled samples.
[0011] According to the method for detecting abnormal events in surveillance videos provided by the present invention, it is detected whether the target surveillance video image includes an abnormal event according to the correction result, including: if the posterior probability of the corrected central object is less than a preset threshold, it is determined that the central object of the target surveillance video image is an abnormal event region; if the posterior probability of the corrected central object is greater than or equal to the preset threshold, it is determined that the central object of the target surveillance video image is a normal event.
[0012] The present invention also provides a device for detecting abnormal events in surveillance videos, including: an acquisition module for acquiring a historical surveillance video and a target surveillance video image, where the historical surveillance video includes video frames with normal event labels; a control and processing module for obtaining the posterior probability of the central object of the target surveillance video image through an event detection model; the control and processing module is further configured to correct the posterior probability of the central object of the target surveillance video image through the label frequency; the control and processing module is further configured to detect whether the target surveillance video image includes an abnormal event according to the correction result; where the event detection model is obtained by training the historical surveillance video, and the event detection model is a binary classification model; the label frequency is obtained by processing through a method of learning with biased positive examples and unlabeled samples after extracting the central object features from the historical surveillance video, and the central object is the image region where the moving target is located.
[0013] According to the device for detecting abnormal events in surveillance videos provided by the present invention, the control and processing module is configured to extract labeled image data and unlabeled image data by performing moving target extraction on the historical surveillance video; the control and processing module is further configured to extract the central object features of the labeled image data and extract the central object features of the unlabeled image data; the control and processing module is further configured to process the central object features of the labeled image data and the central object features of the unlabeled image data through a method of learning with biased positive examples and unlabeled samples to obtain the label frequency.
[0014] According to the detection device for abnormal events in surveillance videos provided by the present invention, the control and processing module is used to train an event detection model based on the central object features of the labeled image data and the central object features of the unlabeled image data.
[0015] According to the detection device for abnormal events in surveillance videos provided by the present invention, the control and processing module is used to extract the moving target data of the video frames with normal event labels in the historical surveillance video to obtain the labeled image data, and extract the moving target data of the video frames without labels in the historical surveillance video to obtain the unlabeled image data.
[0016] According to the detection device for abnormal events in surveillance videos provided by the present invention, the control and processing module is used to extract the autoencoder network features of the central object of the labeled image data; the acquisition module is further used to acquire the previous video frame and the next video frame of the video frame corresponding to the labeled image data in the historical surveillance video; the control and processing module is further used to obtain the histogram of oriented gradients features at the corresponding positions of the central object of the labeled image data according to the previous video frame and the next video frame; the control and processing module is further used to splice and combine the autoencoder network features and the histogram of oriented gradients features in a chronological relationship to obtain the central object features of the labeled image data.
[0017] According to the detection device for abnormal events in surveillance videos provided by the present invention, the control and processing module is used to use a hierarchical partitioning method to obtain the attribute subdomains containing normal event labels in the central object features of the labeled image data and the central object features of the unlabeled image data, where the attribute subdomains are sample regions formed by partitioning the labeled image data and the unlabeled image data according to feature attributes; the control and processing module is further used to count the total number of all samples and the total number of labeled samples in the attribute subdomains; the control and processing module is further used to obtain the label frequency based on a preset inequality, the total number of all samples, and the total number of labeled samples.
[0018] According to the detection device for abnormal events in surveillance videos provided by the present invention, the control and processing module is used to determine that the central object of the target surveillance video image is an abnormal event if the posterior probability of the corrected central object is less than a preset threshold; and determine that the central object of the target surveillance video image is a normal event if the posterior probability of the corrected central object is greater than or equal to the preset threshold.
[0019] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of the detection method for abnormal events in surveillance videos as described in any one of the above are implemented.
[0020] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for detecting abnormal events in a monitored video as described in any one of the above are implemented.
[0021] The method, device, equipment and storage medium for detecting abnormal events in a monitored video provided by the present invention can obtain the posterior probability of a target monitored video (such as a real-time monitored video) through an event detection model trained with a historical monitored video including normal event tags, and then correct the posterior probability based on the tag frequency obtained by processing the historical monitored video. According to the correction result, it can be detected whether an abnormal event occurs in the target monitored video. By processing continuous video stream images and video analysis, the present invention can detect abnormal events in a monitored video in real time, so as to transmit the alarm information of the abnormal events to relevant personnel, enabling the relevant personnel to make corresponding processing in time. Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 is a flowchart of the method for detecting abnormal events in a monitored video provided by the present invention;
[0024] Figure 2 is a structural block diagram of the device for detecting abnormal events in a monitored video provided by the present invention;
[0025] Figure 3 is a schematic structural diagram of an electronic device in an example of the present invention. Detailed Embodiments
[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0027] It should be understood that the "embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, the "in an embodiment" or "in one embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner.
[0028] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0029] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the term "connection" should be understood in a broad sense. For example, it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meaning of the above terms in the present invention can be understood according to specific circumstances.
[0030] The following combines Figure 1 to describe the method for detecting abnormal events in surveillance videos of the present invention.
[0031] As Figure 1 shown, the method for detecting abnormal events in surveillance videos of the present invention includes:
[0032] S1: Obtain a target surveillance video image. Among them, the target surveillance video image can be a real-time video image captured by a certain surveillance camera.
[0033] S2: Obtain the posterior probability of the central object of the target surveillance video image through an event detection model. Among them, the event detection model is obtained by training based on the historical surveillance video including normal event labels. The event detection model is a binary classification model, and the central object is the image area where the moving target is located.
[0034] In one embodiment of the present invention, step S2 includes:
[0035] S2-1: Obtain the historical surveillance video, and extract the moving target according to the historical surveillance video to obtain labeled image data and unlabeled image data.
[0036] Specifically, obtain the historical surveillance video, extract the moving targets in the video frames of the historical surveillance video, and construct a semi-supervised data set. Among them, the historical surveillance video is the historical footage recorded by the surveillance camera. The video frame is each frame of the historical surveillance video. The moving target is a person or object with an action behavior in the surveillance video. The semi-supervised data set is a data set of labeled image data and unlabeled image data.
[0037] In an embodiment of the present invention, the extraction of moving object from historical surveillance videos results in labeled image data and unlabeled image data, including:
[0038] Extract the moving object data of the video frames with normal event labels in the historical surveillance videos to obtain labeled image data. Among them, the normal event labels are manually marked. In this embodiment, only a small number of video frames with moving objects can be set with normal event labels, and the following specific embodiments of the present invention are used for abnormal event detection. Since only a small number of labels are set, the cost consumption of manual annotation can be greatly reduced.
[0039] Extract the moving object data in the video frames without labels in the historical surveillance videos to obtain unlabeled image data.
[0040] In this embodiment, the SSD network is used to extract the moving object images in the video frames of the historical surveillance videos. The SSD network can process about 13 frames per second on a single GPU and can accurately detect smaller objects. For constructing a semi-supervised dataset, labeled video data and a large amount of unlabeled video data can be automatically obtained, reducing the workload of manually annotating samples and improving the training effect of the event detection model.
[0041] S2-2: Extract the central object features of the labeled image data and extract the central object features of the unlabeled image data. Among them, the central object features are the features of the image area where the moving object is located.
[0042] In an embodiment of the present invention, step S2-2 includes:
[0043] S2-2-A-1: Extract the autoencoder network features of the central object of the labeled image data.
[0044] S2-2-A-2: Obtain the previous video frame and the next video frame of the video frame corresponding to the labeled image data in the historical surveillance video, and the histogram of oriented gradients features at the corresponding positions of the central object in the labeled image data according to the previous video frame and the next video frame.
[0045] S2-A-3: Concatenate and combine the autoencoder network features and the histogram of oriented gradients features in chronological order to obtain the central object features of the labeled image data.
[0046] S2-2-B-1: Extract the autoencoder network features of the central object of the unlabeled image data.
[0047] S2-2-B-2: Obtain the previous video frame and the next video frame of the video frame corresponding to the unlabeled image data in the historical surveillance video, and the histogram of oriented gradients features at the corresponding positions of the central object in the unlabeled image data according to the previous video frame and the next video frame.
[0048] S2-B-3: Concatenate and combine the auto-encoding network features and the histogram of oriented gradients (HOG) features according to the temporal relationship to obtain the central object features of the labeled image data.
[0049] Specifically, for the central object of the video frames in the historical surveillance video, convolutional auto-encoding network is used to extract features. Herein, the central object is the image area where the moving target is located. This embodiment can inherently learn the potential appearance features. This auto-encoding is based on a lightweight structure and consists of an encoder with 3 convolutional layers and a max-pooling layer, a decoder with 3 upsampling layers and convolutional layers, and an additional convolutional layer for the final output. Considering the continuity characteristic of the video itself, in this embodiment, the HOG method is used to extract features at the corresponding positions of the central object in the previous frame and the next frame in the labeled image data / unlabeled image data, further strengthening the appearance features. The HOG can not only extract appearance features but also contain motion features. The HOG features and the convolutional auto-encoding network features are concatenated and combined in the temporal order to form the central object features.
[0050] S2-3: Train an event detection model according to the central object features of the labeled image data and the central object features of the unlabeled image data. Herein, the event detection model is a binary classification model.
[0051] Specifically, based on the support vector machine classification method, the central object features of the labeled image data and the central object features of the unlabeled image data are used for training to initially construct a binary classification model as the event detection model. Then, the posterior probability of the central object features of each unlabeled image data is corrected using the label frequency and a new pseudo-label is assigned to it. Repeat the above steps to iteratively train the support vector machine model. This method effectively fuses the central object features of the labeled image data and the central object features of the unlabeled image data, and quickly trains a binary classification model with good robustness as the event detection model.
[0052] S3: Correct the posterior probability of the central object of the target surveillance video image through the label frequency. Herein, the label frequency is obtained after processing the historical surveillance video by using the method of learning with biased positive examples and unlabeled samples for central object feature extraction.
[0053] In an embodiment of the present invention, step S3 includes:
[0054] S3-1: Use the hierarchical partitioning method to obtain the attribute sub-domains containing the normal event labels among the central object features of the labeled image data and the central object features of the unlabeled image data. Herein, the attribute domain sub is the sample area formed after partitioning the labeled image data and the unlabeled image data according to the feature attributes.
[0055] S3-2: Count the total number of all examples T and the total number of labeled examples L in the statistical attribute subdomain.
[0056] S3-3: Based on a preset inequality, obtain the label frequency c from the total number of all examples T and the total number of labeled examples L.
[0057] Specifically, based on the Learning from Positive and Unlabled Example (PU) method, when unlabeled data is regarded as negative class samples, a non-traditional probability classifier can be directly trained to obtain the probability Pr(y = 1|x) that a sample is labeled. In the present invention, an abnormal event detection method for correcting the prediction probability by using the label frequency c is proposed, mainly based on the following: First, in the attribute subdomain of positive unlabeled data, the proportion of labeled examples provides a lower bound for the label frequency c; Second, the attribute subdomain with more labeled examples can produce a more accurate estimate of the label frequency c. Use P, L, and T to represent the number of positive examples, the total number of positive labeled samples, and the total number of examples in the semi-supervised dataset respectively. Since the total number of positive labeled samples is unknown, the label frequency c cannot be directly calculated. If almost all examples in an attribute subdomain A are positive examples, then the lower bound value c of the label frequency:
[0058] c≥Pr(s = 1|x∈A) = L / T
[0059] To accurately estimate the lower bound of the label frequency, an attribute subdomain containing as many positive examples as possible needs to be selected. For this purpose, a multi-level partitioning method is adopted in this embodiment to find the attribute subdomain containing as many labeled data as possible.
[0060] In the attribute subdomain, the label frequency c is predicted by using the number of labeled examples and the total number of examples in the attribute subdomain. However, due to the randomness of labeling, more positive examples may be labeled. Therefore, an error term is introduced, which is derived from the Hoeffding inequality. The probability that each positive example is labeled in the attribute subdomain is constant, and the positive examples follow a Bernoulli distribution. Therefore, the expected value of positive examples being selected as labeled examples is μ = cP, and the variance is σ 2 = c(1 - c)P. Then the probability that the number of positive examples selected as labeled examples is not less than K times is:
[0061]
[0062] where ε>0 and K = (c + ε)P, and then let
[0063]
[0064] Also, since P ≤ T, the lower limit value of the label frequency can be obtained:
[0065]
[0066] It can be concluded therefrom that the magnitude of the error term depends on the number of samples included in the attribute subdomain.
[0067] S3-4: Correct the posterior probability of the central object in the target surveillance video image through the label frequency.
[0068] Specifically, the labels and attributes of the labeled image data are completely independent. Under the assumption of completely random selection, with the help of the equation:
[0069]
[0070] where Pr(y = 1|x) represents the probability that the sample is labeled, and Pr(s = 1|x) represents the probability that the sample is labeled after correction. The label frequency c can be directly estimated from the training set, so only the class posterior probability needs to be processed under the assumption of completely random selection. Among them, the posterior probability of the target surveillance video is a value between 0 and 1, where 0 represents an anomaly and 1 represents normal. The present invention corrects the posterior probability of the target surveillance video through the previously obtained label frequency.
[0071] S4: Detect whether the target surveillance video image includes an abnormal event according to the correction result.
[0072] Specifically, it is determined whether it is abnormal based on the posterior probability corrected by the label frequency. If the posterior probability is lower than a predetermined threshold, it is considered that the central object is abnormal, and the position of the central object is restored to the original video frame to achieve the positioning of the abnormal event.
[0073] The detection device for abnormal events in a surveillance video provided by the present invention will be described below. The detection device for abnormal events in a surveillance video described below can be mutually corresponding and referred to the detection method for abnormal events in a surveillance video described above.
[0074] Figure 2 is the structural block diagram of the detection device for abnormal events in a surveillance video provided by the present invention. As Figure 2 shown, the detection device for abnormal events in a surveillance video provided by the present invention includes: an acquisition module 210 and a control and processing module 220.
[0075] The acquisition module 210 is used to acquire historical monitoring videos and target monitoring video images, where the historical monitoring videos include video frames with normal event labels. The control processing module 220 is used to obtain the posterior probability of the central object of the target monitoring video image through an event detection model. The control processing module 220 is also used to correct the posterior probability of the central object of the target monitoring video image according to the label frequency. The control processing module 220 is also used to detect whether the target monitoring video image includes an abnormal event according to the correction result. Among them, the event detection model is obtained by training the historical monitoring video, and the event detection model is a binary classification model. The label frequency is obtained by processing the central object features of the historical monitoring video using a method of learning with biased positive examples and unlabeled samples. The central object is the image area where the moving target is located.
[0076] In an embodiment of the present invention, the control processing module 220 is used to extract labeled image data and unlabeled image data by extracting moving target data from the historical monitoring video. The control processing module 220 is also used to extract the central object features of the labeled image data and extract the central object features of the unlabeled image data. The control processing module 220 is also used to process the central object features of the labeled image data and the central object features of the unlabeled image data by a method of learning with biased positive examples and unlabeled samples to obtain the label frequency.
[0077] In an embodiment of the present invention, the control processing module 220 is used to train an event detection model according to the central object features of the labeled image data and the central object features of the unlabeled image data.
[0078] In an embodiment of the present invention, the control processing module 220 is used to extract the moving target data of the video frames with normal event labels in the historical monitoring video to obtain labeled image data, and extract the moving target data of the video frames without labels in the historical monitoring video to obtain unlabeled image data.
[0079] In an embodiment of the present invention, the control processing module 220 is used to extract the auto-encoder network features of the central object of the labeled image data. The acquisition module 210 is also used to acquire the previous video frame and the subsequent video frame of the video frame corresponding to the labeled image data in the historical monitoring video. The control processing module 220 is also used to obtain the histogram of oriented gradients features at the corresponding position of the central object of the labeled image data according to the previous video frame and the subsequent video frame. The control processing module 220 is also used to splice and combine the auto-encoder network features and the histogram of oriented gradients features in a chronological relationship to obtain the central object features of the labeled image data.
[0080] In an embodiment of the present invention, the control processing module 220 is configured to use a hierarchical partitioning method to obtain an attribute sub-domain containing a normal event label among the central object features of the labeled image data and the central object features of the unlabeled image data. Wherein, the attribute sub-domain is a sample area formed by partitioning the labeled image data and the unlabeled image data according to feature attributes. The control processing module 220 is further configured to count the total number of all samples and the total number of labeled samples in the attribute sub-domain. The control processing module 220 is further configured to obtain a label frequency based on a preset inequality, the total number of all samples, and the total number of labeled samples.
[0081] In an embodiment of the present invention, if the posterior probability of the corrected central object is less than a preset threshold, the control processing module 220 determines that the central object of the target surveillance video image is an abnormal event; if the posterior probability of the corrected central object is greater than or equal to the preset threshold, the control processing module 220 determines that the central object of the target surveillance video image is a normal event.
[0082] It should be noted that the specific implementation of the surveillance video abnormal event detection device in the embodiment of the present invention is similar to the specific implementation of the surveillance video abnormal event detection method in the embodiment of the present invention. For details, please refer to the description in the part of the surveillance video abnormal event detection method. To reduce redundancy, it will not be elaborated here.
[0083] In addition, the other components and functions of the surveillance video abnormal event detection device in the embodiment of the present invention are known to those skilled in the art. To reduce redundancy, it will not be elaborated here.
[0084] Figure 3 This is a schematic structural diagram of an electronic device in an example of the present invention. As Figure 3 shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 can call the logical instructions in the memory 330 to execute the surveillance video abnormal event detection method, which includes: obtaining a target surveillance video image; obtaining the posterior probability of the central object of the target surveillance video image through an event detection model, where the event detection model is obtained by training with a historical surveillance video including normal event labels; correcting the posterior probability of the central object of the target surveillance video image through a label frequency, where the label frequency is obtained by processing the central object features extracted from the historical surveillance video using a method of learning with biased positive examples and unlabeled samples, and the central object is the image area where the moving target is located; detecting whether the target surveillance video image includes an abnormal event according to the correction result.
[0085] In an embodiment of the present invention, the processor may be an integrated circuit chip with signal processing capabilities. The processor may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0086] It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. The processor reads the information in the storage medium and combines its hardware to complete the steps of the above method.
[0087] In addition, when the logical instructions in the above-mentioned memory 330 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs, etc., which can store program codes.
[0088] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for detecting abnormal events in a monitored video provided above. The method includes: obtaining a target monitored video image; obtaining the posterior probability of the central object of the target monitored video image through an event detection model, where the event detection model is obtained by training based on a historical monitored video including video frames with normal event labels; correcting the posterior probability of the central object of the target monitored video image through label frequency, where the label frequency is obtained by processing the historical monitored video through a method of learning with biased positive examples and unlabeled samples after extracting central object features, and the central object is the image area where a moving target is located; detecting whether the target monitored video image includes an abnormal event according to the correction result.
[0089] The storage medium may be a memory, for example, it may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories.
[0090] Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory.
[0091] The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).
[0092] The storage medium described in the embodiments of the present invention is intended to include but not be limited to these and any other suitable types of memories.
[0093] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by a combination of hardware and software. When applying software, the corresponding functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0094] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0095] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general-purpose hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A method for detecting abnormal events in surveillance videos, characterized in that, Including: Obtain a target monitored video image; Obtain the posterior probability of the central object in the target monitored video image through an event detection model, where the event detection model is obtained by training based on a historical monitored video of video frames including normal event labels, the event detection model is a binary classification model, and the central object is the image area where the moving target is located; Revise the posterior probability of the central object in the target monitored video image through label frequency: Obtain the attribute subdomain of the normal event label by using the hierarchical partitioning method, and based on the preset inequality calculate the lower bound of the label frequency, and then revise the posterior probability, where c is the label frequency, T is the total number of all examples in the attribute subdomain, L is the total number of label examples in the attribute subdomain; wherein, the label frequency is obtained after processing the central object features through the historical monitored video by using the method of learning with biased positive examples and unlabeled samples; Detect whether the target monitored video image includes an abnormal event according to the correction result.
2. The detection method of abnormal events in surveillance videos according to claim 1, wherein Before correcting the posterior probability of the central object in the target monitored video image through label frequency, it further includes: Obtain the historical monitored video, and extract labeled image data and unlabeled image data by performing moving target extraction on the historical monitored video; Extract the central object features of the labeled image data, and extract the central object features of the unlabeled image data; Process the central object features of the labeled image data and the central object features of the unlabeled image data through a method of learning with biased positive examples and unlabeled samples to obtain the label frequency.
3. The detection method of abnormal events in surveillance videos according to claim 2, characterized in that, Before obtaining the posterior probability of the central object in the target monitored video image through the event detection model, and after extracting the central object features of the labeled image data and extracting the central object features of the unlabeled image data, it further includes: Train an event detection model according to the central object features of the labeled image data and the central object features of the unlabeled image data.
4. The detection method of abnormal events in surveillance videos according to claim 2, characterized in that, Extracting labeled image data and unlabeled image data by performing moving target extraction on the historical monitored video includes: Extract the moving target data of the video frames with normal event labels in the historical monitored video to obtain the labeled image data; Extract the moving target data in the video frames without labels in the historical monitored video to obtain the unlabeled image data.
5. The detection method of abnormal events in surveillance videos according to claim 2, characterized in that, Extracting the central object features of the labeled image data includes: Extract the autoencoder network features of the central object in the labeled image data; Obtain the previous video frame and the next video frame of the video frame corresponding to the labeled image data in the historical monitored video, and obtain the histogram of oriented gradients features at the corresponding position of the central object in the labeled image data according to the previous video frame and the next video frame; Splice and combine the autoencoder network features and the histogram of oriented gradients features in chronological order to obtain the central object features of the labeled image data.
6. The detection method of abnormal events in surveillance videos according to claim 2, wherein Processing the central object features of the labeled image data and the central object features of the unlabeled image data through a method of learning with biased positive examples and unlabeled samples to obtain the label frequency includes: Use a hierarchical partitioning method to obtain the attribute subdomains containing normal event labels in the central object features of the labeled image data and the central object features of the unlabeled image data, where the attribute subdomains are sample regions formed by partitioning the labeled image data and the unlabeled image data according to feature attributes; Count the total number of all samples and the total number of labeled samples in the attribute subdomains; Based on a preset inequality , the label frequency is obtained from the total number of all samples and the total number of labeled samples, where c is the label frequency, T is the total number of all examples in the attribute subdomain, L is the total number of labeled examples in the attribute subdomain.
7. The detection method for abnormal events in surveillance videos according to claim 1, characterized in that Detecting whether the target monitored video image includes an abnormal event according to the correction result includes: If the posterior probability of the corrected central object is less than a preset threshold, it is determined that the central object of the target surveillance video image is an abnormal event; If the posterior probability of the corrected central object is greater than or equal to the preset threshold, it is determined that the central object of the target surveillance video image is a normal event.
8. A detection device for monitoring abnormal events in videos, characterized in that, Including: An acquisition module, configured to acquire a historical surveillance video and a target surveillance video image, where the historical surveillance video includes video frames with normal event labels; A control processing module is configured to obtain the posterior probability of the central object of the target monitored video image through an event detection model; the control processing module is further configured to correct the posterior probability of the central object of the target monitored video image by using the label frequency: obtain the attribute sub-domains of the normal event labels by using a hierarchical partitioning method, and based on a preset inequality calculate the lower bound of the label frequency, and further correct the posterior probability, where c is the label frequency, T is the total number of all examples in the attribute sub-domain, L is the total number of label examples in the attribute sub-domain; the control processing module is further configured to detect whether the target monitored video image includes an abnormal event according to the correction result; Wherein, the event detection model is obtained by training the historical surveillance video; the label frequency is obtained by performing central object feature extraction on the historical surveillance video and then processing it using a method of learning with biased positive examples and unlabeled samples, and the central object is the image area where the moving target is located.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for detecting abnormal events in a surveillance video according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for detecting abnormal events in a surveillance video according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for detecting road traffic abnormal events in real time
CN103971521A
Command information system state monitoring method and system, medium, equipment and terminal
CN112508068A