Abnormal event detection method and device, equipment, storage medium and product
Optimizing video features and clustering through deep neural network encoder and expectation maximization algorithm, the problem of detection deviation in the prior art is solved and more accurate abnormal event detection is achieved.
Patent Information
- Application Number
- CN202510042597.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Existing abnormal event detection techniques are prone to detection deviations, especially the number of clusters needs to be pre-specified in advance, resulting in distribution estimation deviations. The detection model relies on the distribution likelihood function of normal sample estimation, and it is easy to misjudgment that samples similar to normal samples but are actually abnormal are normal.
The video event samples are mapped to feature space using a deep neural network encoder, clustering and deep neural network encoder are optimized through the expected maximization algorithm, and a priori distribution is learned to improve the accuracy and adaptability of clustering.
By iteratively optimizing clustering and deep neural network encoder, video features can be extracted more accurately, the ability to distinguish abnormal events from normal events, reduce detection deviations, and improve detection accuracy.
Smart Images

Figure CN119964055A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to an abnormal event detection method, device, equipment, storage medium and product. Background Art
[0002] The determination of abnormal events in real-life scenarios depends on people's expected pattern of event occurrence, which is often highly subjective and abstract. However, existing studies have shown that video data can be regarded as sampled from a mixed distribution of multiple event types, and the expected pattern subjectively defined by humans can be characterized by the common prior distribution of event types.
[0003] When solving the prior distribution, the existing abnormal event detection technology solutions need to pre-specify the number of clusters in the clustering algorithm. The specified value usually does not match the actual number of events, resulting in distribution estimation bias. In addition, the detection model completely relies on the distribution likelihood function estimated by normal samples to identify abnormal events, which can easily misjudge abnormal samples that have a certain similarity with normal samples as normal, thus causing detection bias. Summary of the invention
[0004] The main purpose of the present application is to provide an abnormal event detection method, device, equipment, storage medium and product, aiming to solve the technical problem that existing abnormal event detection methods are prone to detection deviation.
[0005] To achieve the above objectives, the present application proposes a method for detecting abnormal events, the method comprising:
[0006] Based on the deep neural network encoder, video features and initial clusters are obtained by mapping video event samples to feature space;
[0007] Based on the expectation maximization algorithm, obtaining optimized clusters, updated deep neural network encoders and prior distributions according to the video features and the initial clusters;
[0008] Abnormal events in a video are detected based on the optimized clustering, the updated deep neural network encoder, and the prior distribution.
[0009] In one embodiment, the video features include normal event features and abnormal event features, and the step of obtaining optimized clusters, updating deep neural network encoders and prior distributions according to the video features and the initial clusters based on the expectation maximization algorithm includes:
[0010] Based on the expectation maximization algorithm, according to the relative relationship between the video features and the initial prior distribution, and the relative relationship between the video features and the initial clusters, the best clusters are obtained, wherein the best clusters include the best normal clusters and the best abnormal clusters;
[0011] Based on the normal event features, the abnormal event features, the optimal clustering and the preset prior distribution, the deep neural network encoder and the optimal clustering are optimized according to the comparison clustering loss and the prior consistency constraint loss to obtain an optimization result;
[0012] When the optimization result meets the convergence condition, the optimized clustering, the updated deep neural network encoder and the prior distribution are obtained.
[0013] In one embodiment, the step of obtaining the best clustering based on the expectation maximization algorithm according to the relative relationship between the video features and the initial prior distribution, and the relative relationship between the video features and the initial clustering, wherein the best clustering includes the best normal clustering and the best abnormal clustering, includes:
[0014] Based on the expectation maximization algorithm, combined with the a priori-guided cluster search mechanism, the relative relationship between the video feature and the initial a priori distribution, as well as the relative relationship between the video feature and the initial cluster are determined to obtain a determination result;
[0015] Based on the judgment result, if the measure of the normal event feature and the nearest distance to the initial cluster center is less than the measure of the normal event feature and the initial prior distribution center, the normal event feature is assigned to the initial cluster, and the initial cluster is used as the best normal cluster;
[0016] If the measure of the normal event feature and the center of the initial prior distribution is less than the measure of any distance from the center of the initial cluster, a new cluster is established for the normal event feature, and the new cluster is used as the best normal cluster;
[0017] Selecting the abnormal event feature with the highest abnormal score as the abnormal candidate feature, and obtaining the best abnormal cluster according to the abnormal candidate feature;
[0018] Based on the best normal cluster and the best abnormal cluster, an optimal cluster corresponding to the video feature is obtained.
[0019] In one embodiment, the step of optimizing the deep neural network encoder and the optimal clustering based on the normal event features, the abnormal event features, the optimal clustering and the preset prior distribution according to the comparison clustering loss and the prior consistency constraint loss to obtain the optimization result includes:
[0020] The normal event feature with the lowest anomaly score is selected as the normal candidate feature;
[0021] Increasing the similarity between the normal candidate features and the best normal cluster and reducing the similarity between the normal candidate features and the best abnormal cluster based on the contrast cluster loss, to obtain a first cluster optimization result;
[0022] Based on the comparative clustering loss, reducing the similarity between the abnormal candidate features and the best normal cluster, and increasing the similarity between the abnormal candidate features and the best abnormal cluster, to obtain a second clustering optimization result;
[0023] Optimizing the deep neural network encoder and the preset prior distribution according to the prior consistency constraint loss to obtain an encoder optimization result and a prior distribution optimization result;
[0024] An optimization result is obtained based on the first clustering optimization result, the second clustering optimization result, the encoder optimization result and the prior distribution optimization result.
[0025] In one embodiment, the step of detecting abnormal events in a video according to the optimized clustering, the updated deep neural network encoder and the prior distribution comprises:
[0026] Extracting a feature vector of the video using the updated deep neural network encoder;
[0027] Based on the optimized clustering, calculating the distance between the feature vector and the prior distribution to obtain an anomaly score;
[0028] Abnormal events in the video are obtained according to the abnormality score.
[0029] In one embodiment, the step of obtaining video features and initial clusters by mapping video event samples to feature space based on a deep neural network encoder includes:
[0030] Based on the deep neural network encoder, the initial video features are obtained by mapping event samples to the feature space;
[0031] Normalizing the initial video features using the Euclidean norm to obtain video features; and obtaining initial clusters by performing preliminary clustering on the video features.
[0032] In addition, to achieve the above objectives, the present application also proposes an abnormal event detection device, the abnormal event detection device comprising:
[0033] A feature extraction module is used to obtain video features and initial clusters by mapping video event samples to feature space based on a deep neural network encoder;
[0034] A parameter iteration module, used for obtaining optimized clusters, updated deep neural network encoders and prior distributions according to the video features and the initial clusters;
[0035] An anomaly detection module is used to detect abnormal events in the video based on the optimized clusters, the updated deep neural network encoder and the prior distribution.
[0036] In addition, to achieve the above-mentioned purpose, the present application also proposes an abnormal event detection device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, and the computer program is configured to implement the steps of the abnormal event detection method described above.
[0037] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the abnormal event detection method described above are implemented.
[0038] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the abnormal event detection method described above are implemented.
[0039] The technical solution proposed in this application is based on a deep neural network encoder. By mapping video event samples to a feature space, video features and initial clusters are obtained. Based on the expectation-maximization algorithm, optimized clusters, updated deep neural network encoders and prior distributions are obtained according to video features and initial clusters. Abnormal events in the video are detected based on the optimized clusters, updated deep neural network encoders and prior distributions. Through the expectation-maximization algorithm, this application can iteratively optimize the initial clusters and deep neural network encoders according to video features and initial clusters, and learn prior distributions, thereby improving clustering accuracy and being able to better adapt to data characteristics, thereby more accurately extracting video features, better distinguishing normal events from abnormal events during anomaly detection, and improving detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0042] Figure 1 A flowchart of the first embodiment of the abnormal event detection method of the present application is provided;
[0043] Figure 2 A flowchart of the second embodiment of the abnormal event detection method of the present application is provided;
[0044] Figure 3 This is a flowchart of the expectation maximization algorithm for this application;
[0045] Figure 4 A flowchart of the third embodiment of the abnormal event detection method of the present application is provided;
[0046] Figure 5 This is a schematic diagram of the module structure of the abnormal event detection device according to an embodiment of the present application;
[0047] Figure 6 Schematic diagram of the device structure of the hardware operating environment involved in the abnormal event detection method in the embodiment of the present application.
[0048] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0049] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0050] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0051] When solving the prior distribution, the existing abnormal event detection technology solutions need to pre-specify the number of clusters in the clustering algorithm. The specified value usually does not match the actual number of events, resulting in distribution estimation bias. In addition, the detection model completely relies on the distribution likelihood function estimated by normal samples to identify abnormal events, which can easily misjudge abnormal samples that have a certain similarity with normal samples as normal, thus causing detection bias.
[0052] Therefore, in order to overcome the above-mentioned defects, the present application provides a solution, which, through the expectation maximization algorithm, can iteratively optimize the initial clustering and the deep neural network encoder according to the video features and the initial clustering, and learn the prior distribution, thereby improving the clustering accuracy and being able to better adapt to the data characteristics, thereby more accurately extracting video features, better distinguishing normal events from abnormal events during the anomaly detection process, and improving the accuracy of detection.
[0053] It should be noted that the execution subject of each embodiment of the present application can be a computing service system with data processing, network communication and program running functions, such as an electronic system, abnormal event detection system or model that can realize the above functions. The following takes the abnormal event detection model as an example (hereinafter referred to as "model") to illustrate the following embodiments.
[0054] Based on this, the present application embodiment provides a method for detecting abnormal events. Figure 1 , Figure 1 This is a flow chart of the first embodiment of the abnormal event detection method of the present application.
[0055] In this embodiment, the abnormal event detection method includes steps S10 to S30:
[0056] Step S10, based on a deep neural network encoder, obtains video features and initial clusters by mapping video event samples to a feature space.
[0057] It should be noted that the deep neural network encoder is a deep learning model that compresses input data into a low-dimensional feature representation through a multi-layer neural network. This feature representation can be used for subsequent classification, regression or clustering tasks; the feature space is a high-dimensional space in which each dimension represents a feature, in which similar data points (such as video frames with similar attributes) are clustered together; video features are representative information extracted from video event samples, such as edges, textures, colors, shapes, etc., which can be used to describe the attributes and patterns of video content; initial clustering refers to the clustering center and its corresponding data point set formed according to certain initial conditions (such as random selection or based on certain rules) before the clustering algorithm starts.
[0058] It can be understood that a deep neural network encoder is used to map video event samples (which may contain video frames, video clips, or other forms of data) to a feature space, thereby extracting representative video features. After the video event samples are mapped to the feature space, they can be clustered or classified according to the video features to obtain initial clusters, which will serve as the starting point for the subsequent optimization process.
[0059] Step S20, based on the expectation-maximization algorithm, obtaining optimized clusters, updated deep neural network encoders and prior distributions according to the video features and the initial clusters.
[0060] It should be noted that the expectation-maximization (EM) algorithm proposed in this application is an iterative optimization algorithm, which is often used in problems with hidden variables, especially when estimating model parameters. In this application, the prior distribution solution and abnormal event detection are both based on the expectation-maximization algorithm, which mainly includes two steps:
[0061] Step E (Expectation Step): Under the current parameter estimation, calculate the expected value of the initial clustering parameters (i.e., estimate the distribution of the latent variables). Specifically, this step calculates the conditional probability distribution of the initial clustering parameters through the current parameter estimation, that is, given the observed data and the current parameters, estimate the expected value of the initial clustering parameters.
[0062] M step (Maximization Step): Based on the expected value of the initial clustering parameters obtained in the E step, the likelihood function is maximized and the parameters of the model are updated, that is, the estimated value of the initial clustering parameters obtained in the E step is used to re-estimate the parameters so that the deep neural network encoder can better fit the observed data.
[0063] It should be understood that the E-step and the M-step of the expectation-maximization algorithm are continuously iterated alternately, and when the iterative optimization reaches the expected convergence condition, the optimized clustering, the updated deep neural network encoder and the prior distribution can be obtained.
[0064] It is understandable that, since this application involves hidden variables (i.e., parameters of initial clustering) set based on subjective prior assumptions, the traditional maximum likelihood estimation method is difficult to use, and directly solving the maximum likelihood estimation may be very complicated or even impossible to solve. The expectation maximization algorithm simplifies this process by dividing it into steps (E step and M step), so that the deep neural network encoder can be indirectly optimized under the condition of initial clustering, making the calculation of each step more efficient, especially in large-scale data sets and complex models, the expectation maximization algorithm can usually effectively find a local optimal solution.
[0065] Step S30, detecting abnormal events in the video according to the optimized clustering, the updated deep neural network encoder and the prior distribution.
[0066] It should be understood that each frame in the video is first converted into a feature vector using the updated deep neural network encoder, and these feature vectors are input into the optimized clusters, so as to determine whether the frame contains normal events or abnormal events. Then, according to the prior distribution, the algorithm is made more sensitive to certain types of abnormal events, while ignoring those events that are unlikely to be abnormal. Among them, the optimized clusters accurately reflect the similarities of different events and can more effectively identify abnormal patterns. The updated deep neural network encoder maps video frames into feature space through its feature extraction capabilities, providing high-quality input for anomaly detection. The prior distribution, as a preliminary assumption or expectation of video event samples, can more accurately judge normal events and abnormal events.
[0067] It should be noted that the present application realizes anomaly detection training under sparse annotation conditions, that is, it proposes a method for distinguishing and screening abnormal events in abnormal videos under weak supervision conditions (using only video-level anomaly annotations). It can complete the learning of normal and abnormal semantic frameworks with only video-level abnormal event annotations, greatly reducing the dependence on the annotation scale and improving the universality of the model.
[0068] This embodiment is based on a deep neural network encoder, and by mapping video event samples to a feature space, video features and initial clusters are obtained. Based on the expectation maximization algorithm, optimized clusters, updated deep neural network encoders and prior distributions are obtained according to the video features and initial clusters, and abnormal events in the video are detected according to the optimized clusters, updated deep neural network encoders and prior distributions. Through the expectation maximization algorithm, the initial clusters and deep neural network encoders can be iteratively optimized according to the video features and initial clusters, and the prior distribution can be learned, which improves the clustering accuracy and can better adapt to data characteristics, thereby extracting video features more accurately, better distinguishing normal events from abnormal events during the abnormality detection process, and improving the accuracy of detection.
[0069] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 2 , the step S20 may include steps S201 to S203:
[0070] Step S201, based on the expectation maximization algorithm, according to the relative relationship between the video features and the initial prior distribution, and the relative relationship between the video features and the initial clusters, the best clusters are obtained, and the best clusters include the best normal clusters and the best abnormal clusters.
[0071] It should be noted that the initial prior distribution is the distribution of beliefs or assumptions about the possible values of a parameter or variable before any data is observed; the best cluster is the cluster that best reflects the intrinsic structure of the data set obtained through the optimization algorithm during the clustering process; the best normal cluster refers to the cluster that contains the characteristics of normal events in the best cluster; the best abnormal cluster refers to the cluster that contains the characteristics of abnormal events in the best cluster.
[0072] It can be understood that in the actual operation of step S201, in order to accurately distinguish and classify normal event features and abnormal event features in the video, it is necessary not only to rely on the direct relationship between the video features themselves and the initial clustering, but also to integrate the additional information provided by the prior distribution to improve the accuracy and robustness of clustering, especially when facing complex and changeable video data. Therefore, a priori-guided cluster search mechanism is introduced to connect the relative relationship between the video feature and the initial prior distribution and the initial cluster. The step S201 may include: based on the expectation maximization algorithm, combined with the priori-guided cluster search mechanism, the relative relationship between the video feature and the initial prior distribution, and the relative relationship between the video feature and the initial cluster are judged to obtain a judgment result; based on the judgment result, if the measure of the normal event feature and the nearest distance to the initial cluster center is less than the measure of the normal event feature and the initial prior distribution center, the normal event feature is assigned to the initial cluster, and the initial cluster is used as the best normal cluster; if the measure of the normal event feature and the initial prior distribution center is less than the measure of any distance to the initial cluster center, a new cluster is established for the normal event feature, and the new cluster is used as the best normal cluster; the abnormal event feature with the highest abnormality score is selected as the abnormal candidate feature, and the best abnormal cluster is obtained according to the abnormal candidate feature; based on the best normal cluster and the best abnormal cluster, the best cluster corresponding to the video feature is obtained.
[0073] It should be understood that the prior-guided clustering search mechanism first uses the core idea of the expectation-maximization algorithm to iteratively update the clustering parameters and model parameters to maximize the likelihood function of the data. In this process, not only the distance from the video features to the centers of each initial cluster is considered (i.e., the metric), but also the degree of deviation of these features from the center of the initial prior distribution is evaluated. This dual consideration ensures that the clustering process can respect the natural distribution of the data and be guided by prior knowledge to a certain extent, thereby avoiding falling into a local optimal solution.
[0074] Specifically, for each normal event feature, the measure of its initial cluster center with the closest distance and the measure of its initial prior distribution center are calculated. If the measure of the normal event feature and the initial cluster center with the closest distance is smaller than the measure of the normal event feature and the initial prior distribution center, it means that the feature is more consistent with the distribution pattern of the existing clusters, so it is assigned to the corresponding initial cluster and the cluster is considered as part of the best normal cluster. On the contrary, if the measure of the feature and the initial prior distribution center is smaller, it means that it may represent a new normal pattern that has not been fully expressed. At this time, a new cluster should be established for it and marked as the best normal cluster.
[0075] At the same time, a strategy based on anomaly scores is adopted for the identification and clustering of abnormal event features. Through the preset scoring criteria or preset scoring model, the features with the highest anomaly scores are screened as abnormal candidate features. These features often deviate from the expectations of the normal mode and indicate potential abnormal events. Subsequently, based on these abnormal candidate features, the expectation maximization algorithm is used for further iterative optimization to form the best abnormal clusters that can closely surround the abnormal features.
[0076] Step S202, based on the normal event features, the abnormal event features, the optimal clustering and the preset prior distribution, according to the comparative clustering loss and the prior consistency constraint loss, optimize the deep neural network encoder and the optimal clustering to obtain an optimization result.
[0077] It should be noted that clustering loss is usually used to evaluate the performance of clustering algorithms, that is, to measure the distance or difference between a data point and the center of its cluster. It can be used in unsupervised learning or semi-supervised learning scenarios, where the labels of the data points are unknown or only partially known. The goal of clustering loss is to make the data points in the same cluster as close as possible, and the data points in different clusters as far apart as possible.
[0078] In addition, prior consistency constraint loss is often used in semi-supervised learning or transfer learning scenarios, where there is prior knowledge or assumptions about the data distribution, and this loss function aims to ensure that the output of the model is consistent with this prior knowledge. For example, in image classification tasks, if the similarity relationship between certain categories is known (such as cats and dogs are more similar than cats and cars), this prior knowledge can be added to the loss function to constrain the output of the model.
[0079] It can be understood that in order to further improve the accuracy and generalization ability of the model, the contrast clustering loss and the prior consistency constraint loss are introduced in step S202 to fine-tune the deep neural network encoder and the optimal clustering to strengthen the association between the normal event features and the optimal normal clustering, while also ensuring that the abnormal event features can be accurately classified into the optimal abnormal clustering.
[0080] Furthermore, the step S202 may also include: selecting the normal event feature with the lowest abnormality score as the normal candidate feature; increasing the similarity between the normal candidate feature and the best normal cluster based on the comparative clustering loss, and reducing the similarity between the normal candidate feature and the best abnormal cluster, to obtain a first clustering optimization result; reducing the similarity between the abnormal candidate feature and the best normal cluster based on the comparative clustering loss, and increasing the similarity between the abnormal candidate feature and the best abnormal cluster, to obtain a second clustering optimization result; optimizing the deep neural network encoder and the preset prior distribution according to the prior consistency constraint loss, to obtain an encoder optimization result and a prior distribution optimization result; obtaining an optimization result based on the first clustering optimization result, the second clustering optimization result, the encoder optimization result and the prior distribution optimization result.
[0081] It can be understood that the features with the lowest anomaly scores are selected from all normal event features. These features are most consistent with the expectations of normal events and are therefore selected as normal candidate features. Next, the contrastive clustering loss is used to enhance the similarity between the normal candidate features and the best normal cluster, while weakening the association with the best abnormal cluster, so that the normal event features are more closely around the best normal cluster in the feature space, thereby enhancing the model's ability to recognize normal events. At the same time, a similar strategy is adopted for the abnormal candidate features, but in the opposite direction, that is, the similarity between the abnormal candidate features and the best normal cluster is reduced, and the association with the best abnormal cluster is increased, ensuring that the abnormal event features can be accurately identified and classified into the best abnormal cluster.
[0082] Furthermore, the prior consistency constraint loss is introduced to further optimize the deep neural network encoder and the preset prior distribution to ensure that the model can not only make full use of the current data information, but also maintain consistency with the prior knowledge during the optimization process, thereby avoiding overfitting and falling into the local optimal solution. Finally, the results of all the above optimization steps, including the first clustering optimization result, the second clustering optimization result, the encoder optimization result and the prior distribution optimization result, are combined to obtain the final optimization result.
[0083] Step S203, when the optimization result meets the convergence condition, obtaining the optimized clustering, the updated deep neural network encoder and the prior distribution.
[0084] It can be understood that when the optimization process reaches the preset convergence condition, the change in the clustering results is less than a preset threshold or reaches the maximum number of iterations, thereby obtaining the optimized clusters; during the iterative optimization process, the deep neural network encoder can adjust its weights and biases according to the input data and the corresponding clustering labels to better capture the intrinsic structure and pattern of the data. When the optimization process converges, a fully trained encoder is obtained, that is, the updated deep neural network encoder; and the prior distribution acts as a constraint or regularization term to affect the clustering results and the training of the encoder. When the optimization process converges, the obtained prior distribution is the version that best matches the current data and clustering results.
[0085] For ease of understanding, refer to Figure 3 This invention is provided for illustration, but is not intended to limit the present application. Figure 3 This is a flow chart of the expectation maximization algorithm of this application. In the E step of the expectation maximization algorithm, the deep neural network encoder (referred to as "encoder") remains frozen and extracts features of the video. Then, a prior-guided clustering search mechanism is applied to determine the clusters of normal event / abnormal event features, and then the prior distribution is updated. In the M step of the expectation maximization algorithm, the encoder is updated by an anomaly-aware clustering method based on prior consistency, which maximizes the log-likelihood of the identified clusters and maps abnormal times to low-probability areas of the prior distribution while ensuring the consistency of the prior.
[0086] exist Figure 3 In the example, the encoder E(·) is first applied to convert normal events / abnormal events V n and V a Mapping to the feature space, we get the video feature X n =E(V n ) and X a =E(V a ), where X n and X a The Euclidean norm (i.e., L2 norm) is used for normalization. Under the Gaussian assumption, the mean parameters of each cluster are determined in the feature space by the expectation maximization algorithm. (N is the number of clusters, M is the number of anomaly categories) and the prior distribution mean parameter μ0, and iteratively updates the encoder E(·).
[0087] In step E, the present application develops a cluster search method inspired by the Dirichlet process, namely, a priori guided cluster search mechanism, which dynamically determines the number of possible clusters in the data set during the training process, and determines the prior distribution through these cluster parameters, breaking through the limitation of the prior art that requires hyperparameters to specify the number of events. The specific algorithm of the priori guided cluster search mechanism can be:
[0088]
[0089] in, and There are the following forms:
[0090]
[0091]
[0092]
[0093] Where n j is the number of elements in the jth cluster; n -i,j is removed from the jth cluster The number of elements after ; It refers to the feature vector obtained after the encoder feature extraction of the i-th slice in the normal event video; refers to the cluster center of the jth normal cluster; It is a metric proposed in this application to express the distance of the feature vector of a video slice relative to the principal component cluster; It is also a metric proposed in this application, which is used to indicate the distance between the feature vector of the video slice and the center of the prior distribution; α is a pre-set hyperparameter; x 1 Refers to any feature vector in the cluster. In addition, for the jth abnormal cluster Using the features of each abnormal video with an abnormal label j The K video features with the highest D0 (indicating a low probability density of the feature in the prior) are selected as anomaly candidates and added to the queue Q. j After traversal, abnormal clustering By Q j The mean of is updated to estimate abnormal clusters, which enhances the subsequent learning of the prior distribution and helps to improve the distinction between normal and abnormal events.
[0094] In the M-step, the main task is to maximize the likelihood of clustering parameters while ensuring prior consistency. The encoder E(·) is updated through a novel prior-consistent anomaly-aware clustering, which includes two losses: anomaly-aware contrastive clustering and prior consistency constraint.
[0095] In the abnormal perception comparison clustering loss link: Since there is no frame-level label provided for the abnormal fragment, we first choose Select the segment features with the highest potential (i.e., the features most likely to correspond to abnormal segments, mainly measured by abnormal scores in this application) as abnormal candidates. Similar to step E, select K video features with the highest D0 as abnormal candidates At the same time, select K video features with the lowest D0 as normal candidates For normal video features x n And normal candidates By comparing the cluster loss ProtoNCE increases the similarity between it and the normal clusters and reduces the similarity between it and the abnormal clusters. By comparing clustering losses, the similarity between it and normal clusters is reduced, and the similarity between it and abnormal clusters is increased. In this way, with the help of abnormal center initialization in step E and abnormal perception comparative clustering, abnormal clips are successfully screened and the difference between them and normal samples is increased by relying on the relationship between samples and prior distribution in videos with only video-level annotations.
[0096] In the prior consistency constraint loss link, the similarity between normal samples relative to the unified prior μ0 is enhanced by contrasting cluster losses to help the encoder capture global semantics and reduce the model's tendency to overfit to specific clusters.
[0097] It can be understood that the prior distribution learning bias problem is solved by the prior consistent anomaly-aware clustering method: the separation between normal and abnormal instances is enhanced by maximizing the likelihood of the identified clusters and minimizing the likelihood of anomalies in the prior distribution, while capturing the global semantics through the prior consistency constraint. This process allows us to learn an unbiased prior distribution that satisfies the normal event type to be located in its high density area while making the abnormal event type located in its low density area, thereby overcoming the boundary bias.
[0098] This embodiment obtains the best clustering based on the expectation maximization algorithm according to the relative relationship between the video features and the initial prior distribution, and the relative relationship between the video features and the initial clustering. The best clustering includes the best normal clustering and the best abnormal clustering. Based on the normal event features, the abnormal event features, the best clustering and the preset prior distribution, the deep neural network encoder and the best clustering are optimized according to the comparison clustering loss and the prior consistency constraint loss to obtain the optimization result. When the optimization result meets the convergence condition, the optimized clustering, the updated deep neural network encoder and the prior distribution are obtained, so as to accurately distinguish between normal clusters and abnormal clusters, ensure that high-quality clusters and encoders are obtained after convergence, and improve the accuracy of abnormality detection.
[0099] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 4 , the step S30 may include steps S301 to S303:
[0100] Step S301, using the updated deep neural network encoder to extract a feature vector of the video.
[0101] It can be understood that a trained and updated deep neural network encoder is used to process videos. The updated deep neural network encoder can receive videos as input and extract key features in the video through its internal multi-layer neural network structure. These features are encoded into vector form in high-dimensional space, called feature vectors, which can capture visual information such as motion, color, texture, etc. in the video.
[0102] Step S302: Based on the optimized clustering, the distance between the feature vector and the prior distribution is calculated to obtain an abnormality score.
[0103] The feature vector of the video is evaluated by the clustering results obtained by the previous optimization (i.e., optimized clusters) and the prior distribution. First, the feature vector is compared with the cluster center and its distance from each cluster center is calculated. Then, the difference between the feature vector and the normal distribution is further evaluated in combination with the prior distribution. This difference is quantified as an anomaly score to measure the degree of abnormality of the video.
[0104] Step S303: obtaining abnormal events in the video according to the abnormality score.
[0105] Abnormal events in the video are identified based on the anomaly scores obtained in the above steps. First, a threshold can be set to distinguish between normal and abnormal videos, where the selection of the threshold may be based on the statistical characteristics of the training data, domain knowledge, or experimental verification. Then, the anomaly score of each video is compared with the threshold. If the anomaly score is higher than the threshold, the video is considered to contain an abnormal event, otherwise, it is considered to be normal.
[0106] For example, in an open-set scenario, the input anomaly may not conform to the distribution of any anomaly cluster. In this case, we detect anomalies by directly measuring the degree of violation between the input video and the prior distribution, that is, the distance between the input and the prior, and we can get the following anomaly score:
[0107] score open (x t )=1-x t ·μ0 (4)
[0108] In the formula, x t refers to the feature vector corresponding to the test video; μ0 is the center of the prior distribution; score open means x t The distance relative to μ0 represents the anomaly score of the video when performing open set anomaly detection.
[0109] It should be noted that the present application has achieved the most advanced results in the abnormal event detection experiment under the open world setting, and has performed performance tests on the abnormal event detection dataset UCF-Crime. The test set and training set of the dataset contain the same 13 types of abnormal events. Some categories of anomalies are removed from the training dataset so that some anomalies in the test set are invisible to the model, so as to construct an anomaly detection scenario in the open world. The abnormal event detection method of the present application has achieved the anomaly detection results in the open set scenario setting on the UCF-Crime dataset as shown in Table 1 below. The first row indicates the number of visible anomaly classes in the training set. From the performance comparison in the table, it can be seen that the method of the present application has achieved higher performance under any visible anomaly setting, which is far superior to other previous methods in this field. Compared with the method of using K-means for prior distribution estimation, the method of the present application has also achieved a greater advantage, which effectively illustrates the effectiveness of the abnormal event detection method of the present application.
[0110] Table 1 Test results of abnormal event detection methods
[0111]
[0112]
[0113] This embodiment uses an updated deep neural network encoder to extract feature vectors of videos, calculates the distance between feature vectors and prior distributions based on optimized clustering, obtains anomaly scores, and obtains abnormal events in the video based on the anomaly scores. This enables the optimized deep neural network encoder to accurately extract features of video clips, and calculates anomaly scores based on the distance between optimized clusters and prior distributions, thereby achieving efficient and accurate detection of abnormal events in video clips.
[0114] In the third embodiment, the step S10 may include steps S101 to S103:
[0115] Step S101, based on a deep neural network encoder, initial video features are obtained by mapping event samples to feature space.
[0116] Step S102: normalize the initial video features using the Euclidean norm to obtain video features.
[0117] It should be noted that the Euclidean norm is a standard method to measure the length of a vector, and the calculation formula is the square root of the sum of the squares of the elements of the vector; the normalization process is to divide each element of the initial video feature by its Euclidean norm to obtain a feature vector with a unit norm.
[0118] In this step, the initial video features obtained in the previous step are normalized. Normalization is a commonly used data preprocessing technique that can scale the norm (or length) of the feature vector to a specific range, usually between 0 and 1 or with a unit norm. The length of the feature vector is measured using the Euclidean norm, and normalized accordingly. The normalized feature vector is the video feature.
[0119] Step S103, obtaining initial clusters by performing preliminary clustering on the video features.
[0120] It should be noted that clustering is an unsupervised learning method that can classify similar samples into one category, thereby revealing the intrinsic structure and distribution of the data.
[0121] It is understandable that the normalized video features can be initially clustered using a clustering algorithm (such as K-means, DBSCAN, etc.) to obtain preliminary clustering results, namely initial clusters, which represent the natural grouping of video features in the feature space.
[0122] This embodiment, based on a deep neural network encoder, maps event samples to a feature space to obtain initial video features, uses the Euclidean norm to normalize the initial video features to obtain video features, and obtains initial clusters by performing preliminary clustering on the video features. This allows the deep neural network encoder to efficiently extract features of video event samples, and enhances feature stability through normalization. The preliminary clustering step lays the foundation for subsequent optimization, thereby improving the accuracy and efficiency of abnormal event detection as a whole.
[0123] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the abnormal event detection method of the present application. More simple transformations based on this technical concept are all within the protection scope of the present application.
[0124] This application also provides an abnormal event detection device, please refer to Figure 5 , the abnormal event detection device comprises:
[0125] A feature extraction module 10, configured to obtain video features and initial clusters by mapping video event samples to a feature space based on a deep neural network encoder;
[0126] A parameter iteration module 20, configured to obtain optimized clusters, updated deep neural network encoders and prior distributions according to the video features and the initial clusters;
[0127] The anomaly detection module 30 is used to detect abnormal events in the video based on the optimized clustering, the updated deep neural network encoder and the prior distribution.
[0128] The abnormal event detection device provided by the present application adopts the abnormal event detection method in the above embodiment, which can solve the technical problem that the existing abnormal event detection method is prone to detection deviation. Compared with the prior art, the beneficial effects of the abnormal event detection device provided by the present application are the same as the beneficial effects of the abnormal event detection method provided by the above embodiment, and the other technical features in the abnormal event detection device are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0129] The present application provides an abnormal event detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the abnormal event detection method in the above-mentioned embodiment 1.
[0130] Reference below Figure 6 , which shows a schematic diagram of the structure of an abnormal event detection device suitable for implementing the embodiment of the present application. The abnormal event detection device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Desctions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The abnormal event detection device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0131] like Figure 6As shown, the abnormal event detection device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the abnormal event detection device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the abnormal event detection device to communicate wirelessly or wired with other devices to exchange data. Although the abnormal event detection device with various systems is shown in the figure, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or provided instead.
[0132] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0133] The abnormal event detection device provided by the present application adopts the abnormal event detection method in the above embodiment, which can solve the technical problem that the existing abnormal event detection method is prone to detection deviation. Compared with the prior art, the beneficial effects of the abnormal event detection device provided by the present application are the same as the beneficial effects of the abnormal event detection method provided by the above embodiment, and the other technical features in the abnormal event detection device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0134] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0135] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0136] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the abnormal event detection method in the above-mentioned embodiment.
[0137] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0138] The computer-readable storage medium may be included in the abnormal event detection device; or may exist independently without being assembled into the abnormal event detection device.
[0139] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the abnormal event detection device, the abnormal event detection device: based on a deep neural network encoder, obtains video features and initial clusters by mapping video event samples to a feature space; based on an expectation-maximization algorithm, obtains optimized clusters, updated deep neural network encoders and prior distributions according to the video features and initial clusters; and detects abnormal events in the video according to the optimized clusters, updated deep neural network encoders and prior distributions.
[0140] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0141] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0142] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.
[0143] The readable storage medium provided by the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned abnormal event detection method, and can solve the technical problem that the existing abnormal event detection method is prone to detection deviation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as the beneficial effects of the abnormal event detection method provided by the above-mentioned embodiment, and will not be repeated here.
[0144] The present application also provides a computer program product, including a computer program, which implements the steps of the abnormal event detection method as described above when executed by a processor.
[0145] The computer program product provided by the present application can solve the technical problem that the existing abnormal event detection method is prone to detection deviation. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the abnormal event detection method provided by the above embodiment, which will not be repeated here.
[0146] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or directly / indirectly applied in other related technical fields, are included in the patent protection scope of the present application.
Claims
1. A method for detecting abnormal events, characterized in that: The method comprises the following steps: Based on the deep neural network encoder, video features and initial clusters are obtained by mapping video event samples to feature space; Based on the expectation maximization algorithm, obtaining optimized clusters, updated deep neural network encoders and prior distributions according to the video features and the initial clusters; Abnormal events in a video are detected based on the optimized clustering, the updated deep neural network encoder, and the prior distribution.
2. The abnormal event detection method according to claim 1, characterized in that: The video features include normal event features and abnormal event features. The step of obtaining optimized clusters, updating deep neural network encoders and prior distributions based on the video features and the initial clusters based on the expectation maximization algorithm includes: Based on the expectation maximization algorithm, according to the relative relationship between the video features and the initial prior distribution, and the relative relationship between the video features and the initial clusters, the best clusters are obtained, wherein the best clusters include the best normal clusters and the best abnormal clusters; Based on the normal event features, the abnormal event features, the optimal clustering and the preset prior distribution, the deep neural network encoder and the optimal clustering are optimized according to the comparison clustering loss and the prior consistency constraint loss to obtain an optimization result; When the optimization result meets the convergence condition, the optimized clustering, the updated deep neural network encoder and the prior distribution are obtained.
3. The abnormal event detection method according to claim 2, characterized in that: The step of obtaining the best clustering based on the expectation maximization algorithm according to the relative relationship between the video features and the initial prior distribution, and the relative relationship between the video features and the initial clustering, wherein the best clustering includes the best normal clustering and the best abnormal clustering, comprises: Based on the expectation maximization algorithm, combined with the a priori-guided cluster search mechanism, the relative relationship between the video feature and the initial a priori distribution, as well as the relative relationship between the video feature and the initial cluster are determined to obtain a determination result; Based on the judgment result, if the measure of the normal event feature and the nearest distance to the initial cluster center is less than the measure of the normal event feature and the initial prior distribution center, the normal event feature is assigned to the initial cluster, and the initial cluster is used as the best normal cluster; If the measure of the normal event feature and the center of the initial prior distribution is less than the measure of any distance from the center of the initial cluster, a new cluster is established for the normal event feature, and the new cluster is used as the best normal cluster; Selecting the abnormal event feature with the highest abnormal score as the abnormal candidate feature, and obtaining the best abnormal cluster according to the abnormal candidate feature; Based on the best normal cluster and the best abnormal cluster, an optimal cluster corresponding to the video feature is obtained.
4. The abnormal event detection method according to claim 2, characterized in that: The step of optimizing the deep neural network encoder and the optimal clustering based on the normal event features, the abnormal event features, the optimal clustering and the preset prior distribution according to the comparison clustering loss and the prior consistency constraint loss to obtain the optimization result includes: The normal event feature with the lowest anomaly score is selected as the normal candidate feature; Increasing the similarity between the normal candidate features and the best normal cluster and reducing the similarity between the normal candidate features and the best abnormal cluster based on the contrast cluster loss, to obtain a first cluster optimization result; Based on the comparative clustering loss, reducing the similarity between the abnormal candidate features and the best normal cluster, and increasing the similarity between the abnormal candidate features and the best abnormal cluster, to obtain a second clustering optimization result; Optimizing the deep neural network encoder and the preset prior distribution according to the prior consistency constraint loss to obtain an encoder optimization result and a prior distribution optimization result; An optimization result is obtained based on the first clustering optimization result, the second clustering optimization result, the encoder optimization result and the prior distribution optimization result.
5. The abnormal event detection method according to any one of claims 1 to 4, characterized in that: The step of detecting abnormal events in the video according to the optimized clustering, the updated deep neural network encoder and the prior distribution comprises: Extracting a feature vector of the video using the updated deep neural network encoder; Based on the optimized clustering, calculating the distance between the feature vector and the prior distribution to obtain an anomaly score; Abnormal events in the video are obtained according to the abnormality score.
6. The abnormal event detection method according to any one of claims 1 to 4, characterized in that: The step of obtaining video features and initial clustering by mapping video event samples to feature space based on a deep neural network encoder includes: Based on the deep neural network encoder, the initial video features are obtained by mapping event samples to the feature space; Normalizing the initial video features using the Euclidean norm to obtain video features; Initial clusters are obtained by performing preliminary clustering on the video features.
7. An abnormal event detection device, characterized in that: The abnormal event detection device comprises: A feature extraction module is used to obtain video features and initial clusters by mapping video event samples to feature space based on a deep neural network encoder; A parameter iteration module, used for obtaining optimized clusters, updated deep neural network encoders and prior distributions according to the video features and the initial clusters; An anomaly detection module is used to detect abnormal events in the video based on the optimized clusters, the updated deep neural network encoder and the prior distribution.
8. An abnormal event detection device, characterized in that: The abnormal event detection device includes: a memory, a processor, and an abnormal event detection program stored in the memory and executable on the processor. When the abnormal event detection program is executed by the processor, the abnormal event detection method according to any one of claims 1 to 6 is implemented.
9. A storage medium, characterized in that: An abnormal event detection program is stored on the storage medium, and when the abnormal event detection program is executed by the processor, the abnormal event detection method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that The computer program product comprises an abnormal event detection program, and when the abnormal event detection program is executed by a processor, the abnormal event detection method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Lane information extraction method, device and equipment and storage medium
CN111341103A
Abnormal behavior detection method based on video monitoring
CN111680614A
Ethereum abnormal transaction behavior detection method based on unsupervised machine learning
CN118114179A