Abnormal event detection method, device and system, and data processing method

Through the unsupervised algorithm, the processing model and spatiotemporal feature extraction technology are used to detect abnormal events based on the spatiotemporal evolution law of normal events, solving the problem of low detection accuracy caused by insufficient samples in the existing technology, and improving the ability to identify abnormal events.

CN113536855BActive Publication Date: 2025-08-22ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010312900.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-20
Publication Date
2025-08-22
Estimated Expiration
2040-04-20

AI Technical Summary

Technical Problem

In the prior art, the detection method of abnormal events lacks sufficient sample data, resulting in the model's detection accuracy of abnormal events, especially the recognition accuracy of occasional and harmful abnormal events.

Method used

By obtaining the multimedia data of the target object, using the processing model to predict and extract spatiotemporal features, detecting abnormal events based on the spatiotemporal evolution law of normal events, avoiding dependence on abnormal event data, and using an unsupervised algorithm to train the processing model.

Benefits of technology

It improves the accuracy of detection of abnormal events, enhances the ability to perceive low-probability abnormal events, and solves the problem of low accuracy of recognition of occasional and harmful abnormal events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113536855B_ABST
    Figure CN113536855B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device and system for detecting abnormal events, and a data processing method. The detection method includes: obtaining first multimedia data and second multimedia data of a target object; processing the first multimedia data using a processing model to predict the first spatiotemporal features of the second multimedia data, and processing the second multimedia data using the processing model to obtain the second spatiotemporal features of the second multimedia data. The processing model is used to input the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, input the spatiotemporal features of the first multimedia data into a prediction network to obtain the first spatiotemporal features of the second multimedia data; based on the first spatiotemporal features and the second spatiotemporal features, determine whether an abnormal event has occurred in the target object. The present application solves the technical problem of low recognition accuracy of occasional and highly harmful abnormal events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and more specifically, to a method, device, and system for detecting abnormal events, and a data processing method. Background Art

[0002] Currently, traditional anomaly detection methods require collecting a large number of samples for model training to uncover unique patterns of abnormal states and behaviors. However, due to the sudden and sporadic nature of abnormal events, data acquisition is costly, and it is often difficult to collect enough samples to model abnormal patterns. This results in low accuracy in the trained models for anomaly detection. Currently, no effective solution has been proposed to address this issue. Summary of the Invention

[0003] The embodiments of the present application provide a method, device, and system for detecting abnormal events, as well as a data processing method, to at least solve the technical problem of low recognition accuracy of occasional and highly harmful abnormal events.

[0004] According to one aspect of an embodiment of the present application, a method for detecting abnormal events is provided, including: acquiring first multimedia data and second multimedia data of a target object, wherein the first multimedia data and the second multimedia data are multimedia data acquired from video data of the target object, and an acquisition time of the first multimedia data is earlier than an acquisition time of the second multimedia data; processing the first multimedia data using a processing model to predict a first spatiotemporal feature of the second multimedia data, and processing the second multimedia data using the processing model to obtain a second spatiotemporal feature of the second multimedia data, wherein the processing model is used to acquire the first multimedia data and the second multimedia data, inputting the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, and inputting the spatiotemporal features of the first multimedia data into a prediction network to obtain the first spatiotemporal features of the second multimedia data; and determining whether an abnormal event occurs in the target object based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data.

[0005] According to another aspect of an embodiment of the present application, a method for detecting abnormal events is also provided, including: acquiring video data of a target object; processing the video data to obtain first multimedia data and second multimedia data, wherein the acquisition time of the first multimedia data is earlier than the acquisition time of the second multimedia data; processing the first multimedia data using a processing model to predict the first spatiotemporal features of the second multimedia data, and processing the second multimedia data using the processing model to obtain the second spatiotemporal features of the second multimedia data, wherein the processing model is used to acquire the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, and input the spatiotemporal features of the first multimedia data into a prediction network to obtain the first spatiotemporal features of the second multimedia data; based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data, determine whether an abnormal event occurs in the target object.

[0006] According to another aspect of an embodiment of the present application, a device for detecting abnormal events is also provided, including: an acquisition module, used to acquire first multimedia data and second multimedia data of a target object, wherein the first multimedia data and the second multimedia data are multimedia data acquired from video data of the target object, and the acquisition time of the first multimedia data is earlier than the acquisition time of the second multimedia data; a processing module, used to process the first multimedia data using a processing model to predict the first spatiotemporal features of the second multimedia data, and process the second multimedia data using the processing model to obtain the second spatiotemporal features of the second multimedia data, wherein the processing model is used to acquire the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, and input the spatiotemporal features of the first multimedia data into a prediction network to obtain the first spatiotemporal features of the second multimedia data; a determination module, used to determine whether an abnormal event occurs to the target object based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data.

[0007] According to another aspect of an embodiment of the present application, a device for detecting abnormal events is also provided, including: an acquisition module for acquiring video data of a target object; a first processing module for processing the video data to obtain first multimedia data and second multimedia data, wherein the acquisition time of the first multimedia data is earlier than the acquisition time of the second multimedia data; a second processing module for processing the first multimedia data using a processing model to predict the first spatiotemporal features of the second multimedia data, and processing the second multimedia data using the processing model to obtain second spatiotemporal features of the second multimedia data, wherein the processing model is used to acquire the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, and input the spatiotemporal features of the first multimedia data into a prediction network to obtain the first spatiotemporal features of the second multimedia data; a determination module for determining whether an abnormal event occurs in the target object based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data.

[0008] According to another aspect of an embodiment of the present application, a data processing method is also provided, including: obtaining video data of a target object; processing the video data to obtain first multimedia data of a first time period and second multimedia data of a second time period; obtaining a first spatiotemporal feature based on the first multimedia data, wherein the first spatiotemporal feature is a prediction feature of the second time period; obtaining a second spatiotemporal feature based on the second multimedia data, wherein the second spatiotemporal feature is a detection feature of the second time period; and judging whether the target object is in an abnormal state based on the first spatiotemporal feature and the second spatiotemporal feature.

[0009] According to another aspect of an embodiment of the present application, a storage medium is further provided, which includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the above-mentioned abnormal event detection method and data processing method.

[0010] According to another aspect of an embodiment of the present application, a computing device is further provided, including: a processor and a memory, the processor being configured to run a program stored in the memory, wherein the above-mentioned abnormal event detection method and data processing method are executed when the program is running.

[0011] According to another aspect of an embodiment of the present application, a system for detecting abnormal events is also provided, including: a processor; and a memory connected to the processor, for providing the processor with instructions for processing the following processing steps: obtaining first multimedia data and second multimedia data of a target object, wherein the first multimedia data and the second multimedia data are multimedia data obtained from video data of the target object, and the acquisition time of the first multimedia data is earlier than the acquisition time of the second multimedia data; processing the first multimedia data using a processing model to predict the first spatiotemporal features of the second multimedia data, and processing the second multimedia data using the processing model to obtain the second spatiotemporal features of the second multimedia data, wherein the processing model is used to obtain the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, and input the spatiotemporal features of the first multimedia data into a prediction network to obtain the first spatiotemporal features of the second multimedia data; based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data, determine whether an abnormal event occurs in the target object.

[0012] In an embodiment of the present application, after obtaining the first multimedia data and the second multimedia data of the target object, the first multimedia data can be processed using a processing model to predict the first spatiotemporal features of the second multimedia data, and the second multimedia data can be processed using the processing model to obtain the second spatiotemporal features of the second multimedia data. Further, through the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data, it can be determined whether an abnormal event has occurred in the target object. It is easy to notice that abnormal events that do not conform to the "normal change law" can be perceived and predicted by mining the spatiotemporal evolution law of the normal state, and the processing model can be trained without obtaining data on the abnormal events, thereby enhancing the perception ability of abnormal events and improving the detection accuracy of abnormal events with a low probability of occurrence, thereby solving the technical problem of low recognition accuracy of occasional and highly harmful abnormal events. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0014] Figure 1 is a hardware structure block diagram of a computer terminal for implementing a method for detecting abnormal events according to an embodiment of the present application;

[0015] Figure 2 is a flow chart of a method for detecting abnormal events according to an embodiment of the present application;

[0016] Figure 3 is a schematic diagram of an optional abnormal event detection method according to an embodiment of the present application;

[0017] Figure 4 is a flow chart of another abnormal event detection method according to an embodiment of the present application;

[0018] Figure 5 is a schematic diagram of a detection device for abnormal events according to an embodiment of the present application;

[0019] Figure 6 is a schematic diagram of another abnormal event detection device according to an embodiment of the present application;

[0020] Figure 7 is a flow chart of a data processing method according to an embodiment of the present application; and

[0021] Figure 8 This is a structural block diagram of a computer terminal according to an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0023] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0024] Due to the sudden and sporadic nature of abnormal events, it is difficult to collect enough samples to model abnormal patterns. Moreover, due to the diversity of abnormalities and the complexity of the scenarios involved, different types of abnormalities usually require training corresponding models.

[0025] Considering that the evolution of normal events follows certain dynamic laws, while abnormal events are usually unpredictable, this application proposes an abnormal event detection method based on unsupervised deep spatiotemporal feature prediction. A large number of normal samples can be collected, and by mining the evolution laws of normal events, predictions can be made for the future. By comparing the predicted results with the actual situation, it can be determined whether abnormal events have occurred. This method is an unsupervised algorithm that overcomes the defect of supervised algorithms that require the collection and annotation of large amounts of abnormal data.

[0026] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:

[0027] Video clip: A series of consecutive frames in a video.

[0028] CNN (Convolutional Neural Network): Convolutional neural network is a type of feedforward neural network that includes convolution calculations and has a deep structure.

[0029] LSTM (Long Short Term Memory): Long short-term memory model, a time recurrent neural network.

[0030] ResNet (Residual Neural Network): A residual network constructed by residual blocks that use skip connections.

[0031] Example 1

[0032] According to an embodiment of the present application, a method for detecting abnormal events is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0033] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for detecting abnormal events is shown. Figure 1As shown, the computer terminal 10 (or mobile device 10) may include one or more (illustrated as 102a, 102b, ..., 102n) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0034] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0035] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the detection method of abnormal events in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned detection method of abnormal events. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0036] The transmission device 106 is used to receive or send data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0037] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0038] It should be noted that, in some optional embodiments, the above Figure 1 The computer device (or mobile device) shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the aforementioned computer device (or mobile device).

[0039] Under the above operating environment, this application provides Figure 2 The abnormal event detection method shown. Figure 2 FIG. 1 is a flow chart of a method for detecting abnormal events according to an embodiment of the present application. Figure 2 As shown, the method includes the following steps:

[0040] Step S202: Acquire first multimedia data and second multimedia data of a target object, wherein the first multimedia data and the second multimedia data are multimedia data acquired from video data of the target object, and the acquisition time of the first multimedia data is earlier than the acquisition time of the second multimedia data;

[0041] The target objects in the above steps can be objects requiring spatiotemporal anomaly event detection, such as, but not limited to, open flame identification, fire hazards, and explosions in the field of safety supervision. They can also be objects in other application scenarios, such as, for example, objects or areas requiring anti-theft monitoring in the field of security. The target objects in the above steps can also be users requiring behavior monitoring, such as, but not limited to, patients requiring physical status monitoring in the medical field.

[0042] The multimedia data in the above steps (including the first multimedia data and the second multimedia data) can be pictures or video clips taken of the target object. In order to detect abnormal events more accurately, in the embodiment of the present application, the video clips of the target object are used as an example for illustration.

[0043] In this embodiment, the future can be predicted by mining the evolution patterns of normal events, and by comparing the predicted results with the actual situation, it can be determined whether there are abnormal events. On this basis, the video data of the target object can be directly obtained and divided into two parts. That is, the continuous video frames are divided into two parts, the part with the earlier acquisition time is used as the above-mentioned first multimedia data, and the part with the later acquisition time is used as the above-mentioned second multimedia data. Among them, by processing the first multimedia data, the spatiotemporal characteristics of the second multimedia data can be predicted. For example, taking a video clip as an example, the first multimedia data can refer to the preceding video clip, and the second multimedia data can refer to the subsequent video clip.

[0044] Step S204: Processing the first multimedia data using the processing model to predict first spatiotemporal features of the second multimedia data, and processing the second multimedia data using the processing model to obtain second spatiotemporal features of the second multimedia data, wherein the processing model is used to obtain the first multimedia data and the second multimedia data, inputting the first multimedia data and the second multimedia data into the extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, and inputting the spatiotemporal features of the first multimedia data into the prediction network to obtain the first spatiotemporal features of the second multimedia data;

[0045] The existing technology uses a video prediction method to detect abnormal events. The specific implementation method is as follows: by giving T frames of images, the next frame of image is predicted and generated, and by comparing the error between the generated image and the actual image, it is determined whether an abnormality occurs.

[0046] However, this method has the following two problems: first, the images of actual monitoring scenes are complex and diverse, and pixel-level image reconstruction is difficult to converge; second, the global image prediction loss may overwhelm local outliers, thereby reducing the sensitivity of perception of abnormal events.

[0047] To address the above issues, in this embodiment, the spatiotemporal features of the second multimedia data can be reconstructed without reconstructing the second multimedia data itself. This allows the processing model to converge more easily and achieve stronger generalization performance. Furthermore, the focus is placed on mining the laws of spatiotemporal evolution, enhancing the ability to perceive abnormal events. On this basis, the processing model in the above steps can be trained using a large number of collected normal samples. The first spatiotemporal features in the above steps can be the spatiotemporal features of the second multimedia data predicted by mining the spatiotemporal evolution laws of the first multimedia data. The second spatiotemporal features can be spatiotemporal features mined directly from the second multimedia data.

[0048] The aforementioned extraction network can be a model for extracting spatiotemporal features corresponding to multimedia data, and can mine the spatiotemporal evolution patterns corresponding to the multimedia data. The prediction network can be a model for predicting the spatiotemporal features of future multimedia data. Optionally, the prediction model can employ two sets of convolutional layers. The first set of convolutional layers downsamples the number of feature channels to compress the feature space and focus on the essential patterns of feature evolution. The second set of convolutional layers upsamples the number of feature channels to restore the original dimensions, facilitating subsequent feature comparison.

[0049] For example, Figure 3 As shown, still taking the video clip as an example, in an optional embodiment, after obtaining the preceding video clip and the subsequent video clip, the two video clips can be input into the extraction model to respectively mine the spatiotemporal features of the two video clips, and the spatiotemporal features mined out from the preceding video clip are further input into the prediction model, and the spatiotemporal features of the subsequent video clip are predicted by the prediction model, thereby obtaining the predicted spatiotemporal features and the extracted spatiotemporal features.

[0050] Step S206 : determining whether an abnormal event occurs to the target object based on the first spatiotemporal feature of the second multimedia data and the second spatiotemporal feature of the second multimedia data.

[0051] The abnormal events in the above steps can be low-probability abnormal events that occur to the target object, including but not limited to open flame identification, sudden fires, explosions, etc. in the field of safety supervision, theft incidents in the field of security, etc., and can also refer to abnormal behavioral states of the target object.

[0052] In this embodiment, if no abnormal event occurs to the target object, the second spatiotemporal feature of the second multimedia data conforms to the spatiotemporal evolution law of the mined normal event, and therefore, the predicted first spatiotemporal feature and the actual second spatiotemporal feature are the same or similar; if an abnormal event occurs to the target object, the second spatiotemporal feature of the second multimedia data does not conform to the spatiotemporal evolution law of the mined normal event, and therefore, the predicted first spatiotemporal feature and the actual second spatiotemporal feature are quite different.

[0053] For example, Figure 3 As shown, taking a video clip as an example, in an optional embodiment, a preceding video clip and a subsequent video clip captured by a camera or other acquisition device can be obtained, and the two video clips can be input into a processing model. By processing the preceding video clip, the spatiotemporal features of the subsequent video clip can be predicted. At the same time, by processing the subsequent video clip, the spatiotemporal features of the subsequent video clip can be extracted. By comparing the predicted spatiotemporal features with the extracted spatiotemporal features, if the difference between the two spatiotemporal features is large, it can be determined that an abnormal event has occurred with the target object; if the difference between the two spatiotemporal features is small, it can be determined that no abnormal event has occurred with the target object.

[0054] It should be noted that, in order to achieve better convergence, before determining whether an abnormal event occurs in the target object, both spatiotemporal features will be converted into feature vectors through the average pooling layer.

[0055] Through the solution provided by the above-mentioned embodiment of the present application, after obtaining the first multimedia data and the second multimedia data of the target object, the first multimedia data can be processed using a processing model to predict the first spatiotemporal features of the second multimedia data, and the second multimedia data can be processed using the processing model to obtain the second spatiotemporal features of the second multimedia data. Further, through the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data, it can be determined whether an abnormal event has occurred in the target object. It is easy to notice that by mining the spatiotemporal evolution laws of normal states, abnormal events that do not conform to the "normal change laws" can be perceived and predicted, and the processing model can be trained without obtaining data on abnormal events, thereby enhancing the perception ability of abnormal events and improving the detection accuracy of abnormal events with a low probability of occurrence, thereby solving the technical problem of low recognition accuracy of occasional and highly harmful abnormal events.

[0056] In the above embodiment of the present application, the first multimedia data and the second multimedia data are input into the extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, including: inputting the first multimedia data and the second multimedia data into the spatial feature extraction model in the extraction network for feature extraction to obtain the first spatial features of the first multimedia data and the second spatial features of the second multimedia data; inputting the first spatial features into the first spatiotemporal feature extraction model in the extraction network for feature extraction to obtain the spatiotemporal features of the first multimedia data; inputting the second spatial features into the second spatiotemporal feature extraction model in the extraction network for feature extraction to obtain the second spatiotemporal features of the second multimedia data.

[0057] The above-mentioned spatial feature extraction model can be a shared spatial feature extraction model, which can extract spatial features for each frame image in the video clip. Optionally, the spatial feature extraction model can adopt the first three blocks of ResNet18, but is not limited to this, and other deep feature extraction models can also be used.

[0058] The above-mentioned first spatiotemporal feature extraction model and second spatiotemporal feature extraction model can mine the spatiotemporal evolution law of multimedia data. Optionally, in order to enhance the generalization ability of the system, the first spatiotemporal feature extraction model and the second spatiotemporal feature extraction model have the same model structure but do not share parameters. Moreover, the first spatiotemporal feature extraction model and the second spatiotemporal feature extraction model can adopt a convolutional long short-term memory network (ConvLSTM), but are not limited to this. Other spatiotemporal feature mining models, such as 3D convolution, can also be adopted.

[0059] For convolutional long short-term memory networks, the temporal and spatial feature evolution patterns of multimedia data can be mined through the design of gate structures. The specific gate calculation formula is as follows:

[0060]

[0061]

[0062]

[0063]

[0064] For example, Figure 3 As shown, still taking the video clip as an example, in an optional embodiment, after obtaining the preceding video clip and the subsequent video clip, the preceding video clip and the subsequent video stream can respectively use a shared spatial feature extraction model to extract spatial features for each frame image in the video clip; the spatial features obtained from the preceding video clip pass through the spatiotemporal feature extraction model 1 to mine the spatiotemporal features corresponding to the video clip; the spatial features obtained from the subsequent video clip pass through the spatiotemporal feature extraction model 2 to mine the spatiotemporal features corresponding to the video clip.

[0065] In the above embodiment of the present application, based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data, determining whether an abnormal event has occurred with the target object includes: obtaining the characteristic distance between the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data; comparing the characteristic distance with a first preset threshold; if the characteristic distance is greater than the first preset threshold, determining that an abnormal event has occurred with the target object; if the characteristic distance is less than or equal to the preset threshold, determining that no abnormal event has occurred with the target object.

[0066] The aforementioned characteristic distance can be used to represent an outlier between two multimedia data sets. A larger characteristic distance indicates that the second multimedia data is less consistent with the evolution pattern of the first multimedia data, and is more likely to have an abnormal event. The characteristic distance between two spatiotemporal features can be calculated using existing distance calculation methods, such as Euclidean distance, Manhattan distance, Chebyshev distance, angle cosine distance, etc., but this application does not impose any specific limitations on this.

[0067] The above-mentioned first preset threshold can be a characteristic distance threshold that can determine the occurrence of an abnormal event. For example, in the embodiment of the present application, the first preset threshold is 0.4 as an example for illustration, but it is not limited to this and can be set according to the actual application scenario.

[0068] For example, Figure 3 As shown, still taking the video clip as an example, in an optional embodiment, after predicting the spatiotemporal features of the subsequent video clip and extracting the spatiotemporal features of the subsequent video clip, the feature distance between the two spatiotemporal features can be calculated. When the feature distance is greater than a first preset threshold, it is determined that an abnormal event has occurred; otherwise, it is determined that no abnormal event has occurred.

[0069] In the above embodiment of the present application, obtaining the first multimedia data and the second multimedia data of the target object includes: sampling the video data to obtain multiple frames of multimedia data; dividing the multiple frames of multimedia data to obtain the first multimedia data and the second multimedia data.

[0070] The number of frames of the multimedia data can be set according to the actual application scenario. In the embodiment of the present application, 8 frames are used as an example for illustration, but the invention is not limited thereto. Dividing the multi-frame multimedia data can mean dividing the multi-frame multimedia data into two parts evenly, but the invention is not limited thereto.

[0071] For example, Figure 3 As shown, still taking the video clip as an example, in an optional embodiment, a video stream of the target object can be obtained. After the video stream is input, 8 frames of video are extracted per second, the first 4 frames of video are used as the preceding video clip, and the last 4 frames of video are used as the subsequent video clip.

[0072] In the above embodiment of the present application, the method also includes the following steps: obtaining multiple groups of training samples, wherein each group of training samples includes: a first sample and a second sample, wherein the first sample and the second sample are samples obtained from the collected video data, and the collection time of the first sample is earlier than the collection time of the second sample; using the multiple groups of training samples to train the initial processing model to obtain a processing model.

[0073] For example, Figure 3As shown, still taking the video clip as an example, in an optional embodiment, the training process of the processing model is similar to the actual detection process. A large number of normal samples can be collected and divided into first samples and second samples. The constructed initial processing model is trained with a large number of normal samples to obtain a processing model.

[0074] It should be noted that after each abnormal event detection, the processing model may be updated to ensure the processing accuracy of the processing model and thus the detection accuracy.

[0075] In the above embodiment of the present application, the initial processing model is trained using multiple groups of training samples to obtain a processing model, including: inputting each group of training samples into the initial processing model, predicting the first spatiotemporal features of the second sample, and obtaining the second spatiotemporal features of the second sample; obtaining a loss value of the initial processing model based on the first spatiotemporal features of the second sample and the second spatiotemporal features of the second sample; comparing the loss value with a second preset threshold; if the loss value is greater than the second preset threshold, updating the network weights of the initial processing model; if the loss value is less than or equal to the second preset threshold, obtaining the processing model.

[0076] The above-mentioned second preset threshold can be a loss value threshold that meets the training accuracy requirements, and can be set according to the actual application scenario. This application does not make any specific limitations on this.

[0077] For example, Figure 3 As shown, still taking the video clip as an example, in an optional embodiment, the first sample and the second sample are respectively processed by a shared spatial feature extraction model to extract spatial features, the spatial features of the first sample are processed by spatiotemporal feature extraction model 1 to obtain corresponding spatiotemporal features, and the spatial features of the second sample are processed by spatiotemporal feature extraction model 2 to obtain corresponding spatiotemporal features, and the spatiotemporal features mined from the first sample are further processed by a prediction model to predict the spatiotemporal features corresponding to the second sample. Combining the mined spatiotemporal features and the predicted spatiotemporal features, the corresponding loss value is calculated by the loss function. If the loss value does not meet the loss value threshold required by the training accuracy, the network weights can be updated by backpropagation until the loss value meets the loss value threshold required by the training accuracy, thereby obtaining a processing model.

[0078] In the above embodiment of the present application, based on the first spatiotemporal features of the second sample and the second spatiotemporal features of the second sample, the loss value of the initial processing model is obtained, including: obtaining the characteristic distance between the first spatiotemporal features of the second sample and the second spatiotemporal features of the second sample, and obtaining the loss value of the initial processing model.

[0079] In this embodiment, the loss function used by the processing model mainly consists of two parts: the first part is the distance loss of the spatiotemporal feature vector, for example, the mean square error loss can be used; the second part is the scene classification loss based on the predicted spatiotemporal features. By adding this second loss, the learned network parameters have more semantic information, which can not only predict the spatiotemporal features of future video clips, but also reflect the spatiotemporal evolution of the scene. This allows the processing model to discover universal spatiotemporal evolution laws and can be well extended to new application scenarios.

[0080] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0081] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0082] Example 2

[0083] According to an embodiment of the present application, a method for detecting abnormal events is also provided.

[0084] In the operating environment of the above embodiment 1, the present application provides the following Figure 4 The abnormal event detection method shown. Figure 4 FIG. 1 is a flow chart of another abnormal event detection method according to an embodiment of the present application. Figure 4 As shown, the method includes the following steps:

[0085] Step S402, obtaining video data of the target object;

[0086] The target objects in the above steps can refer to objects requiring abnormal event detection, including but not limited to open flame identification, sudden fires, and explosions in the field of safety supervision, or objects and areas requiring anti-theft monitoring in the field of security. The target objects in the above steps can also refer to users requiring behavior monitoring, such as patients requiring physical status monitoring in the medical field, but are not limited to these.

[0087] The video data in the above steps can be multiple images, video clips, etc. taken of the target object. In order to detect abnormal events more accurately, in the embodiment of the present application, the video clips of the target object are used as an example for illustration.

[0088] Step S404: Process the video data to obtain first multimedia data and second multimedia data, wherein the first multimedia data is collected earlier than the second multimedia data;

[0089] In this embodiment, the video data can be processed and divided into two multimedia data, that is, the continuous video frames are divided into two parts, the multimedia data with an earlier acquisition time is the above-mentioned first multimedia data, and the multimedia data with a later acquisition time is the above-mentioned second multimedia data. Among them, by processing the first multimedia data, the prediction result of the second multimedia data can be predicted.

[0090] Step S406: Processing the first multimedia data using the processing model to predict first spatiotemporal features of the second multimedia data, and processing the second multimedia data using the processing model to obtain second spatiotemporal features of the second multimedia data, wherein the processing model is used to obtain the first multimedia data and the second multimedia data, inputting the first multimedia data and the second multimedia data into the extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, and inputting the spatiotemporal features of the first multimedia data into the prediction network to obtain the first spatiotemporal features of the second multimedia data;

[0091] The processing model in the above steps can be obtained by training a large number of normal samples collected. The first spatiotemporal features in the above steps can be the spatiotemporal features of the second multimedia data predicted by mining the spatiotemporal evolution laws of the first multimedia data. The second spatiotemporal features can be the spatiotemporal features directly mined from the second multimedia data.

[0092] The aforementioned extraction network can be a model for extracting spatiotemporal features corresponding to multimedia data, and can mine the spatiotemporal evolution patterns corresponding to the multimedia data. The prediction network can be a model for predicting the spatiotemporal features of future multimedia data. Optionally, the prediction model can employ two sets of convolutional layers. The first set of convolutional layers downsamples the number of feature channels to compress the feature space and focus on the essential patterns of feature evolution. The second set of convolutional layers upsamples the number of feature channels to restore the original dimensions, facilitating subsequent feature comparison.

[0093] Step S408 : determining whether an abnormal event occurs to the target object based on the first spatiotemporal feature of the second multimedia data and the second spatiotemporal feature of the second multimedia data.

[0094] The abnormal events in the above steps can be low-probability abnormal events that occur to the target object, including but not limited to open flame identification, sudden fires, explosions, etc. in the field of safety supervision, theft incidents in the field of security, etc., and can also refer to abnormal behavioral states of the target object.

[0095] Through the solution provided by the above embodiment of the present application, after obtaining the video data of the target object, the video data can be processed to obtain first multimedia data and second multimedia data, and then the first multimedia data can be processed using the processing model to predict the first spatiotemporal features of the second multimedia data, and the second multimedia data can be processed using the processing model to obtain the second spatiotemporal features of the second multimedia data. Finally, through the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data, it can be determined whether an abnormal event has occurred in the target object. It is easy to notice that by mining the spatiotemporal evolution laws of normal states, abnormal events that do not conform to the "normal change laws" can be perceived and predicted, and the processing model can be trained without obtaining data on abnormal events, thereby enhancing the perception ability of abnormal events and improving the detection accuracy of abnormal events with a low probability of occurrence, thereby solving the technical problem of low recognition accuracy of occasional and highly harmful abnormal events.

[0096] In the above embodiment of the present application, the first multimedia data and the second multimedia data are input into the extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, including: inputting the first multimedia data and the second multimedia data into the spatial feature extraction model in the extraction network for feature extraction to obtain the first spatial features of the first multimedia data and the second spatial features of the second multimedia data; inputting the first spatial features into the first spatiotemporal feature extraction model in the extraction network for feature extraction to obtain the spatiotemporal features of the first multimedia data; inputting the second spatial features into the second spatiotemporal feature extraction model in the extraction network for feature extraction to obtain the second spatiotemporal features of the second multimedia data.

[0097] The above-mentioned spatial feature extraction model can be a shared spatial feature extraction model, which can extract spatial features for each frame image in the video clip. Optionally, the spatial feature extraction model can adopt the first three blocks of ResNet18, but is not limited to this, and other deep feature extraction models can also be used.

[0098] The above-mentioned first spatiotemporal feature extraction model and second spatiotemporal feature extraction model can be spatiotemporal feature extraction models that can mine the spatiotemporal evolution laws of multimedia data. Optionally, in order to enhance the generalization ability of the system, the first spatiotemporal feature extraction model and the second spatiotemporal feature extraction model have the same model structure but do not share parameters. Moreover, the first spatiotemporal feature extraction model and the second spatiotemporal feature extraction model can adopt a convolutional long short-term memory network (ConvLSTM), but are not limited to this, and other spatiotemporal feature mining models, such as 3D convolution, can also be adopted.

[0099] In the above embodiment of the present application, based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data, determining whether an abnormal event has occurred with the target object includes: obtaining the characteristic distance between the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data; comparing the characteristic distance with a first preset threshold; if the characteristic distance is greater than the first preset threshold, determining that an abnormal event has occurred with the target object; if the characteristic distance is less than or equal to the preset threshold, determining that no abnormal event has occurred with the target object.

[0100] The aforementioned characteristic distance can be used to represent an outlier between two multimedia data sets. A larger characteristic distance indicates that the second multimedia data is less consistent with the evolution pattern of the first multimedia data, and is more likely to have an abnormal event. The characteristic distance between two spatiotemporal features can be calculated using existing distance calculation methods, such as Euclidean distance, Manhattan distance, Chebyshev distance, angle cosine distance, etc., but this application does not impose any specific limitations on this.

[0101] The above-mentioned first preset threshold can be a characteristic distance threshold that can determine the occurrence of an abnormal event. For example, in the embodiment of the present application, the first preset threshold is 0.4 as an example for illustration, but it is not limited to this and can be set according to the actual application scenario.

[0102] In the above embodiment of the present application, the video data is processed to obtain the first multimedia data and the second multimedia data, including: sampling the video data to obtain multiple frames of multimedia data; dividing the multiple frames of multimedia data to obtain the first multimedia data and the second multimedia data.

[0103] The number of frames of the multimedia data can be set according to the actual application scenario. In the embodiment of the present application, 8 frames are used as an example for illustration, but the invention is not limited thereto. Dividing the multi-frame multimedia data can mean dividing the multi-frame multimedia data into two parts evenly, but the invention is not limited thereto.

[0104] In the above embodiment of the present application, the method also includes the following steps: obtaining multiple groups of training samples; processing each group of training samples to obtain a first sample and a second sample contained in each group of training samples, wherein the collection time of the first sample is earlier than the collection time of the second sample; and using the multiple groups of training samples to train the initial processing model to obtain a processing model.

[0105] In the above embodiment of the present application, the initial processing model is trained using multiple groups of training samples to obtain a processing model, including: inputting each group of training samples into the initial processing model, predicting the first spatiotemporal features of the second sample, and obtaining the second spatiotemporal features of the second sample; obtaining a loss value of the initial processing model based on the first spatiotemporal features of the second sample and the second spatiotemporal features of the second sample; comparing the loss value with a second preset threshold; if the loss value is greater than the second preset threshold, updating the network weights of the initial processing model; if the loss value is less than or equal to the second preset threshold, obtaining the processing model.

[0106] The above-mentioned second preset threshold can be a loss value threshold that meets the training accuracy requirements, and can be set according to the actual application scenario. This application does not make any specific limitations on this.

[0107] In the above embodiment of the present application, based on the first spatiotemporal features of the second sample and the second spatiotemporal features of the second sample, the loss value of the initial processing model is obtained, including: obtaining the characteristic distance between the first spatiotemporal features of the second sample and the second spatiotemporal features of the second sample, and obtaining the loss value of the initial processing model.

[0108] In this embodiment, the loss function used in the processing model mainly includes two parts. The first part is the distance loss of the spatiotemporal feature vector, for example, the mean square error loss (Mean Square Error) can be used; the second part is the scene classification loss for the predicted spatiotemporal features.

[0109] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0110] Example 3

[0111] According to an embodiment of the present application, a device for detecting abnormal events for implementing the above abnormal event detection method is also provided, such as Figure 5As shown, the apparatus 500 includes: an acquisition module 502 , a processing module 504 and a determination module 506 .

[0112] Among them, the acquisition module 502 is used to acquire the first multimedia data and the second multimedia data of the target object, wherein the first multimedia data and the second multimedia data are multimedia data acquired from the video data of the target object, and the acquisition time of the first multimedia data is earlier than the acquisition time of the second multimedia data; the processing module 504 is used to process the first multimedia data using the processing model, predict the first spatiotemporal features of the second multimedia data, and process the second multimedia data using the processing model to obtain the second spatiotemporal features of the second multimedia data, wherein the processing model is used to acquire the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into the extraction network, obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, input the spatiotemporal features of the first multimedia data into the prediction network, and obtain the first spatiotemporal features of the second multimedia data; the determination module 506 is used to determine whether an abnormal event occurs to the target object based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data.

[0113] It should be noted that the acquisition module 502, processing module 504, and determination module 506 correspond to steps S202 to S206 in Example 1. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.

[0114] In the above embodiment of the present application, the processing module includes: a first extraction unit, a second extraction unit and a third extraction unit.

[0115] Among them, the first extraction unit is used to input the first multimedia data and the second multimedia data into the spatial feature extraction model in the extraction network for feature extraction, so as to obtain the first spatial feature of the first multimedia data and the second spatial feature of the second multimedia data; the second extraction unit is used to input the first spatial feature into the first spatiotemporal feature extraction model in the extraction network for feature extraction, so as to obtain the spatiotemporal feature of the first multimedia data; the third extraction unit is used to input the second spatial feature into the second spatiotemporal feature extraction model in the extraction network for feature extraction, so as to obtain the second spatiotemporal feature of the second multimedia data.

[0116] In the above embodiment of the present application, the determination module includes: an acquisition unit, a first comparison unit, a first determination unit, and a second determination unit.

[0117] Among them, the acquisition unit is used to obtain the first spatiotemporal feature of the second multimedia data and the characteristic distance of the second spatiotemporal feature of the second multimedia data; the first comparison unit is used to compare the characteristic distance with a first preset threshold; the first determination unit is used to determine that an abnormal event has occurred in the target object when the characteristic distance is greater than the first preset threshold; and the second determination unit is used to determine that no abnormal event has occurred in the target object when the characteristic distance is less than or equal to the preset threshold.

[0118] In the above embodiment of the present application, the acquisition module includes: a sampling unit and a division unit.

[0119] The sampling unit is used to sample the video data to obtain multiple frames of multimedia data; the dividing unit is used to divide the multiple frames of multimedia data to obtain first multimedia data and second multimedia data.

[0120] In the above embodiment of the present application, the device further includes: a training module.

[0121] Among them, the acquisition module is also used to obtain multiple groups of training samples, wherein each group of training samples includes: a first sample and a second sample, wherein the first sample and the second sample are samples obtained from the collected video data, and the collection time of the first sample is earlier than the collection time of the second sample; the training module is used to use multiple groups of training samples to train the initial processing model to obtain a processing model.

[0122] In the above embodiment of the present application, the training module includes: an input unit, a processing unit, a second comparison unit, an updating unit and a third determination unit.

[0123] Among them, the input unit inputs each group of training samples into the initial processing model, predicts the first spatiotemporal features of the second sample, and obtains the second spatiotemporal features of the second sample; the processing unit is used to obtain the loss value of the initial processing model based on the first spatiotemporal features of the second sample and the second spatiotemporal features of the second sample; the second comparison unit is used to compare the loss value with the second preset threshold; the update unit is used to update the network weights of the initial processing model if the loss value is greater than the second preset threshold; the third determination unit is used to obtain the processing model if the loss value is less than or equal to the second preset threshold.

[0124] In the above embodiments of the present application, the processing unit includes: a processing sub-unit.

[0125] The processing subunit is used to obtain the feature distance between the first spatiotemporal feature of the second sample and the second spatiotemporal feature of the second sample, and obtain the loss value of the initial processing model.

[0126] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0127] Example 4

[0128] According to an embodiment of the present application, a device for detecting abnormal events for implementing the above abnormal event detection method is also provided, such as Figure 6 As shown, the apparatus 600 includes: an acquisition module 602 , a first processing module 604 , a second processing module 606 and a determination module 608 .

[0129] Among them, the acquisition module 602 is used to acquire video data of the target object; the first processing module 604 is used to process the video data to obtain first multimedia data and second multimedia data, wherein the acquisition time of the first multimedia data is earlier than the acquisition time of the second multimedia data; the second processing module 606 is used to process the first multimedia data using the processing model, predict the first spatiotemporal features of the second multimedia data, and process the second multimedia data using the processing model to obtain the second spatiotemporal features of the second multimedia data, wherein the processing model is used to acquire the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into the extraction network, obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, input the spatiotemporal features of the first multimedia data into the prediction network, and obtain the first spatiotemporal features of the second multimedia data; the determination module 608 is used to determine whether an abnormal event occurs to the target object based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data.

[0130] It should be noted that the acquisition module 602, the first processing module 604, the second processing module 606, and the determination module 608 correspond to steps S402 to S408 in Example 2. The examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.

[0131] In the above embodiment of the present application, the second processing module includes: a first extraction unit, a second extraction unit and a third extraction unit.

[0132] Among them, the first extraction subunit is used to input the first multimedia data and the second multimedia data into the spatial feature extraction model in the extraction network for feature extraction, so as to obtain the first spatial feature of the first multimedia data and the second spatial feature of the second multimedia data; the second extraction subunit is used to input the first spatial feature into the first spatiotemporal feature extraction model in the extraction network for feature extraction, so as to obtain the spatiotemporal feature of the first multimedia data; the third extraction subunit is used to input the second spatial feature into the second spatiotemporal feature extraction model in the extraction network for feature extraction, so as to obtain the second spatiotemporal feature of the second multimedia data.

[0133] In the above embodiment of the present application, the determination module includes: an acquisition unit, a first comparison unit, a first determination unit, and a second determination unit.

[0134] Among them, the acquisition unit is used to obtain the first spatiotemporal feature of the second multimedia data and the characteristic distance of the second spatiotemporal feature of the second multimedia data; the first comparison unit is used to compare the characteristic distance with a first preset threshold; the first determination unit is used to determine that an abnormal event has occurred in the target object when the characteristic distance is greater than the first preset threshold; and the second determination unit is used to determine that no abnormal event has occurred in the target object when the characteristic distance is less than or equal to the preset threshold.

[0135] In the above embodiment of the present application, the second processing module includes: a sampling unit and a dividing unit.

[0136] The sampling unit is used to sample the video data to obtain multiple frames of multimedia data; the dividing unit is used to divide the multiple frames of multimedia data to obtain first multimedia data and second multimedia data.

[0137] In the above embodiment of the present application, the device further includes: a training module.

[0138] Among them, the acquisition module is also used to obtain multiple groups of training samples; the processing module is also used to process each group of training samples to obtain a first sample and a second sample contained in each group of training samples, wherein the collection time of the first sample is earlier than the collection time of the second sample; the training module is used to train the initial processing model using multiple groups of training samples to obtain a processing model.

[0139] In the above embodiment of the present application, the training module includes: an input unit, a processing unit, a second comparison unit, an updating unit and a third determination unit.

[0140] Among them, the input unit inputs each group of training samples into the initial processing model, predicts the first spatiotemporal features of the second sample, and obtains the second spatiotemporal features of the second sample; the processing unit is used to obtain the loss value of the initial processing model based on the first spatiotemporal features of the second sample and the second spatiotemporal features of the second sample; the second comparison unit is used to compare the loss value with the second preset threshold; the update unit is used to update the network weights of the initial processing model if the loss value is greater than the second preset threshold; the third determination unit is used to obtain the processing model if the loss value is less than or equal to the second preset threshold.

[0141] In the above embodiments of the present application, the processing unit includes: a processing sub-unit.

[0142] The processing subunit is used to obtain the feature distance between the first spatiotemporal feature of the second sample and the second spatiotemporal feature of the second sample, and obtain the loss value of the initial processing model.

[0143] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0144] Example 5

[0145] According to an embodiment of the present application, a system for detecting abnormal events is also provided, including:

[0146] processor; and

[0147] A memory is connected to a processor and is used to provide the processor with instructions for processing the following processing steps: obtaining first multimedia data and second multimedia data of a target object, wherein the first multimedia data and the second multimedia data are multimedia data obtained from video data of the target object, and the acquisition time of the first multimedia data is earlier than the acquisition time of the second multimedia data; using a processing model to process the first multimedia data to predict the first spatiotemporal features of the second multimedia data, and using the processing model to process the second multimedia data to obtain the second spatiotemporal features of the second multimedia data; based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data, determining whether an abnormal event occurs in the target object, wherein the processing model is used to obtain the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, and input the spatiotemporal features of the first multimedia data into a prediction network to obtain the first spatiotemporal features of the second multimedia data.

[0148] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0149] Example 6

[0150] According to an embodiment of the present application, a data processing method is also provided.

[0151] In the operating environment of the above embodiment 1, the present application provides the following Figure 7 The data processing method shown. Figure 7 This is a flow chart of a data processing method according to an embodiment of the present application. Figure 7 As shown, the method includes the following steps:

[0152] Step S702, obtaining video data of the target object;

[0153] The target objects in the above steps can refer to objects requiring abnormal event detection, including but not limited to open flame identification, sudden fires, and explosions in the field of safety supervision, or objects and areas requiring anti-theft monitoring in the field of security. The target objects in the above steps can also refer to users requiring behavior monitoring, such as patients requiring physical status monitoring in the medical field, but are not limited to these.

[0154] The video data in the above steps can be multiple images, video clips, etc. taken of the target object. In order to detect abnormal events more accurately, in the embodiment of the present application, the video clips of the target object are used as an example for illustration.

[0155] Step S704: Process the video data to obtain first multimedia data of the first time period and second multimedia data of the second time period;

[0156] The first time period in the above steps is before the second time period.

[0157] Step S706: acquiring a first spatiotemporal feature based on the first multimedia data, wherein the first spatiotemporal feature is a prediction feature of the second time period;

[0158] The first spatiotemporal feature in the above step may be the spatiotemporal feature of the second multimedia data predicted by mining the spatiotemporal evolution law of the first multimedia data.

[0159] Step S708: acquiring a second spatiotemporal feature based on the second multimedia data, wherein the second spatiotemporal feature is a detection feature of a second time period;

[0160] The second spatiotemporal features in the above steps may be spatiotemporal features mined directly from the second multimedia data.

[0161] Step S710: Based on the first spatiotemporal feature and the second spatiotemporal feature, determine whether the target object is in an abnormal state.

[0162] The abnormal state in the above steps may refer to a state in which a low-probability abnormal event occurs in the target object, including but not limited to open flame identification, sudden fire, explosion, etc. in the field of safety supervision, theft incidents in the field of security, etc., and may also refer to an abnormal behavioral state of the target object.

[0163] Through the solution provided by the above-mentioned embodiment of the present application, after obtaining the video data of the target object, the video data can be processed to obtain first multimedia data of the first time period and second multimedia data of the second time period. Then, based on the first multimedia data, the prediction features of the second time period are obtained, and based on the second multimedia data, the detection features of the second time period are obtained. Finally, through the first spatiotemporal features and the second spatiotemporal features, it can be determined whether the target object is in an abnormal state. It is easy to notice that by mining the spatiotemporal evolution law of the normal state, it is possible to perceive and predict abnormal states that do not conform to the "normal change law", thereby enhancing the perception ability of abnormal states and improving the detection accuracy of abnormal states with a low probability of occurrence, thereby solving the technical problem of low recognition accuracy of occasional and highly harmful abnormal events.

[0164] In the above embodiment of the present application, based on the first multimedia data, obtaining the first spatiotemporal feature includes: inputting the first multimedia data into a shared spatial feature extraction model for feature extraction to obtain a first spatial feature, wherein the first spatial feature is a detection feature of the first time period; inputting the first spatial feature into the first spatiotemporal feature extraction model for feature extraction to obtain a third spatiotemporal feature, wherein the third spatiotemporal feature is a detection feature of the first time period; inputting the third spatiotemporal feature into the prediction network to obtain the first spatiotemporal feature.

[0165] The above-mentioned spatial feature extraction model can be a shared spatial feature extraction model, which can extract spatial features for each frame image in the video clip. Optionally, the spatial feature extraction model can adopt the first three blocks of ResNet18, but is not limited to this, and other deep feature extraction models can also be used.

[0166] The first spatiotemporal feature extraction model can be a spatiotemporal feature extraction model that can mine the spatiotemporal evolution patterns of multimedia data. The first spatiotemporal feature extraction model can use a convolutional long short-term memory network (ConvLSTM), but is not limited thereto. Other spatiotemporal feature mining models, such as 3D convolution, can also be used.

[0167] The aforementioned prediction network can be a model for predicting the spatiotemporal features of future multimedia data. Optionally, the prediction model can employ two sets of convolutional layers. The first set of convolutional layers downsamples the number of feature channels to compress the feature space and focus on the essential features' evolution patterns. The second set of convolutional layers upsamples the number of feature channels to restore the original dimensions, facilitating subsequent feature comparison.

[0168] In the above embodiment of the present application, based on the second multimedia data, obtaining the second spatiotemporal feature includes: inputting the second multimedia data into a shared spatial feature extraction model for feature extraction to obtain a second spatial feature, wherein the second spatial feature is a detection feature of the second time period; inputting the second spatial feature into a second spatiotemporal feature extraction model for feature extraction to obtain a second spatiotemporal feature.

[0169] The second spatiotemporal feature extraction model can be a spatiotemporal feature extraction model that can mine the spatiotemporal evolution of multimedia data. Optionally, in order to enhance the generalization capability of the system, the first spatiotemporal feature extraction model and the second spatiotemporal feature extraction model have the same model structure but do not share parameters. The second spatiotemporal feature extraction model can use a convolutional long short-term memory network (ConvLSTM), but is not limited thereto. Other spatiotemporal feature mining models, such as 3D convolution, can also be used.

[0170] In the above embodiment of the present application, based on the first spatiotemporal feature and the second spatiotemporal feature, determining whether the target object is in an abnormal state includes: obtaining the characteristic distance between the first spatiotemporal feature and the second spatiotemporal feature; comparing the characteristic distance with a first preset threshold; when the characteristic distance is greater than the first preset threshold, determining that the target object is in an abnormal state; when the characteristic distance is less than or equal to the preset threshold, determining that the target object is not in an abnormal state.

[0171] In the above embodiments of the present application, video data is processed to obtain first multimedia data of a first time period and second multimedia data of a second time period, including: sampling the video data to obtain multiple frames of multimedia data; dividing the multiple frames of multimedia data to obtain first multimedia data and second multimedia data.

[0172] In the above embodiment of the present application, the method also includes the following steps: obtaining multiple groups of training samples; processing each group of training samples to obtain the first sample and the second sample contained in each group of training samples; using multiple groups of training samples to train the spatial feature extraction model, the first spatiotemporal feature extraction model, the second spatiotemporal feature extraction model and the prediction network.

[0173] In the above embodiment of the present application, multiple groups of training samples are used to train the spatial feature extraction model, the first spatiotemporal feature extraction model, the second spatiotemporal feature extraction model and the prediction network, including: inputting each group of training samples into the spatial feature extraction model to obtain the first spatial feature and the second spatial feature, wherein the first spatial feature is the detection feature of the first sample, and the second spatial feature is the detection feature of the second sample; inputting the first spatial feature into the first spatiotemporal feature extraction model to obtain the third spatiotemporal feature, wherein the third spatiotemporal feature is the detection feature of the first sample; inputting the third spatiotemporal feature into the prediction network to obtain the first spatiotemporal feature, wherein the first spatiotemporal feature is the detection feature of the first sample. The feature is a prediction feature of the second sample; the second spatial feature is input into the second spatiotemporal feature extraction model to obtain the second spatiotemporal feature, wherein the second spatiotemporal feature is the detection feature of the second sample; based on the first spatiotemporal feature and the second spatiotemporal feature, a loss value is obtained; the loss value is compared with a second preset threshold; if the loss value is greater than the second preset threshold, the network weights of the spatial feature extraction model, the first spatiotemporal feature extraction model, the second spatiotemporal feature extraction model and the prediction network are updated; if the loss value is less than or equal to the second preset threshold, it is determined that the training of the spatial feature extraction model, the first spatiotemporal feature extraction model, the second spatiotemporal feature extraction model and the prediction network is completed.

[0174] In the above embodiment of the present application, the loss value is obtained based on the first spatiotemporal feature of the second sample and the second spatiotemporal feature of the second sample, including: obtaining the characteristic distance between the first spatiotemporal feature of the second sample and the second spatiotemporal feature of the second sample to obtain the loss value.

[0175] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0176] Example 7

[0177] The embodiment of the present application can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.

[0178] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.

[0179] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the abnormal event detection method: obtaining first multimedia data and second multimedia data of the target object, wherein the first multimedia data and the second multimedia data are multimedia data obtained from the video data of the target object, and the acquisition time of the first multimedia data is earlier than the acquisition time of the second multimedia data; using the processing model to process the first multimedia data to predict the first spatiotemporal features of the second multimedia data, and using the processing model to process the second multimedia data to obtain the second spatiotemporal features of the second multimedia data, wherein the processing model is used to obtain the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into the extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, and input the spatiotemporal features of the first multimedia data into the prediction network to obtain the first spatiotemporal features of the second multimedia data; based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data, determine whether an abnormal event occurs to the target object.

[0180] Optionally, Figure 8 This is a structural block diagram of a computer terminal according to an embodiment of the present application. Figure 8 As shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 802 and a memory 804.

[0181] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the detection method and device of abnormal events in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned detection method of abnormal events. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0182] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtaining the first multimedia data and the second multimedia data of the target object, wherein the first multimedia data and the second multimedia data are multimedia data obtained from the video data of the target object, and the acquisition time of the first multimedia data is earlier than the acquisition time of the second multimedia data; using the processing model to process the first multimedia data to predict the first spatiotemporal features of the second multimedia data, and using the processing model to process the second multimedia data to obtain the second spatiotemporal features of the second multimedia data, wherein the processing model is used to obtain the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into the extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, and input the spatiotemporal features of the first multimedia data into the prediction network to obtain the first spatiotemporal features of the second multimedia data; based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data, determine whether an abnormal event occurs to the target object.

[0183] Optionally, the processor may also execute the program code of the following steps: inputting the first multimedia data and the second multimedia data into the spatial feature extraction model in the extraction network for feature extraction to obtain the first spatial features of the first multimedia data and the second spatial features of the second multimedia data; inputting the first spatial features into the first spatiotemporal feature extraction model in the extraction network for feature extraction to obtain the spatiotemporal features of the first multimedia data; inputting the second spatial features into the second spatiotemporal feature extraction model in the extraction network for feature extraction to obtain the second spatiotemporal features of the second multimedia data.

[0184] Optionally, the processor may also execute the program code of the following steps: obtaining the characteristic distance between the first spatiotemporal feature of the second multimedia data and the second spatiotemporal feature of the second multimedia data; comparing the characteristic distance with a first preset threshold; determining that an abnormal event has occurred with the target object when the characteristic distance is greater than the first preset threshold; and determining that no abnormal event has occurred with the target object when the characteristic distance is less than or equal to the preset threshold.

[0185] Optionally, the processor may further execute program codes of the following steps: sampling the video data to obtain multiple frames of multimedia data; and dividing the multiple frames of multimedia data to obtain first multimedia data and second multimedia data.

[0186] Optionally, the processor may also execute the program code of the following steps: obtaining multiple groups of training samples, wherein each group of training samples includes: a first sample and a second sample, wherein the first sample and the second sample are multimedia data obtained from the video data of the target object, and the acquisition time of the first sample is earlier than the acquisition time of the second sample; and using the multiple groups of training samples to train the initial processing model to obtain a processing model.

[0187] Optionally, the processor may also execute the program code of the following steps: input each group of training samples into the initial processing model, predict the first spatiotemporal features of the second sample, and obtain the second spatiotemporal features of the second sample; obtain the loss value of the initial processing model based on the first spatiotemporal features of the second sample and the second spatiotemporal features of the second sample; compare the loss value with a second preset threshold; if the loss value is greater than the second preset threshold, update the network weights of the initial processing model; if the loss value is less than or equal to the second preset threshold, obtain the processing model.

[0188] Optionally, the processor may further execute program code of the following steps: obtaining a feature distance between the first spatiotemporal feature of the second sample and the second spatiotemporal feature of the second sample, and obtaining a loss value of the initial processing model.

[0189] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain video data of the target object; process the video data to obtain first multimedia data and second multimedia data, wherein the acquisition time of the first multimedia data is earlier than the acquisition time of the second multimedia data; use the processing model to process the first multimedia data to predict the first spatiotemporal features of the second multimedia data, and use the processing model to process the second multimedia data to obtain the second spatiotemporal features of the second multimedia data, wherein the processing model is used to obtain the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into the extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, input the spatiotemporal features of the first multimedia data into the prediction network to obtain the first spatiotemporal features of the second multimedia data; based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data, determine whether an abnormal event occurs to the target object.

[0190] The embodiments of the present application provide a solution for detecting abnormal events. By exploring the spatiotemporal evolution patterns of normal states, the system can detect and predict abnormal events that do not conform to these patterns. Furthermore, the system can train a processing model without acquiring data on abnormal events. This enhances the ability to detect abnormal events and improves the accuracy of detecting low-probability abnormal events. This solves the technical problem of low recognition accuracy for sporadic but potentially harmful abnormal events.

[0191] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain video data of the target object; process the video data to obtain first multimedia data of a first time period and second multimedia data of a second time period; based on the first multimedia data, obtain a first spatiotemporal feature, wherein the first spatiotemporal feature is a prediction feature of the second time period; based on the second multimedia data, obtain a second spatiotemporal feature, wherein the second spatiotemporal feature is a detection feature of the second time period; based on the first spatiotemporal feature and the second spatiotemporal feature, determine whether the target object is in an abnormal state.

[0192] Optionally, the processor may also execute the program code of the following steps: inputting the first multimedia data into a shared spatial feature extraction model for feature extraction to obtain a first spatial feature, wherein the first spatial feature is a detection feature of a first time period; inputting the first spatial feature into a first spatiotemporal feature extraction model for feature extraction to obtain a third spatiotemporal feature, wherein the third spatiotemporal feature is a detection feature of the first time period; inputting the third spatiotemporal feature into a prediction network to obtain the first spatiotemporal feature.

[0193] Optionally, the above-mentioned processor can also execute the program code of the following steps: input the second multimedia data into the shared spatial feature extraction model for feature extraction to obtain a second spatial feature, wherein the second spatial feature is the detection feature of the second time period; input the second spatial feature into the second spatiotemporal feature extraction model for feature extraction to obtain a second spatiotemporal feature.

[0194] It can be understood by those skilled in the art that Figure 8 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 8 It does not limit the structure of the above electronic device. For example, the computer terminal A may also include Figure 8 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 8 Different configurations shown.

[0195] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0196] Example 8

[0197] The embodiment of the present application further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the abnormal event detection method provided in the above embodiment.

[0198] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0199] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining first multimedia data and second multimedia data of the target object, wherein the first multimedia data and the second multimedia data are multimedia data obtained from the video data of the target object, and the acquisition time of the first multimedia data is earlier than the acquisition time of the second multimedia data; using the processing model to process the first multimedia data to predict the first spatiotemporal features of the second multimedia data, and using the processing model to process the second multimedia data to obtain the second spatiotemporal features of the second multimedia data, wherein the processing model is used to obtain the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into the extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, and input the spatiotemporal features of the first multimedia data into the prediction network to obtain the first spatiotemporal features of the second multimedia data; based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data, determine whether an abnormal event occurs to the target object.

[0200] Optionally, the above-mentioned storage medium is also configured to store program codes for executing the following steps: inputting the first multimedia data and the second multimedia data into the spatial feature extraction model in the extraction network for feature extraction to obtain the first spatial features of the first multimedia data and the second spatial features of the second multimedia data; inputting the first spatial features into the first spatiotemporal feature extraction model in the extraction network for feature extraction to obtain the spatiotemporal features of the first multimedia data; inputting the second spatial features into the second spatiotemporal feature extraction model in the extraction network for feature extraction to obtain the second spatiotemporal features of the second multimedia data.

[0201] Optionally, the storage medium is further configured to store program code for executing the following steps: obtaining the first spatiotemporal feature of the second multimedia data and the characteristic distance of the second spatiotemporal feature of the second multimedia data; comparing the characteristic distance with a first preset threshold; determining that an abnormal event has occurred with the target object when the characteristic distance is greater than the first preset threshold; and determining that no abnormal event has occurred with the target object when the characteristic distance is less than or equal to the preset threshold.

[0202] Optionally, the storage medium is further configured to store program codes for executing the following steps: acquiring video data of a target object; sampling the video data to obtain multiple frames of multimedia data; and dividing the multiple frames of multimedia data to obtain first multimedia data and second multimedia data.

[0203] Optionally, the storage medium is further configured to store program code for executing the following steps: obtaining multiple groups of training samples, wherein each group of training samples includes: a first sample and a second sample, wherein the first sample and the second sample are multimedia data obtained from the video data of the target object, and the acquisition time of the first sample is earlier than the acquisition time of the second sample; using the multiple groups of training samples to train the initial processing model to obtain a processing model.

[0204] Optionally, the storage medium is also configured to store program code for executing the following steps: inputting each group of training samples into the initial processing model, predicting the first spatiotemporal features of the second sample, and obtaining the second spatiotemporal features of the second sample; obtaining the loss value of the initial processing model based on the first spatiotemporal features of the second sample and the second spatiotemporal features of the second sample; comparing the loss value with a second preset threshold; if the loss value is greater than the second preset threshold, updating the network weights of the initial processing model; if the loss value is less than or equal to the second preset threshold, obtaining the processing model.

[0205] Optionally, the storage medium is further configured to store program codes for executing the following steps: obtaining a feature distance between the first spatiotemporal feature of the second sample and the second spatiotemporal feature of the second sample, and obtaining a loss value of the initial processing model.

[0206] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: acquiring video data of the target object; processing the video data to obtain first multimedia data and second multimedia data, wherein the acquisition time of the first multimedia data is earlier than the acquisition time of the second multimedia data; processing the first multimedia data using a processing model to predict the first spatiotemporal features of the second multimedia data, and processing the second multimedia data using the processing model to obtain second spatiotemporal features of the second multimedia data, wherein the processing model is used to acquire the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, and input the spatiotemporal features of the first multimedia data into a prediction network to obtain the first spatiotemporal features of the second multimedia data; and determine whether an abnormal event occurs to the target object based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data.

[0207] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining video data of the target object; processing the video data to obtain first multimedia data of a first time period and second multimedia data of a second time period; obtaining a first spatiotemporal feature based on the first multimedia data, wherein the first spatiotemporal feature is a prediction feature of the second time period; obtaining a second spatiotemporal feature based on the second multimedia data, wherein the second spatiotemporal feature is a detection feature of the second time period; and judging whether the target object is in an abnormal state based on the first spatiotemporal feature and the second spatiotemporal feature.

[0208] Optionally, the storage medium is also configured to store program code for executing the following steps: inputting the first multimedia data into a shared spatial feature extraction model for feature extraction to obtain a first spatial feature, wherein the first spatial feature is a detection feature of a first time period; inputting the first spatial feature into a first spatiotemporal feature extraction model for feature extraction to obtain a third spatiotemporal feature, wherein the third spatiotemporal feature is a detection feature of the first time period; inputting the third spatiotemporal feature into a prediction network to obtain the first spatiotemporal feature.

[0209] Optionally, the above-mentioned storage medium is also configured to store program code for executing the following steps: inputting the second multimedia data into a shared spatial feature extraction model for feature extraction to obtain a second spatial feature, wherein the second spatial feature is a detection feature of a second time period; inputting the second spatial feature into a second spatiotemporal feature extraction model for feature extraction to obtain a second spatiotemporal feature.

[0210] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0211] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0212] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0213] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0214] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0215] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0216] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A data processing method, comprising: Obtain video data of the target object; Processing the video data to obtain first multimedia data of a first time period and second multimedia data of a second time period; Obtaining, using a processing model based on the first multimedia data, a first spatiotemporal feature, wherein the first spatiotemporal feature is a predicted feature for a second time period, the first spatiotemporal feature being used to represent a spatiotemporal feature of the second multimedia data predicted by mining a spatiotemporal evolution law of the first multimedia data, and a loss function employed by the processing model comprising: a distance loss of a spatiotemporal feature vector and a scene classification loss for the predicted spatiotemporal feature; Obtaining a second spatiotemporal feature based on the second multimedia data using the processing model, wherein the second spatiotemporal feature is a detection feature of a second time period, and the second spatiotemporal feature is used to represent a spatiotemporal feature mined directly from the second multimedia data; Based on the first spatiotemporal feature and the second spatiotemporal feature, determining whether the target object is in an abnormal state; The step of obtaining the first spatiotemporal feature based on the first multimedia data by using the processing model includes: Based on the first multimedia data, the shared spatial feature extraction model and the first spatiotemporal feature extraction model in the processing model, a third spatiotemporal feature is obtained, wherein the third spatiotemporal feature is a detection feature of the first time period; the third spatiotemporal feature is input into the prediction network in the processing model to obtain the first spatiotemporal feature; wherein, the second spatiotemporal feature is obtained based on the second multimedia data using the processing model, including: based on the second multimedia data, the shared spatial feature extraction model and the second spatiotemporal feature extraction model in the processing model, the second spatiotemporal feature is obtained, wherein the second spatiotemporal feature extraction model and the first spatiotemporal feature extraction model do not share parameters.

2. The method according to claim 1, wherein Obtaining a third spatiotemporal feature based on the first multimedia data, the shared spatial feature extraction model in the processing model, and the first spatiotemporal feature extraction model, including: Inputting the first multimedia data into the shared spatial feature extraction model to perform feature extraction to obtain a first spatial feature, wherein the first spatial feature is a detection feature of a first time period; The first spatial feature is input into the first spatiotemporal feature extraction model for feature extraction to obtain the third spatiotemporal feature.

3. The method according to claim 1, wherein Obtaining the second spatiotemporal feature based on the second multimedia data, the shared spatial feature extraction model, and the second spatiotemporal feature extraction model in the processing model includes: Inputting the second multimedia data into the shared spatial feature extraction model to perform feature extraction to obtain a second spatial feature, wherein the second spatial feature is a detection feature of a second time period; The second spatial feature is input into the second spatiotemporal feature extraction model for feature extraction to obtain the second spatiotemporal feature.

4. A method for detecting an abnormal event, comprising: Acquire first multimedia data and second multimedia data of a target object, wherein the first multimedia data and the second multimedia data are multimedia data acquired from video data of the target object, and acquisition time of the first multimedia data is earlier than acquisition time of the second multimedia data; The first multimedia data is processed using a processing model to predict the first spatiotemporal features of the second multimedia data, and the second multimedia data is processed using the processing model to obtain the second spatiotemporal features of the second multimedia data, wherein the processing model is used to obtain the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, input the spatiotemporal features of the first multimedia data into a prediction network to obtain the first spatiotemporal features of the second multimedia data, the first spatiotemporal features are used to represent the spatiotemporal features of the second multimedia data predicted by mining the spatiotemporal evolution laws of the first multimedia data, and the second spatiotemporal features are used to represent the spatiotemporal features mined directly from the second multimedia data, and the loss function used by the processing model includes: distance loss of spatiotemporal feature vectors and scene classification loss for the predicted spatiotemporal features; determining whether an abnormal event occurs to the target object based on the first spatiotemporal feature of the second multimedia data and the second spatiotemporal feature of the second multimedia data; The step of inputting the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data includes: Based on the first multimedia data, the spatial feature extraction model shared in the extraction network, and the first spatiotemporal feature extraction model, the spatiotemporal features of the first multimedia data are obtained, wherein the spatiotemporal features of the first multimedia data are detection features of a first time period; and based on the second multimedia data, the shared spatial feature extraction model, and the second spatiotemporal feature extraction model in the extraction network, the second spatiotemporal features are obtained, wherein the second spatiotemporal feature extraction model and the first spatiotemporal feature extraction model do not share parameters.

5. The method according to claim 4, wherein The method includes: obtaining spatiotemporal features of the first multimedia data based on the first multimedia data, the shared spatial feature extraction model in the extraction network, and the first spatiotemporal feature extraction model; and obtaining the second spatiotemporal features based on the second multimedia data, the shared spatial feature extraction model, and the second spatiotemporal feature extraction model in the extraction network, including: Inputting the first multimedia data and the second multimedia data into the shared spatial feature extraction model in the extraction network to perform feature extraction, thereby obtaining first spatial features of the first multimedia data and second spatial features of the second multimedia data; Inputting the first spatial feature into the first spatiotemporal feature extraction model in the extraction network to perform feature extraction to obtain the spatiotemporal feature of the first multimedia data; The second spatial feature is input into the second spatiotemporal feature extraction model in the extraction network to perform feature extraction to obtain the second spatiotemporal feature of the second multimedia data.

6. The method according to claim 4, wherein: Determining whether an abnormal event occurs on the target object based on the first spatiotemporal feature of the second multimedia data and the second spatiotemporal feature of the second multimedia data includes: Obtaining a feature distance between a first spatiotemporal feature of the second multimedia data and a second spatiotemporal feature of the second multimedia data; Comparing the characteristic distance with a first preset threshold; When the characteristic distance is greater than the first preset threshold, determining that an abnormal event occurs on the target object; When the characteristic distance is less than or equal to the preset threshold, it is determined that no abnormal event occurs to the target object.

7. The method according to claim 4, wherein: Acquiring first multimedia data and second multimedia data of a target object includes: Sampling the video data to obtain multiple frames of multimedia data; The multiple frames of multimedia data are divided to obtain the first multimedia data and the second multimedia data.

8. The method according to any one of claims 4 to 7, wherein: The method further comprises: Acquire multiple groups of training samples, wherein each group of training samples includes: a first sample and a second sample, the first sample and the second sample are samples acquired from acquired video data, and acquisition time of the first sample is earlier than acquisition time of the second sample; The initial processing model is trained using the multiple groups of training samples to obtain the processing model.

9. The method according to claim 8, wherein Training the initial processing model using the multiple sets of training samples to obtain the processing model includes: Inputting each set of training samples into the initial processing model, predicting the first spatiotemporal features of the second samples, and obtaining the second spatiotemporal features of the second samples; Obtaining a loss value of the initial processing model based on the first spatiotemporal feature of the second sample and the second spatiotemporal feature of the second sample; comparing the loss value with a second preset threshold; If the loss value is greater than the second preset threshold, updating the network weights of the initial processing model; If the loss value is less than or equal to the second preset threshold, the processing model is obtained.

10. The method according to claim 9, wherein: Obtaining a loss value of the initial processing model based on the first spatiotemporal feature of the second sample and the second spatiotemporal feature of the second sample, comprising: Obtain a feature distance between the first spatiotemporal feature of the second sample and the second spatiotemporal feature of the second sample to obtain a loss value of the initial processing model.

11. A method for detecting an abnormal event, comprising: Obtain video data of the target object; Processing the video data to obtain first multimedia data and second multimedia data, wherein the first multimedia data is collected earlier than the second multimedia data; The first multimedia data is processed using a processing model to predict the first spatiotemporal features of the second multimedia data, and the second multimedia data is processed using the processing model to obtain the second spatiotemporal features of the second multimedia data, wherein the processing model is used to obtain the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, input the spatiotemporal features of the first multimedia data into a prediction network to obtain the first spatiotemporal features of the second multimedia data, the first spatiotemporal features are used to represent the spatiotemporal features of the second multimedia data predicted by mining the spatiotemporal evolution laws of the first multimedia data, and the second spatiotemporal features are used to represent the spatiotemporal features mined directly from the second multimedia data, and the loss function used by the processing model includes: distance loss of spatiotemporal feature vectors and scene classification loss for the predicted spatiotemporal features; determining whether an abnormal event occurs to the target object based on the first spatiotemporal feature of the second multimedia data and the second spatiotemporal feature of the second multimedia data; The step of inputting the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data includes: Based on the first multimedia data, the spatial feature extraction model shared in the extraction network, and the first spatiotemporal feature extraction model, the spatiotemporal features of the first multimedia data are obtained, wherein the spatiotemporal features of the first multimedia data are detection features of a first time period; and based on the second multimedia data, the shared spatial feature extraction model, and the second spatiotemporal feature extraction model in the extraction network, the second spatiotemporal features are obtained, wherein the second spatiotemporal feature extraction model and the first spatiotemporal feature extraction model do not share parameters.

12. The method according to claim 11, wherein The method includes: obtaining spatiotemporal features of the first multimedia data based on the first multimedia data, the shared spatial feature extraction model in the extraction network, and the first spatiotemporal feature extraction model; and obtaining the second spatiotemporal features based on the second multimedia data, the shared spatial feature extraction model, and the second spatiotemporal feature extraction model in the extraction network, including: Inputting the first multimedia data and the second multimedia data into the shared spatial feature extraction model in the extraction network to perform feature extraction, thereby obtaining a first spatial feature of the first multimedia data and a second spatial feature of the second multimedia data; Inputting the first spatial feature into the first spatiotemporal feature extraction model in the extraction network to perform feature extraction to obtain the spatiotemporal feature of the first multimedia data; The second spatial feature is input into the second spatiotemporal feature extraction model in the extraction network to perform feature extraction to obtain the second spatiotemporal feature of the second multimedia data.

13. A device for detecting abnormal events, comprising: an acquisition module, configured to acquire first multimedia data and second multimedia data of a target object, wherein the first multimedia data and the second multimedia data are multimedia data acquired from video data of the target object, and acquisition time of the first multimedia data is earlier than acquisition time of the second multimedia data; a processing module, configured to process the first multimedia data using a processing model to predict first spatiotemporal features of the second multimedia data, and process the second multimedia data using the processing model to obtain second spatiotemporal features of the second multimedia data, wherein the processing model is configured to obtain the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, input the spatiotemporal features of the first multimedia data into a prediction network to obtain the first spatiotemporal features of the second multimedia data, the first spatiotemporal features being used to represent the spatiotemporal features of the second multimedia data predicted by mining the spatiotemporal evolution laws of the first multimedia data, and the second spatiotemporal features being used to represent the spatiotemporal features mined directly from the second multimedia data, and the loss function adopted by the processing model includes: a distance loss of a spatiotemporal feature vector and a scene classification loss for the predicted spatiotemporal features; a determination module, configured to determine whether an abnormal event occurs to the target object based on the first spatiotemporal feature of the second multimedia data and the second spatiotemporal feature of the second multimedia data; In which, the processing model is used to input the first multimedia data and the second multimedia data into the extraction network through the following steps to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data: based on the first multimedia data, the spatial feature extraction model shared in the extraction network and the first spatiotemporal feature extraction model, the spatiotemporal features of the first multimedia data are obtained, wherein the spatiotemporal features of the first multimedia data are detection features of the first time period; and based on the second multimedia data, the shared spatial feature extraction model and the second spatiotemporal feature extraction model in the extraction network, the second spatiotemporal features are obtained, wherein the second spatiotemporal feature extraction model and the first spatiotemporal feature extraction model do not share parameters.

14. A device for detecting abnormal events, comprising: An acquisition module, used to acquire video data of a target object; a first processing module, configured to process the video data to obtain first multimedia data and second multimedia data, wherein the first multimedia data is collected earlier than the second multimedia data; a second processing module, configured to process the first multimedia data using a processing model to predict first spatiotemporal features of the second multimedia data, and process the second multimedia data using the processing model to obtain second spatiotemporal features of the second multimedia data, wherein the processing model is configured to obtain the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, input the spatiotemporal features of the first multimedia data into a prediction network to obtain the first spatiotemporal features of the second multimedia data, the first spatiotemporal features being used to represent the spatiotemporal features of the second multimedia data predicted by mining the spatiotemporal evolution laws of the first multimedia data, and the second spatiotemporal features being used to represent the spatiotemporal features mined directly from the second multimedia data, and the loss function adopted by the processing model includes: a distance loss of a spatiotemporal feature vector and a scene classification loss for the predicted spatiotemporal features; a determination module, configured to determine whether an abnormal event occurs to the target object based on the first spatiotemporal feature of the second multimedia data and the second spatiotemporal feature of the second multimedia data; Among them, the second processing module is used to input the first multimedia data and the second multimedia data into the extraction network through the following steps to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data: based on the first multimedia data, the spatial feature extraction model shared in the extraction network and the first spatiotemporal feature extraction model, obtain the spatiotemporal features of the first multimedia data, wherein the spatiotemporal features of the first multimedia data are the detection features of the first time period; and based on the second multimedia data, the shared spatial feature extraction model and the second spatiotemporal feature extraction model in the extraction network, obtain the second spatiotemporal features, wherein the second spatiotemporal feature extraction model and the first spatiotemporal feature extraction model do not share parameters.

15. A storage medium comprising a stored program, wherein: When the program is running, the device where the storage medium is located is controlled to execute the data processing method described in any one of claims 1 to 3, or the abnormal event detection method described in any one of claims 4 to 9.

16. A computing device comprising: A processor and a memory, wherein the processor is used to run a program stored in the memory, wherein when the program is run, the data processing method described in any one of claims 1 to 3 or the abnormal event detection method described in any one of claims 4 to 9 is executed.

17. A system for detecting abnormal events, comprising: processor; as well as A memory is connected to the processor and is used to provide the processor with instructions for processing the following processing steps: obtaining first multimedia data and second multimedia data of the target object, wherein the first multimedia data and the second multimedia data are multimedia data obtained from the video data of the target object, and the acquisition time of the first multimedia data is earlier than the acquisition time of the second multimedia data; using a processing model to process the first multimedia data to predict the first spatiotemporal features of the second multimedia data, and using the processing model to process the second multimedia data to obtain the second spatiotemporal features of the second multimedia data, wherein the processing model is used to obtain the first multimedia data and the second multimedia data, input the first multimedia data and the second multimedia data into an extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, input the spatiotemporal features of the first multimedia data into a prediction network to obtain the first spatiotemporal features of the second multimedia data, and the first spatiotemporal features are used to represent the prediction obtained by mining the spatiotemporal evolution law of the first multimedia data. The spatiotemporal features of the second multimedia data are used to represent the spatiotemporal features mined directly from the second multimedia data; based on the first spatiotemporal features of the second multimedia data and the second spatiotemporal features of the second multimedia data, determine whether an abnormal event occurs to the target object, and the loss function adopted by the processing model includes: the distance loss of the spatiotemporal feature vector and the scene classification loss for the predicted spatiotemporal features; wherein the first multimedia data and the second multimedia data are input into the extraction network to obtain the spatiotemporal features of the first multimedia data and the second spatiotemporal features of the second multimedia data, including: based on the first multimedia data, the spatial feature extraction model shared in the extraction network and the first spatiotemporal feature extraction model, obtain the spatiotemporal features of the first multimedia data, wherein the spatiotemporal features of the first multimedia data are detection features of the first time period; and based on the second multimedia data, the shared spatial feature extraction model and the second spatiotemporal feature extraction model in the extraction network, obtain the second spatiotemporal features, wherein the second spatiotemporal feature extraction model and the first spatiotemporal feature extraction model do not share parameters.

Citation Information

Patent Citations

  • Method for abnormal region detection based on self-similarity number encoding

    CN103810467A

  • Log data exception detection method and device, terminal and medium

    CN110321371A