Event detection method and device, and event detection model training method and device

By correcting the category sequence of the detection model through an attention mechanism and a feature projection module, the problem of the inability to identify abnormal events with unclear action features in existing technologies is solved, enabling real-time detection and early warning of abnormal events in videos and improving the real-time analysis capability of video data.

CN114863306BActive Publication Date: 2025-10-24ALIBABA GROUP HOLDING LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110062376.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-18
Publication Date
2025-10-24
Estimated Expiration
2041-01-18

AI Technical Summary

Technical Problem

Existing detection models struggle to effectively identify abnormal events with subtle motion characteristics, making it impossible to detect and alert to anomalies in video scenarios in real time. Furthermore, current technologies can only be used for post-event evidence collection and cannot be used for real-time intervention.

Method used

An attention-based detection model training method is adopted. By receiving features from sample videos, a category sequence is generated, and the category sequence is corrected using the attention mechanism and feature projection module. The target loss function is trained to improve the detection capability of the model.

Benefits of technology

It enables comprehensive and accurate detection of abnormal events, both those with obvious and those with subtle motion characteristics. It can detect and warn of abnormal situations in real time, improving the efficiency of information mining from video data and enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114863306B_ABST
    Figure CN114863306B_ABST
Patent Text Reader

Abstract

Embodiments of the present specification provide a detection model training method and device, and an event detection method and device. The detection model training method comprises receiving a sample video containing an event, and extracting a sample feature of the sample video; inputting the sample feature into a first category module to generate a first category sequence of the sample video; correcting the first category sequence based on an attention mechanism method to obtain a corrected first category sequence; obtaining a target loss function based on the corrected first category sequence, and training the detection model according to the target loss function.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present specification relate to the technical field of computer technology, and particularly relate to a detection model training method. One or more embodiments of the present specification also relate to an event detection method, a detection model training apparatus, an event detection apparatus, a computing device, and a computer-readable storage medium. BACKGROUND

[0002] In current urban brain projects such as public security, municipal, and traffic, a large amount of camera video data is generated every day, and how to effectively mine information from the video data is a problem to be solved. The visual analysis based on a single picture cannot effectively utilize the effective information in time sequence, and it is difficult to obtain good results in the video scene. Therefore, the current common practice is to rely on post-supervision, that is, when someone reports an abnormal behavior event (such as theft, fighting, etc.) occurs, the corresponding period of video is retrieved to view the situation at that time. This post-supervision method can only be used for post-evidence, and cannot find abnormal situations in real time and make warnings and interventions, and the existing detection model can only identify abnormal events with obvious action features, and can easily miss abnormal events with non-obvious action features.

[0003] Therefore, there is an urgent need to provide a detection model training method that can detect abnormal events with non-obvious action features. SUMMARY

[0004] Therefore, the embodiments of the present specification provide a detection model training method. One or more embodiments of the present specification also relate to an event detection method, a detection model training apparatus, an event detection apparatus, a computing device, and a computer-readable storage medium to solve the technical defects in the prior art.

[0005] According to a first aspect of the embodiments of the present specification, a detection model training method is provided, comprising:

[0006] receiving a sample video containing an event, and extracting a sample feature of the sample video;

[0007] inputting the sample feature into a first category module to generate a first category sequence of the sample video;

[0008] correcting the first category sequence based on an attention mechanism method to obtain a corrected first category sequence;

[0009] obtaining a target loss function based on the corrected first category sequence, and training the detection model according to the target loss function.

[0010] According to a second aspect of the embodiments of the present specification, a detection model training method is provided, comprising:

[0011] presenting a video input interface to the user based on a calling request of the user;

[0012] receiving a sample video containing an event sent by the user based on the video input interface, and extracting sample features of the sample video;

[0013] inputting the sample features into a first category module to generate a first category sequence of the sample video;

[0014] correcting the first category sequence based on an attention mechanism method to obtain a corrected first category sequence;

[0015] obtaining a target loss function based on the corrected first category sequence, training the detection model according to the target loss function, and returning the detection model to the user.

[0016] According to a third aspect of an embodiment of the present specification, a detection model training method is provided, comprising:

[0017] receiving a calling request sent by a user, wherein the calling request carries a sample video containing an event;

[0018] extracting sample features of the sample video;

[0019] inputting the sample features into a first category module to generate a first category sequence of the sample video;

[0020] correcting the first category sequence based on an attention mechanism method to obtain a corrected first category sequence;

[0021] obtaining a target loss function based on the corrected first category sequence, training the detection model according to the target loss function, and returning the detection model to the user.

[0022] According to a fourth aspect of an embodiment of the present specification, an event detection method is provided, comprising:

[0023] receiving a video containing an event, and inputting the video into a detection model to obtain an event contained in the video and a time when the event occurs, wherein the detection model is obtained by training according to the detection model training method described above;

[0024] generating a corresponding early warning strategy for the event in a case where the event meets a preset early warning condition.

[0025] According to a fifth aspect of an embodiment of the present specification, an event detection method is provided, applied to a city management scenario, comprising:

[0026] receive a video containing a public security event, and input the video into a detection model to obtain the public security event contained in the video and a time when the public security event occurs, wherein the detection model is obtained according to the detection model training method as described above;

[0027] generate a corresponding early warning strategy for the public security event in a case where the public security event meets a preset early warning condition.

[0028] According to a sixth aspect of an embodiment of the present specification, an event detection method is provided, applied to an offline retail scene, comprising:

[0029] receive a video containing a violation event, and input the video into a detection model to obtain the violation event contained in the video and a time when the violation event occurs, wherein the detection model is obtained according to the detection model training method as described above;

[0030] generate a corresponding early warning strategy for the violation event in a case where the violation event meets a preset early warning condition.

[0031] According to a seventh aspect of an embodiment of the present specification, a detection model training device is provided, comprising:

[0032] A first receiving module is configured to receive a sample video containing an event and extract a sample feature of the sample video;

[0033] A first generating module is configured to input the sample feature into a first category module to generate a first category sequence of the sample video;

[0034] A first correcting module is configured to correct the first category sequence based on an attention mechanism method to obtain a corrected first category sequence;

[0035] A first training module is configured to obtain a target loss function based on the corrected first category sequence, and train the detection model according to the target loss function.

[0036] According to an eighth aspect of an embodiment of the present specification, a detection model training device is provided, comprising:

[0037] An interface display module is configured to display a video input interface for a user based on a calling request of the user;

[0038] A second receiving module is configured to receive a sample video containing an event sent by the user based on the video input interface, and extract a sample feature of the sample video;

[0039] A second generating module is configured to input the sample feature into a first category module to generate a first category sequence of the sample video;

[0040] The second correction module is configured to correct the first category sequence based on an attention mechanism method to obtain a corrected first category sequence.

[0041] The second training module is configured to obtain a target loss function based on the corrected first category sequence, train the detection model according to the target loss function, and return the detection model to the user.

[0042] According to a ninth aspect of an embodiment of the present specification, a detection model training apparatus is provided, comprising:

[0043] The request receiving module is configured to receive a calling request sent by a user, wherein the calling request carries a sample video containing an event.

[0044] The feature extraction module is configured to extract sample features of the sample video.

[0045] The third generation module is configured to input the sample features into a first category module to generate a first category sequence of the sample video.

[0046] The third correction module is configured to correct the first category sequence based on an attention mechanism method to obtain a corrected first category sequence.

[0047] The third training module is configured to obtain a target loss function based on the corrected first category sequence, train the detection model according to the target loss function, and return the detection model to the user.

[0048] According to a tenth aspect of an embodiment of the present specification, an event detection apparatus is provided, comprising:

[0049] The first event detection module is configured to receive a video containing an event, input the video into a detection model, and obtain the event contained in the video and the time when the event occurs, wherein the detection model is trained according to the detection model training method described above.

[0050] The first strategy generation module is configured to generate a corresponding early warning strategy for the event if the event meets a preset early warning condition.

[0051] According to an eleventh aspect of an embodiment of the present specification, an event detection apparatus applied to a city management scene is provided, comprising:

[0052] The second event detection module is configured to receive a video containing a public security event, input the video into a detection model, and obtain the public security event contained in the video and a time when the public security event occurs, wherein the detection model is trained according to the detection model training method described above.

[0053] The second strategy generation module is configured to generate a corresponding early warning strategy for the public security event if the public security event meets preset early warning conditions.

[0054] According to a twelfth aspect of an embodiment of the present specification, an event detection device is provided, comprising:

[0055] The third event detection module is configured to receive a video containing a violation event, input the video into a detection model, and obtain the violation event contained in the video and a time when the violation event occurs, wherein the detection model is trained according to the detection model training method described above.

[0056] The third strategy generation module is configured to generate a corresponding early warning strategy for the violation event if the violation event meets preset early warning conditions.

[0057] According to a thirteenth aspect of an embodiment of the present specification, a computing device is provided, comprising:

[0058] a memory and a processor;

[0059] The memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions, which implement steps of the detection model training method or steps of the event detection method when executed by the processor.

[0060] According to a fourteenth aspect of an embodiment of the present specification, a computer readable storage medium is provided, which stores computer executable instructions, which implement steps of the detection model training method or steps of the event detection method when executed by a processor.

[0061] One embodiment of the present specification realizes a detection model training method and device, an event detection method and device. The detection model training method comprises receiving a sample video containing an event, and extracting a sample feature of the sample video; inputting the sample feature into a first category module to generate a first category sequence of the sample video; correcting the first category sequence based on an attention mechanism method to obtain a corrected first category sequence; obtaining a target loss function based on the corrected first category sequence, and training the detection model according to the target loss function. Specifically, the detection model training method corrects the first category sequence generated by the first category module through the attention mechanism method, optimizes the first category sequence by fusing global information, realizes the detection model trained by the detection model training method, and can not only detect the event with obvious action features through the first category module, but also correct the first category sequence of the first category module through the attention mechanism method, realize the detection of the first category module on the non-significant event, so as to realize the comprehensive and accurate mining of the event in the video. BRIEF DESCRIPTION OF DRAWINGS

[0062] Figure 1 is an example diagram of a specific application scene of an event detection method provided by one embodiment of the present specification;

[0063] Figure 2 is a flowchart of a first detection model training method provided by one embodiment of the present specification;

[0064] Figure 3 is a schematic diagram of clustering events in a sample video by using a feature projection method in a detection model training method provided by one embodiment of the present specification;

[0065] Figure 4 is a flowchart of a first event detection method provided by one embodiment of the present specification;

[0066] Figure 5 is a structure schematic diagram of a detection model in an event detection method provided by one embodiment of the present specification;

[0067] Figure 6 is a flowchart of a second detection model training method provided by one embodiment of the present specification;

[0068] Figure 7 is a flowchart of a third detection model training method provided by one embodiment of the present specification;

[0069] Figure 8 is a structure schematic diagram of a first detection model training device provided by one embodiment of the present specification;

[0070] Figure 9is a structural schematic diagram of a second detection model training apparatus provided by an embodiment of the present specification;

[0071] Figure 10 is a structural schematic diagram of a third detection model training apparatus provided by an embodiment of the present specification;

[0072] Figure 11 is a structural schematic diagram of a first event detection apparatus provided by an embodiment of the present specification;

[0073] Figure 12 is a flow chart of a second event detection method provided by an embodiment of the present specification;

[0074] Figure 13 is a flow chart of a third event detection method provided by an embodiment of the present specification;

[0075] Figure 14 is a structural schematic diagram of a second event detection apparatus provided by an embodiment of the present specification;

[0076] Figure 15 is a structural schematic diagram of a third event detection apparatus provided by an embodiment of the present specification;

[0077] Figure 16 is a structural block diagram of a computing device provided by an embodiment of the present specification. DETAILED DESCRIPTION

[0078] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present specification. However, the present specification can be practiced without the specific details, other than in the examples, and it is understood that the scope of the present specification is not limited to the details below.

[0079] The terminology used in one or more embodiments of the present specification is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present specification. As used in one or more embodiments of the present specification and the accompanying claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in one or more embodiments of the present specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0080] It should be understood that, although the terms first, second, etc. can be employed in this specification to describe various information, these information should not be limited to these terms. These terms are only used to differentiate one piece of information from another. For example, a first can be termed a second, and, similarly, a second can be termed a first, without departing from the scope of the one or more embodiments of the present specification. The word "if' as used herein can be interpreted as meaning "when" or "upon" or "in response to determining" depending on the context.

[0081] First, the noun terms related to the one or more embodiments of the present specification are explained.

[0082] CNN: Convolution Neural Network, a type of feedforward neural network containing convolutional computation and having a deep structure.

[0083] MIL: Multi Instance Learning.

[0084] CAS: Class Activation Sequence.

[0085] t-SNE: t-distributed stochastic neighbor embedding, a machine learning algorithm for dimensionality reduction, is one of the algorithms with better effect for visualization.

[0086] THUMOS-14 dataset: contains a large number of open source videos of human actions in real environment.

[0087] ActivityNet dataset: is the largest dataset for current temporal action detection task, covering 200 different daily activities.

[0088] IoU: is a standard for measuring the accuracy of detecting corresponding objects in a specific dataset.

[0089] In the current Pingan, municipal, traffic and other urban brain projects, a large amount of camera video data will be generated every day, how to effectively mine information from these video data is a problem to be solved. The visual analysis based on a single picture cannot effectively utilize the effective information in time sequence, and it is difficult to obtain good results in the video scene. Therefore, the current common practice is to rely on post supervision, that is, when someone reports an abnormal behavior event (such as theft, fighting, etc.), the corresponding period of video is retrieved to view the situation at that time. This post-supervision method can only be used for post-evidence, and cannot find abnormal situations in real time and make warnings and interventions. Therefore, based on computer vision technology, real-time analysis of abnormal behavior in urban scenes has significant practical significance.

[0090] Based on this, the detection model training method provided in the specification is provided. One or more embodiments of the specification also relate to an event detection method, a detection model training device, an event detection device, a computing device, and a computer readable storage medium, which are described in detail one by one in the following embodiments.

[0091] When the detection model training method provided in the specification is applied to analyze abnormal behavior (fighting, arson, theft, etc.) in urban scenes, deep learning technology can be used to analyze the pictures taken by the camera in real time, analyze whether there is abnormal behavior, and once it is found, issue a warning and intervene, which is beneficial to prevent further expansion of the harm in time; In practical applications, the detection model training method can not only be applied to analyze abnormal behavior in urban scenes, but also be applied to other scenes where motion is detected, for example, analyzing sports behavior in sports scenes, using the detection model provided in the specification to extract basketball, swimming, football, running and other sports actions in the video of a certain square. The analysis of the motion action can determine which sports the nearby crowd is more inclined to, so that sports facilities can be added in the square in a targeted manner to enhance the outdoor experience of users.

[0092] In practical applications, the abnormal motion segment detection method specifically refers to inputting a long video, identifying the time period when the abnormal motion occurs through the abnormal motion segment detection method, and returning the start and end time of the motion and the specific category of the motion. The strong supervision based abnormal motion detection method needs a large amount of manually annotated data to train the network, including the start and end time of the abnormal motion in the video and the category of the motion, and the high cost of manual annotation greatly reduces the practicality of the strong supervision based abnormal motion detection method.

[0093] Compared with strong supervision, the weakly supervised abnormal action detection algorithm only needs to mark whether the whole video contains abnormal actions such as fighting, falling, stealing and the like in the training process, without marking the start and end time of the action. Therefore, the weakly supervised abnormal action detection algorithm greatly reduces the annotation workload and is more practical.

[0094] In implementation, the weakly supervised abnormal action detection can be trained by a multi-instance learning (MIL) based method, and specific details are as follows:

[0095] For the input video, first, high-level semantic features are extracted based on a feature extractor, and the extracted video features are marked as X R T×D , where T represents the time dimension, and D represents the action feature dimension of each time point. Then, a two-layer time sequence 1-dimensional convolutional network is used to generate a class activation sequence (CAS), which is represented as follows:

[0096]

[0097] where f conv represents a 1-dimensional convolutional network, and φ1 and φ2 represent network parameters, and the generated A R T×C is a class activation sequence, that is, for each action feature, the corresponding action class probability value is output, where T represents the length of the video in the time dimension, and C is the number of abnormal action classes. For example, if the video is 3 seconds long, the length of the video in the time dimension is 1 second, and the video contains three abnormal actions, that is, three classes, then the generated class activation sequence can obtain the probability that each second of the video belongs to the three classes, such as the probability of belonging to basketball, swimming and running between the 0th second and the 1st second. A 3*3 matrix, that is, a class activation sequence, is generated.

[0098] Since there is no time sequence annotation information of the action segment in the training process, multi-instance learning is usually used to train the network parameters. Specifically, for each action class, the corresponding video action probability (the probability that the video segment contains this class of action) is calculated based on the generated class activation sequence A, and the formula is as follows:

[0099]

[0100] That is, for class C, the k act highest probability values are selected , and the average of c is obtained.

[0101] After obtaining the video action probability for each action category, the network model is trained based on the cross entropy loss function:

[0102]

[0103] Among them, N is the number of samples input in each training, C is the number of abnormal action categories, y=[y1,y2,...,y C ]∈R C It is the video-level action category label provided during training, that is, if action c appears in the video, the corresponding y c is 1 if the value is set, otherwise it is 0.

[0104] Multi-instance learning relies on the cross-entropy loss function generated by action category classification to reversely optimize the entire model parameters. Since classification and detection are two different tasks—the classification task aims to achieve accurate classification based on salient features, while the detection task aims to locate and classify complete action segments—weakly supervised learning methods that rely on multi-instance learning can typically only locate salient regions, and may miss less obvious abnormal actions. For example, for abnormal actions such as fighting, fights with large movements and obvious features are easy to locate and detect, but fights with small movements and less obvious features are often overlooked. Therefore, detection models trained solely through multi-instance learning can only locate some abnormal actions, and the overall effect is not very good.

[0105] See also Figure 1 , Figure 1 An example diagram showing a specific application scenario of an event detection method provided by an embodiment of this specification is shown.

[0106] In the embodiments of this specification, the event detection method is specifically introduced by taking the application of analyzing abnormal behaviors in urban scenarios as an example.

[0107] Figure 1 The application scenario includes an image acquisition terminal 102, an image receiving terminal 104 and a server 106. Specifically, the image receiving terminal 104 receives the video a captured by the image acquisition terminal 102 in real time in an urban scene; after receiving the video a, the image receiving terminal 104 sends the video a to the server 106. After receiving the video a, the server 106 inputs the video a into the detection model. When the detection model detects abnormal behavior (fighting, arson, theft, etc.) in the video a, an early warning is issued and intervention is carried out, which is conducive to timely preventing the further expansion of the harm and maintaining social order.

[0108] In actual application, the detection model is obtained by training based on the related category activation module and the attention mechanism, and can not only detect the abnormal behavior with obvious action features, i.e., high saliency, but also detect the abnormal behavior with unobvious action features in the video a.

[0109] The event detection method provided in the embodiments of the present specification is applied to analyze abnormal behaviors in a city scene, and the detection model obtained by training based on the related category activation module and the attention mechanism can comprehensively and accurately obtain abnormal behaviors in a video, so that corresponding measures can be quickly taken based on the abnormal behaviors to improve user experience.

[0110] Referring to Figure 2 , Figure 2 A flowchart of a first detection model training method according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0111] Step 202: receiving a sample video containing an event, and extracting sample features of the sample video.

[0112] The application scenarios of the detection model training method are different, and the events contained in the sample video are also different; for example, if the detection model training method is applied to analyze abnormal behaviors in a city scene, the event can be understood as a fighting, theft, arson, etc. action event; if the detection model training method is applied to analyze sports categories in a sports scene, the event can be understood as swimming, playing basketball, playing football, running, etc. action event. For ease of understanding, the detection model training method is applied to analyze abnormal behaviors in a city scene in the embodiments of the present specification.

[0113] Specifically, the extraction of the sample features of the sample video includes:

[0114] extracting high-dimensional sample features of the sample video according to a preset feature extractor, wherein the high-dimensional sample features include time dimension features and event dimension features at each time point.

[0115] The preset feature extractor can be set according to actual application, and the present specification does not make any limitation thereon. For example, commonly used feature extractors include I3D, C3D, P3D, etc.

[0116] In specific implementation, the feature extractor is pre-trained, and in the embodiments of the present specification, the high-dimensional features of the sample video can be directly extracted based on the pre-trained parameters. Specifically, for the received sample video containing an event, first, the high-dimensional sample features of each sample video are extracted based on the pre-trained feature extractor, and then the high-dimensional sample features are recorded as X RT×D wherein, T represents a time dimension feature, and D represents an event dimension feature at each time point.

[0117] In the embodiments of the present specification, before the detection model is trained based on the sample video containing events, the high-dimensional features of the sample video are quickly extracted based on the pre-trained feature extractor, and subsequently the detection model can be trained based on the accurate high-dimensional features to improve the speed and accuracy of the detection model training.

[0118] Step 204: inputting the sample feature into the first category module to generate a first category sequence of the sample video.

[0119] Specifically, after the high-dimensional sample feature X of the sample video is extracted, the high-dimensional sample feature X is input into the first category module to generate a first category sequence of the sample video, wherein the first category module can be understood as an initial detection model, and the detection model can be a convolutional neural network.

[0120] Specifically, the sample feature is input into the first category module to generate a first category sequence of the sample video, comprising:

[0121] The sample feature is input into the first category module, and is convolved by a first convolutional layer and a second convolutional layer of the first category module, and a first category sequence of the sample video is generated by a multi-instance learning method.

[0122] In actual application, the first category module can be a related category activation module. After the high-dimensional sample feature X is obtained, the high-dimensional sample feature X is input into the related category activation module, and after the high-dimensional sample feature X is convolved twice by a first convolutional layer and a second convolutional network of the related category activation module, a category activation sequence A (i.e. a first category sequence) of the sample video is generated by a multi-instance learning method. Subsequently, the category activation sequence A obtained by the multi-instance learning method can be used to accurately locate the salient events with obvious action features in the sample video.

[0123] In specific implementation, the sample feature can also be convolved by more than two convolutional networks in the first category module. Specifically, the number of convolutional layers of the sample feature in the first category module can be set according to actual application, and the present specification does not make any limitation in this regard.

[0124] Step 206: correcting the first category sequence based on an attention mechanism method to obtain a corrected first category sequence.

[0125] Specifically, the first category sequence is corrected based on the attention mechanism method to obtain a corrected first category sequence, comprising:

[0126] The sample feature is input into the first category module, and the initial convolution result of the sample feature is obtained by convolution of the first convolution layer of the first category module;

[0127] The initial convolution result is respectively convolved by the third convolution layer and the fourth convolution layer, and the attention matrix of the sample feature is obtained by fusing the convolution results;

[0128] The first category sequence is corrected based on the attention matrix of the sample feature, and a corrected first category sequence is obtained.

[0129] In specific implementation, after obtaining the high-dimensional sample feature X of the sample feature, the high-dimensional sample feature X is first convolved by the first convolution layer and the second convolution layer of the first category module to obtain the category activation sequence A of the sample video; meanwhile, the high-dimensional sample feature X is also convolved by the first convolution layer of the first category module to obtain the feature F after convolution (i.e. the initial convolution result of the sample feature), and then the feature F is respectively convolved by the third convolution layer and the fourth convolution layer of the first category module to obtain the result one after convolution by the third convolution layer and the result two after convolution by the fourth convolution layer, and then the result one or the result two is transposed, and the result one is fused with the transposed result two to obtain the attention matrix of the sample feature, or the transposed result one is fused with the result two to obtain the attention matrix of the sample feature; finally, the first category sequence is corrected based on the attention matrix of the sample feature to obtain a corrected first category sequence.

[0130] In actual application, since the first category sequence can only be positioned to the significant action interval, in order to improve the completeness of action detection, the first category sequence generated by the multi-instance learning can be corrected based on the above attention mechanism method. Specifically, the general formula of the attention mechanism is as follows:

[0131]

[0132] Wherein, x and y are the input and output of the attention mechanism respectively, and i and j are the corresponding position indexes. The output of the attention mechanism is normalized by , and the function g(·) is used to extract the high-dimensional feature representation of each input feature x. f(x i , x j ) is a similarity function, which controls the contribution degree of g(x j ) to the final output y i by the similarity of the outputs x i and y i . Unlike the convolutional neural network (whose output is only determined by the surrounding values), based on the attention mechanism, the global value is based on the similarity function f(·) to y iTherefore, i It has a global receptive field and can capture more effective information. Here, x can be understood as the sample feature, and y can be understood as the attention matrix. The attention mechanism optimizes the category activation sequence A by fusing and comparing the correlations between various motion features. Therefore, motion feature intervals with less pronounced amplitudes can be better captured and mined due to their correlation with salient feature regions in the feature space, further improving the performance of the first category module learned through multi-instance learning.

[0133] In a specific implementation, the initial convolution result is convolved through the third convolution layer and the fourth convolution layer respectively, and the convolution results are fused to obtain the attention matrix of the sample feature, including:

[0134] The initial convolution results are convolved through the third convolution layer and the fourth convolution layer respectively, and the convolution results are fused based on a preset function to obtain the attention matrix of the sample features.

[0135] Among them, the preset functions include but are not limited to the ReLU activation function. The ReLU activation function is an activation function with a rectified linear unit. The ReLU function is actually a piecewise linear function that changes all negative values ​​to 0, while the positive values ​​remain unchanged. This operation is called unilateral inhibition. (That is to say: when the input is a negative value, it will output 0, and then the neuron will not be activated. This means that only some neurons will be activated at the same time, making the network very sparse, which is very efficient for calculation.) It is precisely because of this unilateral inhibition that the neurons in the neural network also have sparse activation. Therefore, the use of this ReLU function can filter out unreasonable similarity outputs, thereby obtaining better attention results (that is, the values ​​in the attention matrix).

[0136] In the embodiment of this specification, the initial convolution result is convolved through the third convolution layer and the fourth convolution layer respectively to obtain result one and result two, and the result one and result two are fused to obtain the initial fusion result. At the same time, the ReLU activation function is introduced to filter out unreasonable similarity outputs, thereby obtaining better attention results.

[0137] In specific implementation, the conventional attention mechanism in Formula 4 is not applicable to this task. In the embodiment of this specification, the correction of the class activation sequence A can be performed using the following Formula 5:

[0138]

[0139] Among them, A i and are the category activation sequences A and F before and after correction respectively. i ∈R T×Dis a 1-dimensional convolutional network that receives X e R T×D As the intermediate semantic features generated after input, each feature has a dimension of D, and the specific details are shown in the above formula 1. The similarity function f(F i , F j ) adopts the form of cosine similarity to correct the category activation sequence, where theta and filter the effective information of the feature F by means of 1*1 convolution. For example, theta and filter the feature dimension of The subsequent R T×T Attention matrix can realize the correction of the category activation sequence A.

[0140] Based on such an implementation form, expanding the receptive field can more easily find abnormal motion intervals outside the salient region. At the same time, in order to avoid overfitting caused by too many parameters, the attention mechanism is further revised in this specification, as shown below:

[0141]

[0142] By the same set of parameters theta to realize feature filtering, the model parameters can be reduced, and overfitting can be inhibited. In addition, when generating the attention matrix, the ReLU activation function is introduced, which can filter unreasonable similarity output and obtain better attention results.

[0143] Step 208: obtaining a target loss function based on the corrected first category sequence, and training the detection model according to the target loss function.

[0144] Specifically, after obtaining the corrected first category sequence, the target loss function of the detection model is obtained based on the first category sequence, and then the network parameters of the detection model are adjusted according to the target loss function, so as to realize the training of the detection model.

[0145] In specific implementation, the attention mechanism can be efficiently embedded into the detection model, and the The cross-entropy loss function of formula 3 is used to train the entire detection model. Since the corrected category activation sequence has a global receptive field at this time, it is not limited to a local area, and therefore it can be better positioned to the non-salient motion segment.

[0146] In the embodiments of the present specification, the detection model training method corrects the first category sequence generated by the first category module through the attention mechanism method, optimizes the first category sequence by fusing global information, and realizes the detection model obtained by training the detection model through the detection model training method. Not only can the first category module detect events with obvious action features, but also can the first category module detect non-significant events through the correction of the first category sequence of the first category module by the attention mechanism method, so as to realize comprehensive and accurate mining of events in the video.

[0147] In another embodiment of the present specification, after the first category sequence is corrected based on the attention mechanism method, the method further includes:

[0148] The sample feature is input into the second category module to generate a second category sequence of the sample video.

[0149] Specifically, after the high-dimensional sample feature X of the sample video is extracted, the high-dimensional sample feature X is input into the first category module to obtain the corrected category activation sequence A, and the high-dimensional sample feature X is also input into the second category module to generate the corrected first category sequence of the sample video and the second category sequence of the sample video. The first category module and the second category module can be understood as two convolutional networks of the entire detection model. For example, if the detection model is a convolutional neural network, the first category module and the second category module are two parallel convolutional network layers of the convolutional neural network.

[0150] In addition, the sample feature is input into the second category module to generate the second category sequence of the sample video, including:

[0151] The sample feature is input into the second category module, and the second category sequence of the sample video generated by the second category module through the feature projection method is obtained.

[0152] In actual application, the second category module can be a feature projection module composed of a projection vector. The projection vector in the feature projection module is pre-trained before use. After obtaining the high-dimensional sample feature X, the high-dimensional sample feature X is input into the feature projection module, and the feature projection module generates the category activation sequence S of the sample video based on the determined projection vector. The feature projection module can be based on the clustering characteristics of the distribution of actions (i.e. events) in the feature space, so that it can also accurately locate abnormal actions with unclear features.

[0153] In the embodiments of the present specification, on the basis of correcting the relevant category activation module based on the attention mechanism method, the feature projection module of the feature projection method is used to assist the relevant category activation module using the multiple instance learning method to train the detection model, so that the detection model obtained by training can realize more accurate detection of events with obvious action features and events with non-obvious action features.

[0154] Still taking the abnormal actions including fighting, arson, theft and the like as examples, for each type of action, the distribution thereof in the feature space has clustering characteristics.

[0155] Referring to Figure 3 , Figure 3 FIG. 1 shows a schematic diagram of clustering events in a sample video according to a detection model training method provided by an embodiment of the present specification.

[0156] Taking the events in Figure 3 as examples, it can be seen from the visualization method based on t-SNE that the features of fighting (Fight) and the features of non-fighting (no fight) are separable in the feature space, i.e., without additional supervision information, fighting and non-fighting can be distinguished by finding a reasonable projection direction. As shown in Figure 3 , the multiple instance learning method can only locate to the salient regions, so it is necessary to use the attention mechanism method to correct the first category sequence generated by the first category module based on the multiple instance learning method to obtain the non-salient regions, and the feature projection method can also relatively completely discover all the actions of this type, so on the basis of correcting the relevant category activation module based on the attention mechanism method, the feature projection module of the feature projection method can be used to assist the relevant category activation module using the multiple instance learning method, which will have more obvious improvement in effect.

[0157] Based on such a finding, for each type of abnormal action c, a corresponding projection vector is learned to separate the target action, and the specific learning method is to learn a reasonable projection vector by training a projection loss function:

[0158]

[0159] Wherein, N is the number of sample videos input each time, T is the length in the time dimension of a single video, C is the number of categories of abnormal actions, y i,c is 1 when the Cth category of action exists in the video, otherwise it is 0, p c is the projection vector corresponding to the category C, and X i,kis an action feature; by minimizing the feature projection loss function, a corresponding projection vector can be learned for each specific abnormal action category, thereby realizing the mining and positioning of the action category.

[0160] In specific implementation, the first category module is a relevant category activation module, and the second category module is a feature projection module. After the high-dimensional sample feature X is input into the relevant category activation module and the feature projection module, the relevant category activation module generates a category activation sequence A of the sample video by using a multi-instance learning method. However, the multi-instance learning method can only locate a significant region, and therefore an attention mechanism method and a feature projection module are used as an auxiliary. Specifically, the feature projection module is composed of a projection vector. The projection vector in the feature projection module is pre-trained. After the high-dimensional sample feature X is input into the feature projection module, the feature projection module generates a category activation sequence S of the sample video based on the determined projection vector. The feature projection module is based on the clustering characteristics of the distribution of actions (i.e., events) in the feature space, and therefore can accurately locate abnormal actions with unobvious features. Therefore, the detection model of the embodiments of the present specification can realize accurate detection of events with obvious action features and unobvious action features after training of the first category module based on the attention mechanism and the second category module.

[0161] In addition, based on the feature projection module that adds the feature projection method assisting the relevant category activation module that uses the multi-instance learning method, the target loss function is obtained based on the modified first category sequence, and the detection model is trained according to the target loss function, including:

[0162] The target loss function is obtained based on the modified first category sequence and the second category sequence, and the detection model is trained according to the target loss function.

[0163] Specifically, after the modified first category sequence and the second category sequence are obtained, the target loss function of the detection model is obtained based on the modified first category sequence and the second category sequence, and then the network parameters of the detection model are adjusted according to the target loss function, so as to realize the training of the detection model.

[0164] Specifically, the target loss function is obtained based on the modified first category sequence and the second category sequence, and the detection model is trained according to the target loss function, including:

[0165] The first loss function of the first category module is obtained based on the modified first category sequence, and the second loss function of the second category module is obtained based on the second category sequence;

[0166] obtain the target loss function based on the first loss function and the second loss function.

[0167] Specifically, in the training process of the detection model, the network parameters of the entire detection model are trained end-to-end, that is, the first category module and the second category module are trained simultaneously.

[0168] Therefore, the target loss function of the detection model can be accurately obtained through the first loss function of the first category module and the second loss function of the second category module.

[0169] In specific implementation, the obtaining of the target loss function based on the first loss function and the second loss function comprises:

[0170] The first loss function and the second loss function are weighted, and the weighted first loss function and the weighted second loss function are added to obtain the target loss function.

[0171] In the above example, still taking the first category module as the relevant category activation module and the second category module as the feature projection module as an example, the relevant category activation module and the feature projection module are trained simultaneously in the training of the detection model. The loss function based on the relevant category activation module and the feature projection module can obtain the loss function of the entire detection model. The loss function of the relevant category activation module is obtained through the corrected category activation sequence A, and the loss function of the feature projection module is obtained through the category activation sequence S. Specifically, the calculation formula of the loss function L activation of the relevant category activation module is shown in the above formula 3, and the calculation formula of the loss function L project of the feature projection module is shown in the above formula 7. Then, the target loss function of the detection model is obtained by adding the loss functions generated by the relevant category activation module and the feature projection module, and is specifically shown as follows:

[0172] L = L activation + ρL project Formula 8

[0173] Wherein, ρ is used to control the relative weight of the two loss functions, and is usually 0.01.

[0174] In the embodiments of the present specification, the target loss function of the detection model is obtained by weighting and adding the first loss function of the first category module and the second loss function of the second category module. Based on the target loss function, the network parameters of the detection model are adjusted, so that the trained detection model can detect events in the video more comprehensively and accurately in subsequent use.

[0175] In the embodiments of the present specification, the method based on the attention mechanism corrects the category activation sequence generated by the related category activation module of the multiple-instance learning, and the feature projection method simultaneously trains the detection model by using the effective information that the action has a clustering distribution in the feature space, which can further improve the performance of the detection model obtained by training. In addition, the second category module based on the feature projection method can be used to replace the method based on the attention mechanism to train the detection model, which can also improve the performance of the detection model as a whole, that is, on the basis of not using the method based on the attention mechanism to correct the category activation sequence generated by the first category module, the second category module based on the feature projection method is used to assist the first category module, and the method for detecting non-significant events in a video can be achieved.

[0176] In the embodiments of the present specification, the first category module in the detection model training method can be regarded as a related category activation module based on the multiple-instance learning method, and the second category module can be regarded as a feature projection module. When the detection model is trained by using the detection model training method, not only can the first category sequence generated by the first category module be corrected by using the method based on the attention mechanism to locate the abnormal events with obvious and non-obvious action features, but also the second category module can be used to project the vector based on the clustering feature of the event distribution in the feature space to further mine the abnormal events with non-obvious action features, so that the comprehensive and accurate positioning and detection of all abnormal events in the video can be achieved.

[0177] Referring to Figure 4 , Figure 4 A flowchart of a first event detection method provided by an embodiment of the present specification is shown, which specifically includes the following steps.

[0178] Step 402: receiving a video containing an event, and inputting the video into a detection model to obtain the event contained in the video and the time when the event occurs, wherein the detection model is obtained by training according to the detection model training method as described above.

[0179] The application scenarios of the event detection method are different, and the events contained in the video are also different. For example, if the event detection method is applied to analyze abnormal behaviors in a city scene, the event can be understood as a motion event such as fighting, theft, and arson. If the event detection method is applied to analyze the types of sports in a sports scene, the event can be understood as a motion event such as swimming, playing basketball, playing football, and running.

[0180] Specifically, after the video is input into the detection model, a plurality of events contained in the video and the start and end time of each event can be obtained.

[0181] Step 404: generating a corresponding early warning strategy for the event in the case that the event meets preset early warning conditions.

[0182] Optionally, the receiving the video containing the event comprises:

[0183] displaying a video input interface for the user based on the calling request of the user, and receiving the video containing the event sent by the user based on the video input interface.

[0184] Optionally, the receiving the video containing the event comprises:

[0185] receiving a calling request sent by the user, wherein the calling request carries the video containing the event.

[0186] The preset early warning conditions are set according to actual applications, and the present specification does not make any limitation on this. For example, the event detection method is applied in analyzing abnormal behaviors in a city scene. The early warning conditions can include arson and robbery. When the event detected from the video contains an event of arson and / or robbery, it is indicated that the event meets the preset early warning conditions. In this case, a corresponding early warning strategy is generated for the event, such as manual intervention or shouting for warning.

[0187] The event detection method provided by the present specification can realize comprehensive, accurate and real-time prediction of events in a video by using a detection model trained by a related category activation module of a multi-instance learning method based on attention mechanism, and then generate a reasonable early warning strategy according to the predicted event, thereby improving user experience.

[0188] Referring to Figure 5 , Figure 5 FIG. 1 shows a structure schematic diagram of a detection model in an event detection method provided by an embodiment of the present specification.

[0189] Specifically, for Figure 5When the detection model in the event detection method is trained, the video acquired according to time is input to the feature extractor, the feature extractor extracts features of the input video, extracts video high-dimensional features of the video, and then inputs the video high-dimensional features to the first convolutional layer of the related category activation module to generate features F. The features F are input to the second convolutional layer to generate category activation sequence A, and at the same time, the features F are input to the third convolutional layer to filter out effective information TxD1 of the features F through 1*1 convolution, and the features F are input to the fourth convolutional layer to filter out effective information D1XT of the features F through 1*1 convolution. Then, the above two effective information are fused to obtain an attention matrix TXT, and the category activation sequence A is corrected based on the attention matrix TXT, and a loss function of the related category activation module is generated based on the corrected category activation sequence A. Finally, the network parameters of the detection model are adjusted according to the loss function of the related category activation module, so as to realize the training of the detection model.

[0190] In actual application, after the video is input to the detection model, high-dimensional feature extraction is also performed, the video is input to the feature projection module, the category activation sequence S is generated, then the corrected category activation sequence A and the category activation sequence S are weighted and fused to obtain a new category activation sequence, and the new category activation sequence can be used to locate more complete abnormal actions, and the recall rate and accuracy of abnormal work are improved.

[0191] Specifically, the category activation sequences generated by the related category activation module and the feature projection module are fused, and are expressed as follows:

[0192] A final =A+γS Formula 9

[0193] Wherein, γ is a hyperparameter for adjusting the weights of the two category activation sequences, and in actual landing process, the value is usually 0.5.

[0194] The detection model in the event detection method provided by the embodiments of the present specification is trained by the related category activation module of the multi-instance learning method of the attention mechanism and the feature projection module of the feature projection method, and when the detection model is used to detect events in the video, the events in the video can be detected more comprehensively, accurately and in real time. The recall rate and accuracy of the events in the video can be greatly improved due to the fusion of the new category activation sequence in the detection model.

[0195] To verify the performance of the detection model obtained by training, the feature projection method and the attention mechanism method are added respectively to the classic model STPN (Sparse Temporal Pooling Network) as a reference on the data sets THUMOS-14 and ActivityNet, and the indicators are mAP (mean average precision) under different IoU (Intersection over Union). The comparison results are shown in the following tables.

[0196] Table 1: Comparison experiment after introducing the attention mechanism method and the feature projection method on the THUMOS-14 data set

[0197]

[0198] Table 2: Comparison experiment after introducing the attention mechanism method and the feature projection method on the ActivityNet data set

[0199] Model name IoU=0.5 IoU=0.75 IoU=0.95 Average STPN 29.3 16.9 2.6 17.3 STPN+attention 32.1 19.4 4.2 20.1 STPN+feature projection 34.0 21.9 5.1 21.9

[0200] As can be seen from Tables 1 and 2, after introducing the attention mechanism method and the feature projection method, the indicators on the THUMOS-14 and ActivityNet data sets are significantly improved; and compared with STPN+attention, the indicators of STPN+feature projection are more significantly improved, so the indicators of STPN+attention+feature projection will be better, which verifies the effectiveness of the detection model training method of the present specification.

[0201] In the embodiments of the present specification, in view of the problem that the current abnormal action detection method can only locate the saliency region, the attention mechanism is used to realize the positioning and detection of the target action. The problem that multiple-instance learning can only find saliency action is avoided, and the performance is significantly improved on the THUMOS-14 and ActivityNet data sets, which verifies the effectiveness of the present scheme.

[0202] Referring to Figure 6 , Figure 6 A flowchart of a second detection model training method provided by an embodiment of the present specification is shown, which specifically includes the following steps.

[0203] Step 602: Based on the calling request of a user, a video input interface is displayed for the user.

[0204] Step 604: Receive a sample video containing an event sent by the user based on the video input interface, and extract sample features of the sample video.

[0205] Step 606: input the sample feature into the first category module to generate a first category sequence of the sample video.

[0206] Step 608: correct the first category sequence based on an attention mechanism method to obtain a corrected first category sequence.

[0207] Step 610: obtain a target loss function based on the corrected first category sequence, train the detection model according to the target loss function, and return the detection model to the user.

[0208] The detection model training method provided by the embodiments of the present specification corrects the first category sequence generated by the first category module through the attention mechanism method, optimizes the first category sequence by fusing global information, and realizes the detection model obtained by training through the detection model training method. Not only can the first category module detect events with obvious action features, but also can the first category module detect non-significant events through the correction of the first category sequence of the first category module by the attention mechanism method, so that comprehensive and accurate mining of events in a video can be realized.

[0209] The above is a schematic scheme of the second detection model training method of the present embodiment. It should be noted that the technical scheme of the second detection model training method belongs to the same concept as the technical scheme of the first detection model training method described above, and the details of the technical scheme of the second detection model training method that are not described in detail can be referred to the description of the technical scheme of the first detection model training method.

[0210] Referring to Figure 7 , Figure 7 A flowchart of a third detection model training method provided by an embodiment of the present specification is shown, which specifically includes the following steps.

[0211] Step 702: receive a calling request sent by a user, wherein the calling request carries a sample video containing an event.

[0212] Step 704: extract a sample feature of the sample video.

[0213] Step 706: input the sample feature into the first category module to generate a first category sequence of the sample video.

[0214] Step 708: correct the first category sequence based on an attention mechanism method to obtain a corrected first category sequence.

[0215] Step 710: obtain a target loss function based on the corrected first category sequence, train the detection model according to the target loss function, and return the detection model to the user.

[0216] The detection model training method provided by the embodiment of the present specification corrects the first category sequence generated by the first category module through the attention mechanism method, optimizes the first category sequence by fusing global information, and realizes the detection model obtained by training the detection model training method. Not only can the first category module detect events with obvious action features, but also the first category module can detect non-significant events through the correction of the first category sequence of the first category module by the attention mechanism method, so that comprehensive and accurate mining of events in the video can be realized.

[0217] The above is a schematic scheme of the third detection model training method of the embodiment. It should be noted that the technical scheme of the third detection model training method belongs to the same concept as the technical scheme of the first detection model training method described above. The details of the technical scheme of the third detection model training method that are not described in detail can be referred to the description of the technical scheme of the first detection model training method.

[0218] Corresponding to the above method embodiment, the present specification also provides a detection model training device embodiment, Figure 8 The structure schematic diagram of the first detection model training device provided by an embodiment of the present specification is shown. As shown in the figure, Figure 8 The device comprises:

[0219] The first receiving module 802 is configured to receive a sample video containing an event and extract sample features of the sample video;

[0220] The first generation module 804 is configured to input the sample features into the first category module to generate a first category sequence of the sample video;

[0221] The first correction module 806 is configured to correct the first category sequence based on the attention mechanism method to obtain a corrected first category sequence;

[0222] The first training module 808 is configured to obtain a target loss function based on the corrected first category sequence, and train the detection model according to the target loss function.

[0223] Optionally, the first receiving module 802 is further configured to:

[0224] extract high-dimensional sample features of the sample video according to a preset feature extractor, wherein the high-dimensional sample features include time dimension features and event dimension features at each time point.

[0225] Optionally, the first generation module 804 is further configured to:

[0226] The sample feature is input into the first category module, is convolved by a first convolutional layer and a second convolutional layer of the first category module, and a first category sequence of the sample video is generated by a multi-instance learning method.

[0227] Optionally, the first correction module 806 is further configured to:

[0228] The sample feature is input into the first category module, is convolved by a first convolutional layer of the first category module to obtain an initial convolutional result of the sample feature;

[0229] The initial convolutional result is respectively convolved by a third convolutional layer and a fourth convolutional layer and the convolutional results are fused to obtain an attention matrix of the sample feature;

[0230] The first category sequence is corrected based on the attention matrix of the sample feature to obtain a corrected first category sequence.

[0231] Optionally, the first correction module 806 is further configured to:

[0232] The initial convolutional result is respectively convolved by a third convolutional layer and a fourth convolutional layer and the convolutional results are fused based on a preset function to obtain an attention matrix of the sample feature.

[0233] Optionally, the apparatus further comprises:

[0234] A fourth generation module configured to input the sample feature into a second category module to generate a second category sequence of the sample video.

[0235] Optionally, the first training module 808 is further configured to:

[0236] A target loss function is obtained based on the corrected first category sequence and the second category sequence, and the detection model is trained according to the target loss function.

[0237] Optionally, the fourth generation module is further configured to:

[0238] The sample feature is input into the second category module, and a second category sequence of the sample video generated by the second category module through a feature projection method is obtained.

[0239] Optionally, the first training module 808 is further configured to:

[0240] A first loss function of the first category module is obtained based on the corrected first category sequence, and a second loss function of the second category module is obtained based on the second category sequence.

[0241] obtaining the target loss function based on the first loss function and the second loss function.

[0242] Optionally, the first training module 808 is further configured to:

[0243] weighting the first loss function and the second loss function, adding the weighted first loss function and the second loss function to obtain the target loss function.

[0244] The detection model training device provided by the embodiment of the present specification can correct the first category sequence generated by the first category module through the attention mechanism method, fuse global information to optimize the first category sequence, and realize the detection model trained by the detection model training method. The detection model trained by the detection model training method can not only detect events with obvious action features through the first category module, but also correct the first category sequence of the first category module through the attention mechanism method, realize the detection of non-significant events by the first category module, and thus can realize comprehensive and accurate mining of events in a video.

[0245] The above is a schematic scheme of the first detection model training device of the embodiment. It should be noted that the technical scheme of the first detection model training device belongs to the same concept as the technical scheme of the first detection model training method described above. The technical scheme of the first detection model training device is not described in detail, and the description of the technical scheme of the first detection model training method can be referred to.

[0246] Corresponding to the method embodiments described above, the present specification also provides detection model training device embodiments, Figure 9 The structure of the second detection model training device provided by an embodiment of the present specification is shown. As shown in the figure, Figure 9 The device comprises:

[0247] The interface display module 902 is configured to display a video input interface for the user based on the user's call request;

[0248] The second receiving module 904 is configured to receive a sample video containing an event sent by the user based on the video input interface, and extract sample features of the sample video;

[0249] The second generation module 906 is configured to input the sample features into the first category module to generate a first category sequence of the sample video;

[0250] The second correction module 908 is configured to correct the first category sequence based on the attention mechanism method to obtain a corrected first category sequence;

[0251] The second training module 910 is configured to obtain a target loss function based on the corrected first category sequence, train the detection model according to the target loss function, and return the detection model to the user.

[0252] The detection model training device provided by the embodiments of the present specification corrects the first category sequence generated by the first category module through the attention mechanism method, optimizes the first category sequence by fusing global information, and realizes the detection model trained by the detection model training method. The detection model can not only detect events with obvious action features through the first category module, but also correct the first category sequence of the first category module through the attention mechanism method, realize the detection of non-significant events by the first category module, and thus can realize comprehensive and accurate mining of events in a video.

[0253] The above is a schematic scheme of the second detection model training device of the present embodiment. It should be noted that the technical scheme of the second detection model training device belongs to the same concept as the technical scheme of the second detection model training method described above. The technical scheme of the second detection model training device, which is not described in detail, can be referred to the description of the technical scheme of the second detection model training method.

[0254] Corresponding to the method embodiments described above, the present specification also provides a detection model training device embodiment, Figure 10 The structure schematic diagram of the third detection model training device provided by an embodiment of the present specification is shown. As shown in the figure, Figure 10 The device comprises:

[0255] The request receiving module 1002 is configured to receive a calling request sent by a user, wherein the calling request carries a sample video containing an event;

[0256] The feature extraction module 1004 is configured to extract sample features of the sample video;

[0257] The third generation module 1006 is configured to input the sample features into the first category module to generate a first category sequence of the sample video;

[0258] The third correction module 1008 is configured to correct the first category sequence based on the attention mechanism method to obtain a corrected first category sequence;

[0259] The third training module 1010 is configured to obtain a target loss function based on the corrected first category sequence, train the detection model according to the target loss function, and return the detection model to the user.

[0260] The detection model training device provided by the embodiment of the present specification corrects the first category sequence generated by the first category module through the attention mechanism method, optimizes the first category sequence by fusing global information, and realizes the detection model obtained by the detection model training method. The detection model not only can detect events with obvious action features through the first category module, but also can correct the first category sequence of the first category module through the attention mechanism method, realize the detection of non-significant events by the first category module, and thus can realize comprehensive and accurate mining of events in the video.

[0261] The above is a schematic scheme of the third detection model training device of the present embodiment. It should be noted that the technical scheme of the third detection model training device belongs to the same concept as the technical scheme of the third detection model training method described above. The technical scheme of the third detection model training device, which is not described in detail, can be referred to the description of the technical scheme of the third detection model training method.

[0262] Corresponding to the method embodiment described above, the present specification also provides an event detection device embodiment, Figure 11 The structure of an event detection device provided by an embodiment of the present specification is shown. As shown in Figure 11 The device comprises:

[0263] The first event detection module 1102 is configured to receive a video containing an event, and input the video into a detection model to obtain the event contained in the video and the time when the event occurs, wherein the detection model is obtained according to the detection model training method described above.

[0264] The first strategy generation module 1104 is configured to generate a corresponding early warning strategy for the event if the event meets a preset early warning condition.

[0265] Optionally, the first event detection module 1102 is further configured to:

[0266] Based on the calling request of the user, a video input interface is displayed to the user, and a video containing an event sent by the user based on the video input interface is received.

[0267] Optionally, the first event detection module 1102 is further configured to:

[0268] Receive a calling request sent by a user, wherein the calling request carries a video containing an event.

[0269] The event detection device provided in the embodiments of the present specification can realize comprehensive, accurate and real-time prediction of events in a video by using a detection model trained by a relevant category activation module of a multi-instance learning method using an attention mechanism, and then generate a reasonable early warning strategy according to the predicted events, thereby improving user experience.

[0270] The above is a schematic scheme of the event detection device of the present embodiment. It should be noted that the technical scheme of the event detection device belongs to the same concept as the technical scheme of the event detection method described above, and the details of the technical scheme of the event detection device that are not described in detail can be referred to the description of the technical scheme of the event detection method.

[0271] Referring to Figure 12 , Figure 12 A flowchart of a second event detection method provided by an embodiment of the present specification is shown, wherein the method is applied to a city management scenario, and specifically includes the following steps.

[0272] Step 1202: receiving a video containing a public security event, and inputting the video into a detection model to obtain a public security event contained in the video and a time at which the public security event occurs, wherein the detection model is obtained by training according to the detection model training method described above.

[0273] The public security event includes but is not limited to action events such as fighting, theft and brawling in a city management scenario.

[0274] Step 1204: generating a corresponding early warning strategy for the public security event in a case where the public security event meets a preset early warning condition.

[0275] The preset early warning condition is set according to actual application, for example, the preset early warning condition is to generate an early warning strategy in a case where the public security event is a arson event.

[0276] The event detection method provided in the embodiments of the present specification is applied to a city management scenario, and can realize comprehensive, accurate and real-time prediction of a public security event in a city management scenario obtained by using a detection model trained by a relevant category activation module of a multi-instance learning method using an attention mechanism, and then generate a reasonable early warning strategy according to the predicted public security event, so as to better maintain the safety of city management.

[0277] The above is a schematic scheme of the second event detection method of the present embodiment. It should be noted that the technical scheme of the event detection method belongs to the same concept as the technical scheme of the first event detection method described above, and the details of the technical scheme of the event detection method that are not described in detail can be referred to the description of the technical scheme of the first event detection method.

[0278] Referring toFigure 13 , Figure 13 A flowchart of a third event detection method provided by an embodiment of the present specification is shown, wherein the method is applied to an offline retail scene and specifically includes the following steps.

[0279] Step 1302: receiving a video containing a violation event and inputting the video into a detection model to obtain the violation event contained in the video and the time when the violation event occurs, wherein the detection model is obtained according to the detection model training method as described above.

[0280] The violation event includes but is not limited to action events such as theft and damage to sold items in the offline retail scene.

[0281] Step 1304: generating a corresponding early warning strategy for the violation event if the violation event meets a preset early warning condition.

[0282] The preset early warning condition is set according to actual application, for example, the preset early warning condition is to generate an early warning strategy when the violation event is a damage to sold items event.

[0283] The event detection method provided by the embodiment of the present specification is applied to an offline retail scene, which can realize comprehensive, accurate and real-time prediction of the violation event in the offline retail scene obtained by the detection model trained according to the related category activation module of the multi-instance learning method with attention mechanism, and then generate a reasonable early warning strategy according to the predicted violation event to avoid loss of various resources and goods in the offline retail scene.

[0284] The above is a schematic scheme of the third event detection method of the present embodiment. It should be noted that the technical scheme of the event detection method belongs to the same concept as the technical scheme of the first event detection method described above, and the details of the technical scheme of the event detection method not described in detail can be referred to the description of the technical scheme of the first event detection method.

[0285] Referring to Figure 14 , Figure 14 A flowchart of a second event detection device provided by an embodiment of the present specification is shown, wherein the device is applied to a city management scene and includes:

[0286] The second event detection module 1402 is configured to receive a video containing a public security event and input the video into a detection model to obtain the public security event contained in the video and the time when the public security event occurs, wherein the detection model is obtained according to the detection model training method as described above.

[0287] The second strategy generation module 1404 is configured to generate a corresponding early warning strategy for the public security event if the public security event meets preset early warning conditions.

[0288] The event detection device provided in the embodiments of the present specification is applied to the urban management scene, and can realize comprehensive, accurate and real-time prediction of the public security events in the urban management scene by using the detection model trained by the relevant category activation module of the multi-instance learning method with the attention mechanism, and then generate reasonable early warning strategies according to the predicted public security events, so as to better maintain the safety of urban management.

[0289] The above is a schematic scheme of the second event detection device of the present embodiment. It should be noted that the technical scheme of the event detection device belongs to the same concept as the technical scheme of the second event detection method described above, and the details of the technical scheme of the event detection device that are not described in detail can be referred to the description of the technical scheme of the second event detection method.

[0290] Referring to Figure 15 , Figure 15 A flowchart of a third event detection device provided by an embodiment of the present specification is shown, wherein the device is applied to an offline retail scene and includes:

[0291] The third event detection module 1502 is configured to receive a video containing a violation event, and input the video into a detection model to obtain the violation event contained in the video and the time when the violation event occurs, wherein the detection model is obtained by training according to the detection model training method described above.

[0292] The third strategy generation module 1504 is configured to generate a corresponding early warning strategy for the violation event if the violation event meets preset early warning conditions.

[0293] The event detection device provided in the embodiments of the present specification is applied to the offline retail scene, and can realize comprehensive, accurate and real-time prediction of the violation events in the offline retail scene by using the detection model trained by the relevant category activation module of the multi-instance learning method with the attention mechanism, and then generate reasonable early warning strategies according to the predicted violation events, so as to avoid the loss of various resources and goods in the offline retail scene.

[0294] The above is a schematic scheme of the third event detection device of the present embodiment. It should be noted that the technical scheme of the event detection device belongs to the same concept as the technical scheme of the third event detection method described above, and the details of the technical scheme of the event detection device that are not described in detail can be referred to the description of the technical scheme of the third event detection method.

[0295] Figure 16 A structural block diagram of a computing device 1600 is shown, according to one embodiment of the present specification. The components of the computing device 1600 include, but are not limited to, a memory 1610 and a processor 1620. The processor 1620 is connected with the memory 1610 through a bus 1630, and a database 1650 is used to save data.

[0296] The computing device 1600 also includes an access device 1640, which enables the computing device 1600 to communicate via one or more networks 1660. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1640 can include one or more of any type of network interface (e.g., a network interface card (NIC)), wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a near-field communication (NFC) interface, and the like.

[0297] In one embodiment of the present specification, the above-mentioned components of the computing device 1600 and other components not shown in the present specification can be connected with each other, for example, through a bus. It should be understood that, Figure 16 the above-mentioned components of the computing device 1600 and other components not shown in the present specification can be connected with each other, for example, through a bus. It should be understood that, Figure 16 The structural block diagram of the computing device shown is merely for the purpose of example, and is not a limitation on the scope of the present specification. Other components can be added or replaced by those skilled in the art as needed.

[0298] The computing device 1600 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smartglasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a PC. The computing device 1600 can also be a mobile or stationary server.

[0299] The processor 1620 is configured to execute computer-executable instructions, which, when executed by the processor, implement the steps of the detection model training method or the steps of the event detection method.

[0300] The above is a schematic scheme of the computing device of the embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the detection model training method or the event detection method described above belong to the same concept, and the details of the technical scheme of the computing device that are not described in detail can be referred to the description of the technical scheme of the detection model training method or the event detection method.

[0301] An embodiment of the present specification also provides a computer readable storage medium storing computer instructions, which, when executed by a processor, implement the steps of the detection model training method or the steps of the event detection method.

[0302] The above is a schematic scheme of the computer readable storage medium of the embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the detection model training method or the event detection method described above belong to the same concept, and the details of the technical scheme of the storage medium that are not described in detail can be referred to the description of the technical scheme of the detection model training method or the event detection method.

[0303] The above describes specific embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.

[0304] The computer instructions include computer program code, which can be in the form of source code, object code, executable code, or some intermediate form. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.

[0305] It should be noted that, for the aforementioned method embodiments, the sequences of the described actions are not necessarily required to implement the present application, and certain actions can be performed in other sequences, or even at the same time, in accordance with the present application. Furthermore, certain actions can not be required to implement the present application. Additionally, the described embodiments are not necessarily the only possible implementation of the present application.

[0306] In the above embodiments, the description of each embodiment is focused on the aspects of the embodiment, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0307] The preferred embodiments of the present application disclosed above are only used to help explain the present application. The alternative embodiments do not describe all the details and limit the present application to the specific embodiments described. Obviously, according to the content of the present application, many modifications and changes can be made. The present application selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present application, so that those skilled in the art can well understand and use the present application. The present application is limited by the claims and their full scope and equivalents.

Claims

1. A detection model training method, comprising: receiving a sample video containing events, and extracting sample features of the sample video; inputting the sample features into a first category module, and generating a first category sequence of the sample video by a multi-instance learning method; correcting the first category sequence based on an attention mechanism method to obtain a corrected first category sequence, wherein the attention mechanism method corrects the first category sequence by optimizing the first category sequence based on the correlation between each action feature in the sample video; obtaining a target loss function based on the corrected first category sequence, and training the detection model according to the target loss function.

2. The detection model training method of claim 1, wherein the extracting the sample features of the sample video comprises: extracting high-dimensional sample features of the sample video according to a preset feature extractor, wherein the high-dimensional sample features include time dimension features and event dimension features at each time point.

3. The detection model training method of claim 1, wherein the inputting the sample features into the first category module and generating the first category sequence of the sample video by the multi-instance learning method comprises: inputting the sample features into the first category module, performing convolution through a first convolution layer and a second convolution layer of the first category module, and generating the first category sequence of the sample video by the multi-instance learning method.

4. The detection model training method of any one of claims 1-3, wherein the correcting the first category sequence based on the attention mechanism method to obtain the corrected first category sequence comprises: inputting the sample features into the first category module, performing convolution through a first convolution layer of the first category module to obtain an initial convolution result of the sample features; performing convolution through a third convolution layer and a fourth convolution layer and fusing the convolution results to obtain an attention matrix of the sample features; correcting the first category sequence based on the attention matrix of the sample features to obtain the corrected first category sequence.

5. The detection model training method of claim 4, wherein the performing convolution through the third convolution layer and the fourth convolution layer and fusing the convolution results to obtain the attention matrix of the sample features comprises: performing convolution through the third convolution layer and the fourth convolution layer and fusing the convolution results based on a preset function to obtain the attention matrix of the sample features.

6. The detection model training method of claim 1, further comprising: inputting the sample features into a second category module to generate a second category sequence of the sample video after the correcting the first category sequence based on the attention mechanism method to obtain the corrected first category sequence.

7. The detection model training method of claim 6, wherein the target loss function is obtained based on the corrected first category sequence and the second category sequence, and the detection model is trained according to the target loss function.

8. The detection model training method of claim 6, wherein the sample feature is input into the second category module to generate the second category sequence of the sample video.

9. The detection model training method of claim 7 or 8, wherein the target loss function is obtained based on the corrected first category sequence and the second category sequence, and the detection model is trained according to the target loss function.

10. The detection model training method of claim 9, wherein the target loss function is obtained based on the first loss function and the second loss function.

11. A detection model training method, comprising: displaying a video input interface for a user based on a calling request of the user; receiving a sample video containing an event sent by the user based on the video input interface, and extracting a sample feature of the sample video; inputting the sample feature into a first category module to generate a first category sequence of the sample video by a multi-instance learning method; correcting the first category sequence based on an attention mechanism method to obtain a corrected first category sequence, wherein the attention mechanism method corrects the first category sequence by optimizing the first category sequence by fusing and comparing the correlation between each action feature in the sample video; obtaining a target loss function based on the corrected first category sequence, and training the detection model according to the target loss function, and returning the detection model to the user.

12. A detection model training method, comprising: receiving a calling request sent by a user, wherein the calling request carries a sample video containing an event; extracting a sample feature of the sample video; inputting the sample feature into a first category module to generate a first category sequence of the sample video by a multi-instance learning method; ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ The attention mechanism-based method corrects the first category sequence to obtain a corrected first category sequence. The attention mechanism-based method corrects the first category sequence by optimizing the first category sequence through fusion and comparison of the correlation between each action feature in the sample video. A target loss function is obtained based on the corrected first category sequence, and the detection model is trained according to the target loss function, and the detection model is returned to the user.

13. An event detection method, comprising: receiving a video containing an event, and inputting the video into a detection model to obtain the event contained in the video and the time when the event occurs, wherein the detection model is trained according to the detection model training method of any one of claims 1-10; generating a corresponding early warning strategy for the event if the event meets a preset early warning condition.

14. The event detection method of claim 13, wherein the receiving a video containing an event comprises: displaying a video input interface to a user based on a call request of the user, and receiving a video containing an event sent by the user based on the video input interface.

15. The event detection method of claim 13, wherein the receiving a video containing an event comprises: receiving a call request sent by a user, wherein the call request carries a video containing an event.

16. An event detection method applied to a city management scenario, comprising: receiving a video containing a public security event, and inputting the video into a detection model to obtain the public security event contained in the video and the time when the public security event occurs, wherein the detection model is trained according to the detection model training method of any one of claims 1-10; generating a corresponding early warning strategy for the public security event if the public security event meets a preset early warning condition.

17. An event detection method applied to an offline retail scenario, comprising: receiving a video containing a violation event, and inputting the video into a detection model to obtain the violation event contained in the video and the time when the violation event occurs, wherein the detection model is trained according to the detection model training method of any one of claims 1-10; generating a corresponding early warning strategy for the violation event if the violation event meets a preset early warning condition.

18. A detection model training apparatus, comprising: a first receiving module configured to receive a sample video containing an event, and extract sample features of the sample video; a first generating module configured to input the sample features into a first category module, and generate a first category sequence of the sample video through a multi-instance learning method. a first correction module configured to correct the first category sequence based on an attention mechanism method, to obtain a corrected first category sequence, wherein the attention mechanism method corrects the first category sequence by optimizing the first category sequence by fusing and comparing the correlations between the action features in the sample video; a first training module configured to obtain a target loss function based on the corrected first category sequence, and train the detection model according to the target loss function.

19. A detection model training apparatus, comprising: an interface display module configured to display a video input interface for a user based on a call request of the user; a second receiving module configured to receive a sample video containing an event sent by the user based on the video input interface, and extract sample features of the sample video; a second generation module configured to input the sample features into a first category module, and generate a first category sequence of the sample video by a multi-instance learning method; a second correction module configured to correct the first category sequence based on an attention mechanism method, to obtain a corrected first category sequence, wherein the attention mechanism method corrects the first category sequence by optimizing the first category sequence by fusing and comparing the correlations between the action features in the sample video; a second training module configured to obtain a target loss function based on the corrected first category sequence, and train the detection model according to the target loss function, and return the detection model to the user.

20. A detection model training apparatus, comprising: a request receiving module configured to receive a call request sent by a user, wherein the call request carries a sample video containing an event; a feature extraction module configured to extract sample features of the sample video; a third generation module configured to input the sample features into a first category module, and generate a first category sequence of the sample video by a multi-instance learning method; a third correction module configured to correct the first category sequence based on an attention mechanism method, to obtain a corrected first category sequence, wherein the attention mechanism method corrects the first category sequence by optimizing the first category sequence by fusing and comparing the correlations between the action features in the sample video; a third training module configured to obtain a target loss function based on the corrected first category sequence, and train the detection model according to the target loss function, and return the detection model to the user.

21. An event detection apparatus, comprising: a first event detection module configured to receive a video containing an event, and input the video into a detection model, to obtain the event contained in the video and the time when the event occurs, wherein the detection model is trained according to the detection model training method in any one of claims 1-10. The first strategy generation module is configured to generate a corresponding early warning strategy for the event if the event meets preset early warning conditions.

22. An event detection device applied to a city management scenario, comprising: The second event detection module is configured to receive a video containing a public security event, input the video into a detection model, and obtain the public security event contained in the video and a time when the public security event occurs, wherein the detection model is trained according to the detection model training method in any one of claims 1-10. The second strategy generation module is configured to generate a corresponding early warning strategy for the public security event if the public security event meets preset early warning conditions.

23. An event detection device applied to an offline retail scenario, comprising: The third event detection module is configured to receive a video containing a violation event, input the video into a detection model, and obtain the violation event contained in the video and a time when the violation event occurs, wherein the detection model is trained according to the detection model training method in any one of claims 1-10. The third strategy generation module is configured to generate a corresponding early warning strategy for the violation event if the violation event meets preset early warning conditions.

24. A computing device, comprising: a memory and a processor; The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions, which realize the steps of the detection model training method in any one of claims 1-10, 11, and 12 and the steps of the event detection method in any one of claims 13-15, 16, and 17.

25. A computer readable storage medium storing computer instructions, which realize the steps of the detection model training method in any one of claims 1-10, 11, and 12 and the steps of the event detection method in any one of claims 13-15, 16, and 17 when executed by a processor.

Citation Information

Patent Citations

  • Target tracking method for unsupervised similarity discriminant learning

    CN110569793A

  • Transformer substation personnel behavior recognition method based on monitoring video time sequence action positioning and anomaly detection

    CN111291699A

  • Weak supervision time sequence action detection method and system based on adaptive sampling

    CN111652083A