Anomaly Behavior Detection Method, Device, Electronic Device and Storage Medium

By detecting and tracking video frames, determining the detection frame area and identifying the trajectory sequence, the problem of being unable to detect single and multi-person behaviors in the prior art is solved, and a more intelligent and efficient abnormal behavior detection is achieved.

CN114842392BActive Publication Date: 2025-07-22SHANGHAI SENSETIME INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210524859.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-13
Publication Date
2025-07-22
Estimated Expiration
2042-05-13

AI Technical Summary

Technical Problem

The prior art cannot adaptively detect abnormalities in single-person behavior and multiple-person behavior simultaneously, resulting in low intelligence in behavior detection.

Method used

By detecting and tracking the current video frame in the video sequence, the detection frame area is determined, the object area is classified according to the overlap, and the trajectory sequence is determined in combination with historical video frames to identify abnormal behaviors.

Benefits of technology

It realizes adaptively detecting abnormal behaviors of single objects and multiple objects simultaneously, improving the intelligence and accuracy of abnormal behavior detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114842392B_ABST
    Figure CN114842392B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose an abnormal behavior detection method, apparatus, electronic device, and storage medium. The method includes: detecting and tracking a current video frame in a video sequence to determine a detection box area of each existing target object; determining an overlap degree between different detection box areas, classifying each detection box area according to the overlap degree to determine a first single object area and a first multiple object areas in the current video frame; determining a single object trajectory sequence to be recognized and a multiple object trajectory sequence to be recognized according to a second single object area, a second multiple object areas, the first single object area, and the first multiple object areas corresponding to historical video frames in the video sequence; and performing abnormal behavior recognition on both the single object trajectory sequence to be recognized and the multiple object trajectory sequence to be recognized to obtain corresponding recognition results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of intelligent terminal technologies, and in particular, to an abnormal behavior detection method, apparatus, electronic device, and storage medium. Background Art

[0002] Abnormal detection in videos is an important issue in the field of computer vision. Currently, abnormal detection in videos has a wide range of applications in aspects such as public safety, such as detecting illegal behaviors, traffic accidents, and some uncommon events, etc. Summary of the Invention

[0003] Embodiments of the present disclosure provide an abnormal behavior detection method, apparatus, electronic device, and storage medium, which can adaptively detect single-object abnormal behaviors and multi-object abnormal behaviors simultaneously, improving the intelligence of abnormal behavior detection.

[0004] The technical solution of the embodiments of the present disclosure is implemented as follows:

[0005] Embodiments of the present disclosure provide an abnormal behavior detection method, including:

[0006] Detect and track the current video frame in the video sequence to determine the detection box area of each target object present;

[0007] Determine the overlap degree between different detection box areas, classify each detection box area according to the overlap degree, and determine the first single-object area and the first multi-object area corresponding to the current video frame;

[0008] Determine the single-object trajectory sequence to be recognized and the multi-object trajectory sequence to be recognized according to the second single-object area, the second multi-object area, the first single-object area, and the first multi-object area corresponding to the historical video frames in the video sequence;

[0009] Perform abnormal behavior recognition on the single-object trajectory sequence to be recognized and the multi-object trajectory sequence to be recognized, and obtain the corresponding recognition results.

[0010] Embodiments of the present disclosure provide an abnormal behavior detection apparatus, including:

[0011] A determination unit is configured to detect and track a current video frame in a video sequence, determine a detection box area of each existing target object; determine an overlap degree between different detection box areas, classify each detection box area according to the overlap degree, and determine a first single-object area and a first multi-object area corresponding to the current video frame; determine a single-object trajectory sequence to be recognized and a multi-object trajectory sequence to be recognized according to a second single-object area, a second multi-object area, the first single-object area, and the second multi-object area corresponding to historical video frames in the video sequence;

[0012] An identification unit is configured to perform abnormal behavior identification on both the single-object trajectory sequence to be recognized and the multi-object trajectory sequence to be recognized, and obtain corresponding identification results.

[0013] An embodiment of the present disclosure provides an electronic device, including: a memory for storing an executable computer program; a processor for implementing the above abnormal behavior detection method when executing the executable computer program stored in the memory.

[0014] An embodiment of the present disclosure provides a computer-readable storage medium storing a computer program for causing a processor to implement the above abnormal behavior detection method when executed.

[0015] The technical solution provided by the embodiment of the present disclosure has the following beneficial technical effects:

[0016] Since in the process of detecting abnormal behaviors in a video, the detection box areas in each video frame in the video sequence can be classified, and according to the classification of the detection box areas in each video frame, a first single-object area and a first multi-object area corresponding to the current video frame are determined. According to the second single-object area and the second multi-object area corresponding to historical video frames in the video sequence, and the first single-object area and the first multi-object area in the current video frame, a single-object trajectory sequence to be recognized and a multi-object trajectory sequence to be recognized are determined, and abnormal behavior identification is performed on both the single-object trajectory sequence to be recognized and the multi-object trajectory sequence to be recognized, so as to obtain the behavior identification result of the single-object trajectory sequence to be recognized and the behavior identification result of the multi-object trajectory sequence to be recognized; therefore, it is possible to adaptively detect single-object abnormal behaviors and multi-object abnormal behaviors simultaneously in the process of detecting abnormal behaviors in a video, thereby improving the intelligence of detecting abnormal behaviors in a video.

[0017] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings here are incorporated into the specification and form a part of this specification. These drawings show embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0019] Figure 1 It is an alternative flowchart diagram of the abnormal behavior detection method provided for the embodiments of the present disclosure;

[0020] Figure 2 It is an alternative flowchart diagram of the abnormal behavior detection method provided for the embodiments of the present disclosure;

[0021] Figure 3 It is a schematic diagram of an exemplary video frame including a single object box region and multiple object box regions provided for the embodiments of the present disclosure;

[0022] Figure 4 It is an alternative flowchart diagram of the abnormal behavior detection method provided for the embodiments of the present disclosure;

[0023] Figure 5 It is an alternative flowchart diagram of the abnormal behavior detection method provided for the embodiments of the present disclosure;

[0024] Figure 6 It is an alternative flowchart diagram of the abnormal behavior detection method provided for the embodiments of the present disclosure;

[0025] Figure 7 It is an alternative flowchart diagram of the abnormal behavior detection method provided for the embodiments of the present disclosure;

[0026] Figure 8 It is an alternative flowchart diagram of the abnormal behavior detection method provided for the embodiments of the present disclosure;

[0027] Figure 9 It is a schematic diagram of the minimum coordinate union of the regional coordinates of the exemplary first candidate region and second candidate region, corresponding to the region in the video frame, provided for the embodiments of the present disclosure;

[0028] Figure 10 It is an alternative flowchart diagram of the abnormal behavior detection method provided for the embodiments of the present disclosure;

[0029] Figure 11 It is an alternative flowchart diagram of the exemplary abnormal classification and single object abnormal behavior recognition for a candidate trajectory sequence provided for the embodiments of the present disclosure;

[0030] Figure 12 It is an alternative flowchart diagram of the exemplary multiple object abnormal behavior recognition when based on 3 video frames provided for the embodiments of the present disclosure;

[0031] Figure 13Schematic structural diagram of the abnormal behavior detection device provided by an embodiment of the present disclosure;

[0032] Figure 14 Schematic structural diagram of the electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0033] In order to make the objectives, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limitations on the present disclosure. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present disclosure.

[0034] Currently, when detecting abnormal behaviors in a video, according to the difference in the number of behavior executors, it can be roughly divided into single-person behaviors and multi-person behaviors. Single-person behaviors usually perform action recognition based on the generated pedestrian trajectories, and multi-person behaviors usually first locate the image area and then perform recognition based on the area sequence. In related technologies, when detecting abnormal behaviors according to a video, single-person behavior detection and multi-person behavior detection usually cannot be performed simultaneously, and moreover, whether to perform single-person behavior detection or multi-person behavior detection needs to be preset in advance. Therefore, in related technologies, it is impossible to adaptively detect abnormal behaviors of single-person and multi-person behaviors simultaneously, resulting in low intelligence in pedestrian behavior detection.

[0035] Based on this, an embodiment of the present disclosure provides an abnormal behavior detection method, which can adaptively detect single-object abnormal behaviors and multi-object abnormal behaviors simultaneously, improving the intelligence of abnormal behavior detection. The abnormal behavior detection method provided by the embodiment of the present disclosure is applied to an electronic device. The electronic device provided by the embodiment of the present disclosure can be implemented as various types of user terminals (hereinafter referred to as terminals) such as AR (Augmented Reality) glasses, laptop computers, tablet computers, desktop computers, set-top boxes, and mobile devices (such as mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), or can also be implemented as a server.

[0036] Figure 1 is an optional flowchart of the abnormal behavior detection method provided by an embodiment of the present disclosure, and will be described in conjunction with Figure 1 the steps shown.

[0037] S101. Detect and track the current video frame in the video sequence, and determine the detection box area of each existing target object.

[0038] In the embodiments of the present disclosure, an electronic device may obtain a video sequence from an image acquisition device. Moreover, for each video frame obtained from the video sequence, the electronic device may perform object detection and tracking on the currently obtained video frame (the current video frame), and in the case where there is an object of interest in the current video frame, obtain the detection box regions of each object of interest present in the current video frame.

[0039] In some embodiments, the image acquisition device may be a device inside the electronic device or an external device communicatively connected to the electronic device. The embodiments of the present disclosure do not limit this.

[0040] In some embodiments, the electronic device may sample the video sequence sent in real time by the image acquisition device to obtain a sampled video sequence, and perform object detection and tracking on each video frame in the sampled video sequence. For example, the electronic device may sample 8 video frames from the video frame sequence sent in real time by the image acquisition device within every 3 seconds, and for each sampled video frame, perform object detection and tracking on the video frame; in this way, the computational load of the electronic device can be reduced, thereby improving the efficiency during abnormal behavior detection. In some embodiments, the electronic device may also perform object detection and tracking on each video frame in the video frame sequence sent in real time by the image acquisition device. The embodiments of the present disclosure do not limit this.

[0041] Here, the object of interest may be a pedestrian, an animal, or the like. The embodiments of the present disclosure do not limit this; while the electronic device performs detection and tracking on the current video frame in the video sequence and determines the detection box regions of each object of interest present, it also determines the object identifier of each detection box region. The object identifier represents the object of interest corresponding to the detection box region, and the objects of interest corresponding to two detection box regions with different object identifiers are different.

[0042] Here, a pre-trained object detection network and object tracking network may be used to perform object detection and tracking on the video frame; the embodiments of the present disclosure do not limit the object detection network and the object tracking network here.

[0043] In the embodiments of the present disclosure, for the current video frame, in the case where there is one or more objects of interest in the video frame, the electronic device may perform object detection and tracking on the video frame to obtain the detection box regions of each object of interest and the object identifiers of the detection box regions; while in the case where there is no object of interest in the video frame, the electronic device cannot obtain the detection box regions of any object of interest and the object identifiers of the detection box regions.

[0044] S102. Determine the overlap degree between different detection box regions, classify each detection box region according to the overlap degree, and determine the first single-object region and the first multi-object region corresponding to the current video frame.

[0045] In the embodiments of the present disclosure, for the convenience of distinction, a single object region in the current video frame is referred to as a first single object region, a single object region in the historical video frame is referred to as a second single object region, and a plurality of object regions in the current video frame are referred to as a first plurality of object regions, and a plurality of object regions in the historical video frame are referred to as a second plurality of object regions.

[0046] In the embodiments of the present disclosure, for the current video frame, the electronic device may determine the degree of region overlap between different detection box regions, classify each detection box region according to the obtained degree of region overlap, and determine the first single object region and the first plurality of object regions corresponding to the current video frame according to the classification of the detection box regions.

[0047] Here, the first single object region is an image region in the current video frame that contains a single target object; the first plurality of object regions is an image region in the current video frame that contains two or more target objects.

[0048] Here, by determining the degree of region overlap between different detection box regions, the electronic device can obtain at least one degree of overlap corresponding to each detection box region, and classify the detection box region according to at least one degree of overlap corresponding to each detection box region.

[0049] In some embodiments, the electronic device may determine the intersection-over-union ratio of the areas between any two different detection box regions to obtain at least one degree of overlap corresponding to each detection box region. That is, for any one detection box region, the electronic device may determine the intersection-over-union ratio of the areas between the detection box region and each remaining detection box region, and use the intersection-over-union ratio of the areas between the detection box region and each remaining detection box region as at least one degree of overlap corresponding to the detection box region; where the remaining detection box regions are the detection box regions in the current video frame other than the detection box region. For example, when there are 3 detection box regions in the current video frame: detection box region 1, detection box region 2, and detection box region 3, for detection box region 1, the intersection-over-union ratio of the areas between detection box region 1 and detection box region 2, and between detection box region 1 and detection box region 3 can be calculated to obtain 2 region overlap degrees corresponding to detection box region 1; for detection box region 2, the intersection-over-union ratio of the areas between detection box region 2 and detection box region 1, and between detection box region 2 and detection box region 3 can be calculated to obtain 2 region overlap degrees corresponding to detection box region 2; for detection box region 3, the intersection-over-union ratio of the areas between detection box region 3 and detection box region 1, and between detection box region 3 and detection box region 2 can be calculated to obtain 2 region overlap degrees corresponding to detection box region 3.

[0050] S103. Determine the single-object trajectory sequence to be recognized and the multi-object trajectory sequence to be recognized based on the second single-object region, the second multi-object region, the first single-object region, and the first multi-object region corresponding to the historical video frames in the video sequence.

[0051] In the embodiments of the present disclosure, the electronic device may determine the single-object trajectory sequence to be recognized and the multi-object trajectory sequence to be recognized based on the first single-object region and the first multi-object region corresponding to the current video frame in the obtained video sequence, and the second single-object region and the second multi-object region corresponding to the historical video frames in the obtained video sequence.

[0052] Here, the electronic device may determine the single-object trajectory sequence to be recognized for each different target object and the multi-object trajectory sequence to be recognized. For example, when the target object is a pedestrian, the single-person trajectory sequences for different pedestrians and the same multi-person trajectory sequence corresponding to multiple pedestrians may be determined.

[0053] In some embodiments, the electronic device may, when a preset condition is met, determine the single-object trajectory sequence to be recognized and the multi-object trajectory sequence to be recognized based on the second single-object region and the second multi-object region corresponding to the historical video frames in the obtained video sequence, and the first single-object region and the first multi-object region corresponding to the current video frame. Exemplarily, the preset condition may include:

[0054] (1) No target object is detected in any of the consecutive X video frames after the current video frame.

[0055] Here, X is an integer greater than 0. For example, it may be 2 or 3, etc., and the embodiments of the present disclosure do not limit this.

[0056] (2) The sum of the video frame numbers of the current video frame and the historical video frame is greater than or equal to a preset number of frames.

[0057] For example, when the electronic device continuously obtains 8 video frames, it determines the single-object trajectory sequence to be recognized and the multi-object trajectory sequence to be recognized corresponding to these 8 video frames based on the first single-object region, the first multi-object region, the second single-object region, and the second multi-object region corresponding to these 8 video frames. And when the current video frame is the 8th video frame obtained, it can determine the single-object trajectory sequence to be recognized and the multi-object trajectory sequence to be recognized based on the 8th video frame and the 7 video frames before the 8th video frame. After that, it continues to obtain the next 8 video frames and continues to determine the single-object trajectory sequence to be recognized and the multi-object trajectory sequence to be recognized corresponding to the next 8 video frames according to the next 8 video frames. This cycle continues until the detection stops.

[0058] S104. Perform abnormal behavior recognition on both the single-object trajectory sequence to be recognized and the multi-object trajectory sequence to be recognized, and obtain the corresponding recognition results.

[0059] In the embodiments of the present disclosure, the electronic device may use a classification network (classifier) to perform abnormal behavior recognition on both the single-object trajectory sequence to be recognized and the multi-object trajectory sequence to be recognized, so as to obtain the recognition result for characterizing whether there is an abnormal behavior in the single-object trajectory sequence to be recognized, and the recognition result for characterizing whether there is an abnormal behavior in the multi-object trajectory sequence to be recognized.

[0060] In some embodiments, as Figure 2 shown, the above S102 may be implemented through S1021 to S1023, and the steps will be described with Figure 2 illustrations.

[0061] S1021. Determine the overlap degree between different detection frame regions, classify each detection frame region according to the overlap degree, and obtain the classification result.

[0062] Here, the electronic device may classify each obtained detection frame region according to the region overlap degree of the detection frame region to determine whether the detection frame region is a single-object frame region or a multi-object frame region, and obtain the classification result.

[0063] In some embodiments, for any detection box region, when there is an overlap degree greater than a preset threshold among at least one overlap degree corresponding to the any detection box region, a classification result indicating that the any detection box region is a plurality of object box regions is obtained; when at least one overlap degree corresponding to the any detection box region is less than or equal to the preset threshold, a classification result indicating that the any detection box region is a single object box region is obtained; thus, by classifying each detection box region in the current video frame, all single object box regions and all multiple object box regions in the current video frame can be obtained.

[0064] Here, the preset threshold can be set according to actual needs. For example, it can be 0 or 0.5, etc. The embodiments of the present disclosure do not limit this.

[0065] Here, a single object box region is a detection box region in the current video frame whose overlap degree with the remaining detection box regions is less than or equal to the preset threshold; a multiple object box region is a detection box region in the current video frame whose overlap degree with the remaining detection box regions is greater than the preset threshold. For example, when the preset threshold is 0, a single object box region is a detection box region in the current video frame that does not overlap with the remaining detection box regions; a multiple object box region is a detection box region in the current video frame that overlaps with the remaining detection box regions. Exemplarily, as Figure 3 shown, among them, detection box region 2-1 and detection box region 2-2 are two different single object box regions, and detection box region 2-3 and detection box region 2-4 are two different multiple object box regions.

[0066] S1022. When the classification result indicates that each detection box region is a single object box region, determine that each detection box region is a first single object region.

[0067] Here, for each detection box region, when the detection box region is a single object box region, the electronic device directly uses the detection box region as a first single object region in the current video frame.

[0068] S1023. When the classification result indicates that each detection box region is a multiple object box region, expand the multiple object box regions by a preset ratio to obtain a first multiple object region; the range of the first multiple object region is larger than the range of the corresponding multiple object box regions.

[0069] Here, for each detection box region, when the detection box region is a multiple object box region, the electronic device expands the detection box region by a preset ratio, and uses the region obtained after expansion and whose range is larger than the range of the detection box region before expansion as a first multiple object region in the current video frame.

[0070] Here, the preset ratio can be set according to actual needs. For example, the preset ratio can be 2 or 3, etc., and the embodiments of the present disclosure do not limit this. Exemplarily, the electronic device may expand the multiple object frame regions 2-4 in the above Figure 3 by 2 times, and the obtained first multiple object regions will be regions that include the multiple object frame regions 2-4 and the multiple object frame regions 2-3.

[0071] In some embodiments, as Figure 4 shown, the above S103 may be implemented through S1031 to S1032. Taking Figure 1 the S103 in which can be implemented through S1031 to S1032 as an example, the steps shown in Figure 4 will be described.

[0072] S1031. Determine a single object trajectory sequence to be recognized according to the second single object region corresponding to the historical video frame and the first single object region.

[0073] In some embodiments, the electronic device may form a sequence of the second single object regions and the first single object regions corresponding to the same target object in the second single object region corresponding to the historical video frame and the first single object region corresponding to the current video frame, so as to obtain a single object trajectory sequence to be recognized. For example, continuing with the above 8 video frames as an example, in the case where the second single object regions corresponding to the target object xx correspond to the 1st, 2nd, 3rd, and 4th video frames, and the second single object regions corresponding to the target object yy correspond to the 4th, 5th, 6th, and 7th video frames, the electronic device may use the sequence composed of the 4 second single object regions corresponding to the 1st, 2nd, 3rd, and 4th video frames as a single object trajectory sequence to be recognized, and use the sequence composed of the 4 second single object regions corresponding to the 4th, 5th, 6th, and 7th video frames as another single object trajectory sequence to be recognized.

[0074] S1032. Determine a multiple object trajectory sequence to be recognized according to the second multiple object regions corresponding to the historical video frame and the first multiple object regions.

[0075] In some embodiments, the electronic device may form a sequence of the second single object regions and the first single object regions corresponding to the same multiple target objects in the second multiple object regions corresponding to the historical video frame and the first multiple object regions corresponding to the current video frame, so as to obtain a multiple object trajectory sequence to be recognized.

[0076] In the embodiments of the present disclosure, S1031 and S1032 may be executed simultaneously or sequentially, and the embodiments of the present disclosure do not limit this execution order.

[0077] In some embodiments, the above S1031 may be implemented through S201 - S202, and the steps shown in Figure 5 will be described.

[0078] S201. Determine a candidate trajectory sequence corresponding to each target object from the second single - object regions corresponding to historical video frames and the first single - object regions, so as to obtain at least one candidate trajectory sequence; each detection - box region corresponds to one target object one by one.

[0079] Here, the electronic device may select the second single - object regions and the first single - object regions belonging to the same target object from the second single - object regions corresponding to historical video frames and the first single - object regions, and form a sequence of the second single - object regions and the first single - object regions belonging to the same target object as the candidate trajectory sequence corresponding to the target object, so as to obtain at least one candidate trajectory sequence corresponding one by one to at least one different target object.

[0080] S202. Obtain a single - object trajectory sequence to be recognized by performing trajectory screening on each candidate trajectory sequence.

[0081] In the case of obtaining at least one candidate trajectory sequence, the electronic device may perform trajectory screening on each candidate trajectory sequence to screen out the single - object trajectory sequence to be recognized.

[0082] Here, the electronic device may screen out the abnormal candidate trajectory sequences from the obtained at least one candidate trajectory sequence as the single - object trajectory sequence to be recognized.

[0083] In some embodiments, the above S202 may be implemented in the following manner: perform binary classification on each candidate trajectory sequence to obtain the label of each candidate trajectory sequence; the label is used to characterize whether each candidate trajectory sequence is abnormal; according to the label, screen out the single - object trajectory sequence to be recognized from at least one candidate trajectory sequence.

[0084] In some embodiments, the electronic device may use a trajectory nomination network to perform binary classification on each candidate trajectory sequence. Here, the trajectory nomination network may be a video classification network. The electronic device may perform binary classification on whether each candidate trajectory sequence is abnormal through a pre - trained video classification network, obtain the label of each candidate trajectory sequence, and screen out the abnormal candidate trajectory sequences from at least one candidate trajectory sequence as the single - object trajectory sequence to be recognized. The video action classification network may be, for example, a Temporal Segment Networks (TSN) network, or a Tempora Shift Module (TSM) network, etc. The embodiments of the present disclosure do not limit this.

[0085] Here, the label can be a score value. For a candidate trajectory sequence, the electronic device can use a video classification network to determine the anomaly score value corresponding to the candidate trajectory sequence, and determine whether the anomaly score value is greater than or equal to a score threshold. When the anomaly score value is greater than or equal to the score threshold, the electronic device determines that the candidate trajectory sequence is an abnormal sequence and outputs the corresponding score value; when the anomaly score value is less than the score threshold, the electronic device determines that the candidate trajectory sequence is a normal sequence and outputs the corresponding score value. The score value can be, for example, "0" or "1", or it can also be other numerical values, etc. The embodiments of the present disclosure do not limit the specific numerical value of the score value. Exemplarily, when the score value is "0" or "1", and "0" represents normal and "1" represents abnormal, when the anomaly score value is greater than or equal to the score threshold, the electronic device determines that the candidate trajectory sequence is an abnormal sequence and outputs the score value "1"; when the anomaly score value is less than the score threshold, the electronic device determines that the candidate trajectory sequence is a normal sequence and outputs the score value "0".

[0086] Here, the score threshold can be obtained by training the video classification network. The electronic device can use a training set to train the initial video classification network by methods such as gradient descent, so as to finally obtain the score threshold; Exemplarily, the score threshold can be 0.5. The embodiments of the present disclosure do not limit the specific numerical value of the score threshold.

[0087] In the embodiments of the present disclosure, by classifying the candidate trajectory sequences and using the candidate trajectory sequences representing anomalies as single object trajectory sequences to be recognized, normal candidate trajectory sequences can be filtered out, so as to perform subsequent single object abnormal behavior recognition on the remaining abnormal candidate trajectory sequences, thereby reducing the computational amount when recognizing the trajectory sequences of a single object and improving the accuracy and efficiency of abnormal behavior recognition.

[0088] In some embodiments, the above S1032 can be implemented through S301 to S302, and will be described in combination with Figure 6 the steps shown.

[0089] S301. Perform region screening on the first multiple object regions corresponding to the current video frame to obtain a first candidate region, and perform region screening on the second multiple object regions corresponding to each historical video frame to obtain a second candidate region.

[0090] Here, the electronic device can perform region screening on at least one of the first multiple object regions corresponding to the current video frame to obtain the first candidate region corresponding to the current video frame, and perform region screening on at least one of the second multiple object regions corresponding to each historical video frame to obtain the second candidate region corresponding to the historical video frame; In this way, through region screening, it is beneficial to improve the accuracy of recognizing the abnormal behaviors of multiple objects.

[0091] In some embodiments, the electronic device may classify each of the first plurality of object regions and screen out first candidate regions according to the region classification results; and classify each of the second plurality of object regions and screen out second candidate regions according to the region classification results.

[0092] In some embodiments, the region screening of the first plurality of object regions corresponding to the current video frame in S301 above to obtain the first candidate regions may be implemented through S3011 to S3012:

[0093] S3011. Determine the outliers of each of the first plurality of object regions corresponding to the current video frame.

[0094] S3012. According to the outliers, screen out the first plurality of object regions with the highest degree of abnormality from the first plurality of object regions corresponding to the current video frame as the first candidate regions.

[0095] In some embodiments, the region screening of the second plurality of object regions corresponding to each historical video frame in S301 above to obtain the second candidate regions may be implemented through S3013 to S3014:

[0096] S3013. Determine the outliers of each of the second plurality of object regions corresponding to each historical video frame.

[0097] S3014. According to the outliers, screen out the second plurality of object regions with the highest degree of abnormality from the second plurality of object regions corresponding to each historical video frame as the second candidate regions.

[0098] Here, the electronic device may determine the outliers of each of the first plurality of object regions corresponding to the current video frame through a region nomination network, and determine the outliers of each of the second plurality of object regions corresponding to each historical video frame through the region nomination network. The region nomination network may be an image classification network, and the electronic device may use a pre-trained image classification network to determine the outliers of each of the first plurality of object regions or the second plurality of object regions. The image classification network may be, for example, a Visual Geometry Group (VGG) network, a Residual Network (ResNet), etc., and the embodiments of the present disclosure do not limit this.

[0099] Here, the outliers of each of the first plurality of object regions or the second plurality of object regions represent the degree of abnormality of the first plurality of object regions or the second plurality of object regions. Exemplarily, the outlier may be a value between 0 and 1, and the larger the value, the higher the degree of abnormality. The embodiments of the present disclosure do not limit the numerical range of the outliers.

[0100] The electronic device may, according to the outlier values corresponding to each of the first plurality of object regions or the second plurality of object regions, select, from the first plurality of object regions or the second plurality of object regions corresponding to a video frame, the first plurality of object regions or the second plurality of object regions with the highest degree of abnormality as the first candidate region or the second candidate region corresponding to the video frame. For example, when there are 3 first plurality of object regions in video frame Z and the outlier values corresponding to these 3 first plurality of object regions are different, one of the first plurality of object regions with the highest degree of abnormality may be selected from these 3 first plurality of object regions as the first candidate region in video frame Z; in this way, the obtained region is more accurate, so that the plurality of object trajectory sequences to be recognized determined according to the obtained region are more accurate.

[0101] S302. Obtain a plurality of object trajectory sequences to be recognized according to the first candidate region and the second candidate region.

[0102] In some embodiments, the electronic device may form a sequence of the first candidate region and the second candidate region belonging to the same target object and use this sequence as the plurality of object trajectory sequences to be recognized.

[0103] In the embodiments of the present disclosure, by performing region screening on the first plurality of object regions and the second plurality of object regions to obtain the first candidate region and the second candidate region, irrelevant first plurality of object regions and second plurality of object regions can be filtered out (that is, there is no interaction behavior between the target objects in the plurality of object regions, or the interaction behavior does not belong to an abnormal behavior), and subsequent abnormal behavior recognition is performed on the plurality of object trajectory sequences determined by the remaining first candidate region and second candidate region, so as to filter out the sequences composed of the irrelevant first plurality of object regions and second plurality of object regions, reducing the recognition of the sequences composed of irrelevant plurality of object regions when performing multiple-object abnormal behavior recognition, thereby reducing the computational amount during behavior recognition and improving the accuracy and efficiency during abnormal behavior recognition.

[0104] In some embodiments, the above S301 may be implemented through S3015 to S3018, and will be described in combination with Figure 7 the steps shown.

[0105] S3015. When there are a plurality of first plurality of object regions corresponding to the current video frame, perform non-maximum suppression processing on the plurality of first plurality of object regions to obtain at least one first standard region.

[0106] In some embodiments, when the electronic device determines that there are two or more first multiple object regions in the current video frame, it can perform Non-Maximum Suppression (NMS) processing on these two or more first multiple object regions to eliminate the redundant first multiple object regions in the current video frame and retain the best one or more first multiple object regions. Thus, the retained first multiple object regions are used as the first standard regions.

[0107] S3016. When each historical video frame corresponds to multiple second multiple object regions, perform non-maximum suppression processing on the multiple second multiple object regions to obtain at least one second standard region.

[0108] In some embodiments, when the electronic device determines that there are two or more second multiple object regions in each historical video frame, it can perform NMS processing on these two or more second multiple object regions to eliminate the redundant second multiple object regions in the historical video frame and retain the best one or more second multiple object regions. Thus, the retained second multiple object regions are used as the second standard regions.

[0109] S3017. Perform region screening on at least one first standard region to obtain a first candidate region.

[0110] S3018. Perform region screening on at least one second standard region to obtain a second candidate region.

[0111] When the first standard region corresponding to the current video frame is obtained, the electronic device can use the region nomination network in the above embodiments to perform region screening on each first standard region to obtain a first candidate region; and when the second standard region corresponding to each video frame is obtained, the electronic device can use the region nomination network in the above embodiments to perform region screening on each second standard region to obtain a second candidate region corresponding to each video frame.

[0112] In the embodiments of the present disclosure, by performing NMS processing on the first multiple object regions or the second multiple object regions corresponding to each video frame to obtain the first standard region or the second standard region, and using the first standard region and the second standard region to determine multiple object trajectory sequences to be recognized, the regions used to determine the multiple object trajectory sequences to be recognized can be made more accurate, thereby improving the accuracy of the obtained multiple object trajectory sequences to be recognized, which is beneficial to the correct recognition of abnormal behaviors of the multiple object trajectory sequences to be recognized in the subsequent process.

[0113] In some embodiments, the above S302 can also be implemented through S3021 to S3023, which will be Figure 8For example, an illustration will be given.

[0114] S3021. Determine the merging position information according to the position information of the first candidate region in the video frame where it is located and the position information of the second candidate region in the video frame where it is located; the region corresponding to the merging position information includes any one of the first candidate regions or any one of the second candidate regions.

[0115] The electronic device can determine a merging position information according to the position information of the first candidate region in the current video frame where it is located and the position information of the second candidate region in the historical video frame where it is located, and the region corresponding to the merging position information includes any one of the first candidate regions or any one of the second candidate regions.

[0116] Here, the position information can be region coordinates, and the merging position information can be the union of the region coordinates corresponding to the first candidate region or the second candidate region, for example, the smallest union. For example, when the region coordinates of one first candidate region H1 in the current video frame where it is located are {(x11, y11), (x12, y12)}, and the region coordinates of one second candidate region H2 in the historical video frame where it is located are {(x21, y21), (x22, y22)}, and x11 < x12 < x21 < x22, y21 < y22 < y11 < y12, the smallest coordinate union of the region coordinates corresponding to these two candidate regions is: {(x11, y21), (x22, y12)}; among them, the region corresponding to the smallest coordinate union in the current video frame where the first candidate region H1 is located includes the first candidate region H1, and the region corresponding to the smallest coordinate union in the historical video frame where the second candidate region H2 is located includes the second candidate region H2. For example, Figure 9 Shows the region 9-1 corresponding to the smallest coordinate union between the region coordinates {(x11, y11), (x12, y12)} of one first candidate region in the current video frame where it is located and the region coordinates {(x21, y21), (x22, y22)} of one second candidate region in the historical video frame where it is located, where the region 9-1 includes the first candidate region 9-2 and includes the second candidate region 9-3.

[0117] S3022. Use the regions corresponding to the merging position information in each of the historical video frame and the current video frame as the regions to be recognized.

[0118] S3023. Use the sequence composed of at least one region to be recognized as the multiple object trajectory sequences to be recognized.

[0119] For example, in the case where the historical video frame and the current video frame are 8 video frames, the electronic device may use the region corresponding to the merging position information in each of the 8 video frames as the region to be recognized, so as to obtain 8 regions to be recognized, and form a sequence of these 8 regions to be recognized according to the frame order of the corresponding 8 video frames, thereby obtaining multiple object trajectory sequences to be recognized.

[0120] In the embodiments of the present disclosure, according to the position information of the first candidate region and the second candidate region in the video frame where they are located, the merging position information is determined, and the region corresponding to the merging position information in each of the historical video frame and the current video frame is used as the region to be recognized, and the sequence composed of at least one region to be recognized is used as the multiple object trajectory sequences to be recognized, so that the effective range of each region constituting the multiple object trajectory sequences to be recognized is increased, and more effective content is included, thereby making the obtained multiple object trajectory sequences to be recognized more accurate, and finally improving the recognition accuracy when performing multiple object abnormal behavior recognition on the multiple object trajectory sequences to be recognized subsequently.

[0121] In some embodiments, as Figure 10 shown, the above S104 may be implemented through S1041 to S1042, and will be described by taking Figure 10 as an example.

[0122] S1041. Respectively use at least one different first classification network to perform abnormal behavior classification on a single object trajectory sequence to be recognized, and obtain corresponding first classification results; wherein, the at least one different first classification network is used to classify at least one different first abnormal behavior event.

[0123] For each single object trajectory sequence to be recognized, the electronic device may input the single object trajectory sequence to be recognized into a first classification network (first classifier), or input it into multiple different first classifiers respectively, and perform corresponding action classification, so as to realize the recognition of at least one different first abnormal behavior for each single object trajectory sequence to be recognized, wherein each first classifier may classify a first abnormal behavior event.

[0124] Here, the at least one different first abnormal behavior events classified by the at least one different first classifiers may be set according to actual needs. For example, they may be respectively used to classify falling behaviors, climbing behaviors, etc., and the embodiments of the present disclosure do not limit this.

[0125] In some embodiments, before using at least one different first classification network to perform different first abnormal behavior event recognition on each single object trajectory sequence to be recognized, the electronic device may first use a video feature extraction network to extract features from the single object trajectory sequence to be recognized, obtain the sequence features of the single object trajectory sequence to be recognized, and then use at least one different first classification network to perform recognition of at least one different first abnormal behavior event on the sequence features; in this way, the accuracy of abnormal behavior recognition can be improved. The video feature extraction network may be, for example, a pre-trained convolutional neural network (CNN), or other pre-trained networks for feature extraction, and the embodiments of the present disclosure do not limit this.

[0126] Figure 11 FIG. is a schematic flowchart of an exemplary process for abnormal classification and single object abnormal behavior recognition of a candidate trajectory sequence provided by an embodiment of the present disclosure. As Figure 11 shown, for a candidate trajectory sequence I, the electronic device first performs binary classification on whether the candidate trajectory sequence I is abnormal through a trajectory nomination network, and when the label is 0 (indicating normal), no processing is performed on the candidate trajectory sequence I, while when the label is 1 (indicating abnormal), the candidate trajectory sequence I is used as a single object trajectory sequence to be recognized. Then, a video feature extraction network is used to extract features from the single object trajectory sequence to be recognized to obtain sequence features. Finally, n different classifiers are used to perform recognition of n different abnormal behaviors such as falling behavior, railing climbing behavior, and waving for help behavior on the sequence features, so as to identify which abnormal behavior among the n different abnormal behaviors exists in the single object trajectory sequence to be recognized. For example, there is a railing climbing behavior.

[0127] S1042. Respectively use at least one different second classification network to perform abnormal behavior classification on multiple object trajectory sequences to be recognized, and obtain corresponding second classification results; wherein, the at least one different second classification network is used to classify at least one different second abnormal behavior event.

[0128] For each of the multiple object trajectory sequences to be recognized, the electronic device may input the multiple object trajectory sequences to be recognized into a second classification network (second classifier), or input them into multiple different second classifiers respectively, to perform corresponding action classification, so as to realize the recognition of at least one different second abnormal behavior for each of the multiple object trajectory sequences to be recognized, wherein each second classifier can classify a second abnormal behavior event.

[0129] Here, at least one different first abnormal behavior event to be classified by at least one different second classifier can be set according to actual needs. For example, it can be used to classify fighting behavior, gathering behavior, etc. The embodiments of the present disclosure do not limit this.

[0130] In the embodiments of the present disclosure, S1041 to S1042 above can be executed simultaneously or successively. The embodiments of the present disclosure do not limit this.

[0131] In some embodiments, the first abnormal behavior event can be a behavior event for a single object, and the second abnormal behavior event can be a behavior event for multiple objects.

[0132] In some embodiments, before the electronic device uses at least one different second classifier to identify different second abnormal behavior events for each of the multiple object trajectory sequences to be recognized, it can also first use a video feature extraction network to extract features from the multiple object trajectory sequences to be recognized, obtain the sequence features of the multiple object trajectory sequences to be recognized, and then use at least one different second classifier to identify at least one different second abnormal behavior event for the sequence features; in this way, the correctness of abnormal behavior recognition can be improved.

[0133] Figure 12 This is a schematic flowchart for recognizing abnormal behaviors of multiple objects according to 3 video frames provided by the embodiments of the present disclosure. As Figure 12 shown, for each of the 3 video frames (the first 2 video frames are historical video frames, and the 3rd video frame is the current video frame) containing detection frames, after the electronic device detects each detection frame area in the video frame, it classifies each detection frame area to obtain multiple object frame areas in the video frame, and then expands each multiple object frame area by 2 times to obtain a first multiple object area and a second multiple object area; uses a region nomination network to perform region classification on each first multiple object area and second multiple object area to obtain the abnormal value of each multiple object area, and filters out the first candidate area and the second candidate area according to the abnormal value. Then, according to the position information of all the first candidate areas and second candidate areas corresponding to these 3 video frames in the video frames where they are located, a combined position information is determined, and the area corresponding to the combined position information in each of these 3 video frames is used as the area to be recognized, so as to obtain 3 areas to be recognized. According to the frame order of these 3 video frames, these 3 areas to be recognized are formed into a sequence, and this sequence is used as the multiple object trajectory sequence to be recognized. Finally, an action recognition network (m classifiers) is used to recognize m different second abnormal behavior events for the multiple object trajectory sequence to be recognized, and the recognition result is obtained.

[0134] The present disclosure also provides an abnormal behavior detection device. Figure 13 It is a schematic structural diagram of the abnormal behavior detection device provided by the embodiment of the present disclosure; as Figure 13 shown, the abnormal behavior detection device 1 includes: a determination unit 10, configured to detect and track a current video frame in a video sequence, and determine a detection box area of each existing target object; determine an overlap degree between different detection box areas, classify each detection box area according to the overlap degree, and determine a first single object area and a first multiple object area corresponding to the current video frame; determine a single object trajectory sequence to be recognized and a multiple object trajectory sequence to be recognized according to a second single object area, a second multiple object area, the first single object area, and the second multiple object area corresponding to historical video frames in the video sequence; an identification unit 20, configured to perform abnormal behavior identification on the single object trajectory sequence to be recognized and the multiple object trajectory sequence to be recognized, and obtain corresponding recognition results.

[0135] In some embodiments of the present disclosure, the determination unit 10 is further configured to determine an area intersection-over-union ratio between any two different detection box areas, and obtain at least one of the overlap degrees corresponding to each detection box area.

[0136] In some embodiments of the present disclosure, the determination unit 10 is further configured to classify each detection box area according to the overlap degree to obtain a classification result; in the case where the classification result indicates that each detection box area is a single object box area, determine each detection box area as the first single object area; in the case where the classification result indicates that each detection box area is a multiple object box area, expand the multiple object box area by a preset ratio to obtain the first multiple object area; the range of the first multiple object area is larger than the range of the corresponding multiple object box area.

[0137] In some embodiments of the present disclosure, the determination unit 10 is further configured to, for any one detection box area, in the case where there is an overlap degree greater than a preset threshold among at least one of the overlap degrees, obtain the classification result indicating that the any one detection box area is a multiple object box area; in the case where at least one of the overlap degrees is less than or equal to the preset threshold, obtain the classification result indicating that the any one detection box area is a single object box area.

[0138] In some embodiments of the present disclosure, the determination unit 10 is further configured to determine the single object trajectory sequence to be recognized according to the second single object area corresponding to the historical video frame and the first single object area; determine the multiple object trajectory sequence to be recognized according to the second multiple object area corresponding to the historical video frame and the first multiple object area.

[0139] In some embodiments of the present disclosure, the determining unit 10 is further configured to determine, from the second single object region corresponding to the historical video frame and the first single object region, a candidate trajectory sequence corresponding to each target object, so as to obtain at least one candidate trajectory sequence; each detection box region corresponds to a target object one by one; by performing trajectory screening on each candidate trajectory sequence, the single object trajectory sequence to be recognized is obtained.

[0140] In some embodiments of the present disclosure, the determining unit 10 is further configured to perform binary classification on each candidate trajectory sequence to obtain a label of each candidate trajectory sequence; the label is used to characterize whether each candidate trajectory sequence is abnormal; according to the label, the single object trajectory sequence to be recognized is screened from the at least one candidate trajectory sequence.

[0141] In some embodiments of the present disclosure, the determining unit 10 is further configured to perform region screening on the first multiple object regions corresponding to the current video frame to obtain a first candidate region, and perform region screening on the second multiple object regions corresponding to each historical video frame to obtain a second candidate region; according to the first candidate region and the second candidate region, the multiple object trajectory sequences to be recognized are obtained.

[0142] In some embodiments of the present disclosure, the determining unit 10 is further configured to determine an outlier value of each of the first multiple object regions corresponding to the current video frame; according to the outlier value, the first multiple object regions with the highest degree of abnormality are screened out from the first multiple object regions corresponding to the current video frame as the first candidate region.

[0143] In some embodiments of the present disclosure, the determining unit 10 is further configured to determine combined position information according to the position information of the first candidate region in the video frame where it is located and the position information of the second candidate region in the video frame where it is located; the region corresponding to the combined position information includes any one of the first candidate regions or any one of the second candidate regions; the regions corresponding to the combined position information in each of the historical video frame and the current video frame are used as regions to be recognized; a sequence composed of at least one of the regions to be recognized is used as the multiple object trajectory sequences to be recognized.

[0144] In some embodiments of the present disclosure, the determining unit 10 is further configured to, when the current video frame corresponds to a plurality of first object regions, perform non-maximum suppression processing on the plurality of first object regions to obtain at least one first standard region; when each historical video frame corresponds to a plurality of second object regions, perform the non-maximum suppression processing on the plurality of second object regions to obtain at least one second standard region; perform region screening on the at least one first standard region to obtain the first candidate region; and perform region screening on the at least one second standard region to obtain the second candidate region.

[0145] In some embodiments of the present disclosure, the recognition result includes: a first classification result and a second classification result; the recognition unit 20 is further configured to respectively use at least one different first classification network to perform abnormal behavior classification on the single object trajectory sequence to be recognized to obtain a corresponding first classification result; wherein the at least one first classification network is used to classify at least one different first abnormal behavior event; respectively use at least one different second classification network to perform abnormal behavior classification on the plurality of object trajectory sequences to be recognized to obtain a corresponding second classification result; wherein the at least one second classification network is used to classify at least one different second abnormal behavior event.

[0146] Embodiments of the present disclosure further provide an electronic device. Figure 14 As shown in the structural schematic diagram of the electronic device provided by the embodiments of the present disclosure, Figure 14 it includes: a memory 22 and a processor 23, wherein the memory 22 and the processor 23 are connected through a communication bus 21; the memory 22 is used to store an executable computer program; when the processor 23 executes the executable computer program stored in the memory 22, it implements the abnormal behavior detection method provided by the embodiments of the present disclosure.

[0147] Embodiments of the present disclosure provide a computer-readable storage medium storing a computer program, which when executed by the processor 23, implements the abnormal behavior detection method provided by the embodiments of the present disclosure.

[0148] In some embodiments of the present disclosure, a storage medium may be a tangible device that can hold and store instructions used by an instruction execution device, and may be a volatile storage medium or a non-volatile storage medium. A computer-readable storage medium may be, for example (but not limited to), an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices, such as punch cards or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated through a waveguide or other transmission medium (e.g., optical pulses through an optical fiber cable), or electrical signals transmitted through wires; it may also be various devices including one or any combination of the foregoing memories; it may also be various devices including one or any combination of the foregoing memories.

[0149] In some embodiments of the present disclosure, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a stand-alone program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0150] As an example, executable instructions may or may not correspond to a file in a file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program under discussion, or, stored in multiple cooperating files (e.g., files storing one or more modules, subroutines, or portions of code).

[0151] As an example, executable instructions may be deployed to execute on one computing device, or on multiple computing devices located at one location, or, on multiple computing devices distributed at multiple locations and interconnected by a communication network.

[0152] In summary, by adopting the implementation solution of the present technology, it is possible to adaptively detect the abnormal behaviors of a single object and multiple objects simultaneously during the process of detecting abnormal behaviors in a video, thereby improving the intelligence of detecting abnormal behaviors in a video; and, it is possible to filter out the trajectory sequences of normal single objects and perform subsequent abnormal behavior recognition on the remaining abnormal trajectory sequences of single objects, thereby reducing the computational amount during the abnormal behavior recognition of the trajectory sequences of single objects and improving the accuracy and efficiency during the abnormal behavior recognition; and, by performing region screening on the first multiple object regions and the second multiple object regions to obtain the first candidate region and the second candidate region, it is possible to filter out irrelevant multiple object regions (that is, there is no interaction behavior between the target objects in the multiple object regions, or the interaction behavior does not belong to an abnormal behavior), and perform subsequent abnormal behavior recognition on the remaining first candidate region and the second candidate region, thereby filtering out the sequences composed of the irrelevant first multiple object regions and the second multiple object regions, so that when performing abnormal behavior recognition of multiple objects, the recognition of the sequences composed of the irrelevant first multiple object regions and the second multiple object regions is reduced, thereby reducing the computational amount during the behavior recognition and improving the accuracy and efficiency during the abnormal behavior recognition.

[0153] As described above, the foregoing is only an embodiment of the present disclosure and is not intended to limit the protection scope of the present disclosure. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present disclosure are included within the protection scope of the present disclosure.

Claims

1. An abnormal behavior detection method, characterized in that, Including: Detecting and tracking the current video frame in the video sequence to determine the detection box area of each existing target object; Determining the overlap degree between different detection box areas, classifying each detection box area according to the overlap degree, and determining the first single object area and the first multiple object areas corresponding to the current video frame; Determining the single object trajectory sequence to be recognized according to the second single object area corresponding to the historical video frame and the first single object area; Determining the multiple object trajectory sequences to be recognized according to the second multiple object areas corresponding to the historical video frames and the first multiple object areas; Performing abnormal behavior recognition on both the single object trajectory sequence to be recognized and the multiple object trajectory sequences to be recognized to obtain corresponding recognition results; The determining the single object trajectory sequence to be recognized according to the second single object area corresponding to the historical video frame and the first single object area includes: Determining a candidate trajectory sequence corresponding to each target object from the second single object area corresponding to the historical video frame and the first single object area, so as to obtain at least one candidate trajectory sequence; each detection box area corresponds to one target object one by one; Performing binary classification on each candidate trajectory sequence to obtain the label of each candidate trajectory sequence; the label is used to characterize whether each candidate trajectory sequence is abnormal; According to the label, screening the single object trajectory sequence to be recognized from the at least one candidate trajectory sequence.

2. The method according to claim 1, characterized in that, The determining the overlap degree between different detection box areas includes: Determining the intersection-over-union ratio of the areas between any two different detection box areas to obtain at least one of the overlap degrees corresponding to each detection box area.

3. The method according to claim 1 or 2, characterized in that, The classifying each detection box area according to the overlap degree and determining the first single object area and the first multiple object areas corresponding to the current video frame includes: Classifying each detection box area according to the overlap degree to obtain a classification result; When the classification result indicates that each detection box area is a single object box area, determining each detection box area as the first single object area; When the classification result indicates that each detection box area is a multiple object box area, expanding the multiple object box areas by a preset ratio to obtain the first multiple object areas; the range of the first multiple object areas is larger than the range of the corresponding multiple object box areas.

4. The method according to claim 3, characterized in that, The classifying each detection box area according to the overlap degree to obtain a classification result includes: For any one detection box area, when there is an overlap degree greater than a preset threshold among at least one of the overlap degrees, obtaining the classification result indicating that the any one detection box area is a multiple object box area; When all of the at least one overlap degree is less than or equal to the preset threshold, obtaining the classification result indicating that the any one detection box area is a single object box area.

5. The method according to claim 1, wherein The determining the multiple object trajectory sequences to be recognized according to the second multiple object areas corresponding to the historical video frames and the first multiple object areas includes: Perform region screening on the first multiple object regions corresponding to the current video frame to obtain first candidate regions, and perform region screening on the second multiple object regions corresponding to each historical video frame to obtain second candidate regions; Obtain the multiple object trajectory sequences to be recognized according to the first candidate regions and the second candidate regions.

6. The method according to claim 5, characterized in that, The performing region screening on the first multiple object regions corresponding to the current video frame to obtain first candidate regions includes: Determine the outliers of each of the first multiple object regions corresponding to the current video frame; According to the outliers, screen out the first multiple object regions with the highest degree of abnormality from the first multiple object regions corresponding to the current video frame as the first candidate regions.

7. The method according to claim 5, characterized in that The obtaining the multiple object trajectory sequences to be recognized according to the first candidate regions and the second candidate regions includes: Determine combined position information according to the position information of the first candidate regions in the video frames where they are located and the position information of the second candidate regions in the video frames where they are located; the region corresponding to the combined position information contains any one of the first candidate regions or any one of the second candidate regions; Use the regions corresponding to the combined position information in each of the historical video frames and the current video frame as regions to be recognized; Use the sequence composed of at least one of the regions to be recognized as the multiple object trajectory sequences to be recognized.

8. The method according to claim 6, characterized in that The performing region screening on the first multiple object regions corresponding to the current video frame to obtain first candidate regions, and performing region screening on the second multiple object regions corresponding to each historical video frame to obtain second candidate regions includes: When there are multiple first multiple object regions corresponding to the current video frame, perform non-maximum suppression processing on the multiple first multiple object regions to obtain at least one first standard region; When there are multiple second multiple object regions corresponding to each historical video frame, perform the non-maximum suppression processing on the multiple second multiple object regions to obtain at least one second standard region; Perform region screening on the at least one first standard region to obtain the first candidate regions; Perform region screening on the at least one second standard region to obtain the second candidate regions.

9. The method according to any one of claims 1-2, 4-8, characterized in that, The recognition results include: a first classification result and a second classification result; the performing abnormal behavior recognition on the single object trajectory sequence to be recognized and the multiple object trajectory sequences to be recognized to obtain corresponding recognition results includes: Respectively use at least one first classification network to perform abnormal behavior classification on the single object trajectory sequence to be recognized to obtain the corresponding first classification result; the at least one first classification network is used to classify at least one different first abnormal behavior event; Respectively use at least one second classification network to perform abnormal behavior classification on the multiple object trajectory sequences to be recognized to obtain the corresponding second classification result; the at least one second classification network is used to classify at least one different second abnormal behavior event.

10. An abnormal behavior detection device, characterized in that, Includes: A determining unit, configured to detect and track a current video frame in a video sequence, and determine a detection box area of each existing target object; Determine the overlap degree between different detection box areas, classify each detection box area according to the overlap degree, and determine a first single-object area and a first multi-object area corresponding to the current video frame; determine a single-object trajectory sequence to be recognized according to a second single-object area corresponding to a historical video frame and the first single-object area; determine a multi-object trajectory sequence to be recognized according to a second multi-object area corresponding to a historical video frame and the first multi-object area; An identifying unit, configured to perform abnormal behavior identification on both the single-object trajectory sequence to be recognized and the multi-object trajectory sequence to be recognized, and obtain corresponding recognition results; The determining unit is further configured to determine a candidate trajectory sequence corresponding to each target object from the second single-object area corresponding to the historical video frame and the first single-object area, so as to obtain at least one candidate trajectory sequence; each detection box area corresponds to one target object one by one; perform binary classification on each candidate trajectory sequence to obtain a label of each candidate trajectory sequence; The label is used to characterize whether each candidate trajectory sequence is abnormal; According to the label, screen the single-object trajectory sequence to be recognized from the at least one candidate trajectory sequence.

11. An electronic device, characterized in that, Comprising: A memory, configured to store an executable computer program; A processor, configured to implement the method according to any one of claims 1 to 9 when executing the executable computer program stored in the memory.

12. A computer-readable storage medium, characterized in that, A computer program is stored, which is used to cause a processor to implement the method according to any one of claims 1 to 9 when executed.

Citation Information

Patent Citations

  • Behavior recognition method and device, equipment and storage medium

    CN113111838A

  • Method, program, and system for determining whether abnormal behavior occurs, on basis of behavior sequence

    WO2021100919A1