Group abnormal event detection method, computer program product, equipment and medium
By constructing a group relationship feature map and using preset graphs to predict the network to calculate the peak signal-to-noise ratio, the problem of inaccurate detection in the existing group abnormal event detection methods is solved, and accurate judgment and early warning of group abnormal events is achieved.
Patent Information
- Application Number
- CN202510821463.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
Most existing group abnormal event detection methods mostly have the problem of inaccurate detection and fail to fully consider the characteristics of group relationships.
Based on the individual object detection information of each frame monitoring screen, an actual group relationship feature map is constructed, and the predicted group relationship feature map of future frames is determined through the preset graph prediction network, and the peak signal-to-noise ratio is calculated to generate the target anomaly score to judge group anomaly events.
The accuracy of group abnormal events detection is improved, and effective early warning treatment of group abnormal events is achieved, thereby avoiding the inaccurate detection caused by traditional methods due to insufficient consideration of group relationship characteristics.
Smart Images

Figure CN120339922A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a method for detecting group abnormal events, a computer program product, a device, and a medium. Background Art
[0002] The task of video abnormal event detection can be roughly divided into group abnormal event detection and individual abnormal event detection. Most of the existing methods for detecting group abnormal events extract and model features using the apparent information and motion information of video frames, and then judge whether an abnormal event occurs by calculating the abnormal score. However, most of the existing abnormal detection methods have the problem of inaccurate detection, which urgently needs to be solved by those skilled in the art. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a method for detecting group abnormal events, a computer program product, a device, and a medium, so as to improve the accuracy of detecting group abnormal events. The specific solutions are as follows: In a first aspect, the present application discloses a method for detecting group abnormal events, including: Based on the detection information of individual targets in each frame of the monitoring video, construct the actual group relationship feature map corresponding to the frame, and determine the sequence of group relationship feature maps based on the actual group relationship feature maps of each frame; Input the sequence of group relationship feature maps into a preset graph prediction network, and determine the predicted group relationship feature map of the future frame through the preset graph prediction network; Determine the peak signal-to-noise ratio between the predicted group relationship feature map of the future frame and the actual group relationship feature map of the future frame, and generate a target abnormal score according to the peak signal-to-noise ratio; Judge whether a group abnormal event occurs according to the target abnormal score, so as to perform early warning processing when a group abnormal event occurs.
[0004] Optionally, constructing the actual group relationship feature map based on the detection information of individual targets in each frame of the monitoring video includes: Identify the individual targets in each frame of the monitoring video, and obtain the detection information of the individual targets; Use the individual targets as nodes in the actual group relationship feature map, and use the detection information of the individual targets as the node attributes of the nodes to obtain a node set; Obtain an edge set based on the interaction intensity between any two individual targets; Construct the actual group relationship feature map based on the node set and the edge set.
[0005] Optionally, the detection information includes the confidence, position information, apparent features, and motion features of the individual targets.
[0006] Optionally, identify individual targets in each frame of the monitoring video and obtain detection information of the individual targets, including: Obtain the monitoring video from the surveillance video; Perform multi-scale scaling processing on the monitoring video to generate input images of multiple scales, and input the input images of multiple scales into a multi-scale object detection model in parallel; Identify individual targets in the monitoring video through the multi-scale object detection model, and obtain the confidence and position information of the individual targets.
[0007] Optionally, identify individual targets in each frame of the monitoring video and obtain detection information of the individual targets, including: Extract the corresponding regional image from the monitoring video based on the position information of the individual target; Preprocess the regional image and input the preprocessed regional image into a pre-trained convolutional neural network model to extract the apparent features of the individual target through the convolutional neural network model; where the apparent features represent the visual appearance features of the individual target.
[0008] Optionally, identify individual targets in each frame of the monitoring video and obtain detection information of the individual targets, including: Extract the corresponding regional image from the monitoring video based on the position information of the individual target; Extract the optical flow information at the regional image through an optical flow estimation network, and determine the motion characteristics of the individual target according to the optical flow information at the regional image.
[0009] Optionally, the method for detecting group abnormal events further includes: Based on the position information of any two individual targets, determine the spatial distance between any two individual targets, and calculate the edge weight attribute between any two individual targets according to the spatial distance; where the edge weight attribute represents the interaction intensity between any two individual targets.
[0010] Optionally, based on the interaction intensity between any two individual targets, obtain an edge set, including: Use the edge weight attribute between any two individual targets as the weight value between any two individual targets, and construct a target matrix according to the weight value between any two individual targets; Determine the edge set based on the target matrix.
[0011] Optionally, determine whether a group abnormal event occurs according to the target anomaly score, including: Judge whether the target anomaly score exceeds a preset score threshold; If the target anomaly score exceeds the preset score threshold, it is determined that a group abnormal event has occurred.
[0012] Optionally, the preset graph prediction network includes a temporal convolutional block and a spatial convolutional block; among them, the temporal convolutional block processes the temporal dimension features of the sequence of group relationship feature graphs, and the spatial convolutional block processes the node topology features of the actual group relationship feature graph of a single frame in the sequence of group relationship feature graphs.
[0013] Optionally, determining the peak signal-to-noise ratio between the predicted group relationship feature graph of the future frame and the actual group relationship feature graph of the future frame, and generating a target anomaly score according to the peak signal-to-noise ratio, includes: Determining the target peak signal-to-noise ratio between the predicted group relationship feature graph of the future frame and the actual group relationship feature graph of the future frame; Generating a target anomaly score according to the target peak signal-to-noise ratio, the minimum value of the historical peak signal-to-noise ratio, and the maximum value of the historical peak signal-to-noise ratio.
[0014] Optionally, generating a target anomaly score according to the target peak signal-to-noise ratio, the minimum value of the historical peak signal-to-noise ratio, and the maximum value of the historical peak signal-to-noise ratio, includes: Calculating a first difference between the target peak signal-to-noise ratio and the minimum value of the historical peak signal-to-noise ratio; Calculating a second difference between the maximum value of the historical peak signal-to-noise ratio and the minimum value of the historical peak signal-to-noise ratio; Normalizing the ratio of the first difference to the second difference to obtain the target anomaly score.
[0015] In a second aspect, the present application discloses a computer program product, and when the computer program is executed by a processor, the steps of the group anomaly event detection method disclosed above are implemented.
[0016] In a third aspect, the present application discloses an electronic device, including: A memory for storing a computer program; A processor for executing the computer program to implement the group anomaly event detection method disclosed above.
[0017] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the group anomaly event detection method disclosed above is implemented.
[0018] It can be seen that the present application proposes a method for detecting group abnormal events, including: constructing an actual group relationship feature map corresponding to each frame based on the detection information of individual targets in each frame of the monitoring video, and determining a sequence of group relationship feature maps based on the actual group relationship feature maps of each frame; inputting the sequence of group relationship feature maps into a preset graph prediction network, and determining a predicted group relationship feature map of a future frame through the preset graph prediction network; determining the peak signal-to-noise ratio between the predicted group relationship feature map of the future frame and the actual group relationship feature map of the future frame, and generating a target abnormal score according to the peak signal-to-noise ratio; judging whether a group abnormal event occurs according to the target abnormal score, so as to perform early warning processing when a group abnormal event occurs.
[0019] Beneficial effects: Based on the detection information of individual targets in each frame of the monitoring video, the present application constructs an actual group relationship feature map corresponding to each frame, and determines a sequence of group relationship feature maps based on the actual group relationship feature maps of each frame. In this way, the group relationship can be quantitatively characterized in the form of a graph structure. Compared with the traditional method of detecting abnormal events only based on apparent information and motion information, the present application captures the key feature of group relationship more comprehensively, provides richer feature inputs for abnormal detection, and thus improves the accuracy of abnormal detection. Further, the present application inputs the sequence of group relationship feature maps into a preset graph prediction network, and thus determines a predicted group relationship feature map of a future frame through the preset graph prediction network, realizing the modeling and prediction of group behavior, and generating a target abnormal score by calculating the peak signal-to-noise ratio between the predicted group relationship feature map of the future frame and the actual group relationship feature map of the future frame. In this way, the difference degree between the predicted group relationship feature map and the actual group relationship feature map can be quantified, so as to accurately judge whether a group abnormal event occurs, avoiding the problem of inaccurate detection caused by the traditional method not fully considering the group relationship feature, and at the same time realizing the effective early warning processing of group abnormal events. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0021] Figure 1 It is a flowchart of a method for detecting group abnormal events disclosed in the present application; Figure 2 It is a flowchart of constructing an actual group relationship feature map disclosed in the present application; Figure 3 It is a schematic diagram of a graph prediction encoding and decoding process disclosed in the present application; Figure 4 Schematic diagram of a spatio-temporal convolution module disclosed in the present application; Figure 5 Flow chart of a specific method for detecting group abnormal events disclosed in the present application; Figure 6 Schematic diagram of the structure of a device for detecting group abnormal events disclosed in the present application; Figure 7 Schematic diagram of the structure of an electronic device disclosed in the present application. Detailed implementation manners
[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] The video abnormal event detection task can be roughly divided into group abnormal event detection and individual abnormal event detection. Most of the existing group abnormal event detection methods use the apparent information and motion information of video frames for feature extraction and modeling, and then judge whether an abnormal event occurs by calculating the abnormal score. However, most of the existing abnormal detection methods have the problem of inaccurate detection, which urgently needs to be solved by those skilled in the art.
[0024] Therefore, the embodiments of the present application propose a group abnormal event detection scheme, which can improve the accuracy of group abnormal event detection.
[0025] The embodiments of the present application disclose a method for detecting group abnormal events. Refer to Figure 1 , the method includes: Step S11: Based on the detection information of individual targets in each frame of the monitoring screen, construct the actual group relationship feature map corresponding to the frame, and determine the sequence of group relationship feature maps based on the actual group relationship feature maps of each frame.
[0026] In this embodiment, identify the individual targets in each frame of the monitoring screen and obtain the detection information of the individual targets; use the individual targets as the nodes in the actual group relationship feature map, and use the detection information of the individual targets as the node attributes of the nodes to obtain a node set; based on the interaction intensity between any two individual targets, obtain an edge set; construct the actual group relationship feature map based on the node set and the edge set. Among them, the detection information includes the confidence, position information, apparent feature, and motion feature of the individual target.
[0027] The following is an expanded description. Refer to Figure 2 , first analyze how to obtain the detection information of individual targets.
[0028] In the first specific implementation, a surveillance image is obtained from a surveillance video, and the surveillance image is processed by multi-scale scaling to generate input images of multiple scales, and the input images of multiple scales are input in parallel into a multi-scale object detection model (You Only Look Once). Further, individual objects in the surveillance image are identified by the multi-scale object detection model, and the confidence and position information of the individual objects are obtained (the confidence and position information are collectively referred to as Figure 2 the detection results in). Specifically, the YOLO detection model focuses on the global information of the entire image during the training and inference phases. Compared with object detection methods based on region proposals (such as the Fast Region-based Convolutional Neural Network, which only focuses on the local image information within the candidate boxes during detection), the YOLO detection model can effectively reduce the background misdetection rate. Assume that a single-frame image of the surveillance image is , and the scaled image is input into the YOLO detection model. Assume that after passing through the YOLO detection model, individual objects are obtained, , and at the same time, confidences of the individual objects are obtained and position information . Among them, the confidence of the individual object is used to evaluate the reliability of the individual object. Assume that the confidence threshold is set to 0.5. If the confidence of the individual object is 0.3, it means that the YOLO detection model has low recognition reliability for this individual object and may be a misdetection and needs to be excluded. If the confidence of the individual object is 0.9, it means that the YOLO detection model has high recognition reliability for this individual object and will be included in the subsequent processing. Further, the position information of the individual object includes the center point coordinates of the individual object, the height of the individual object, and the width of the individual object.
[0029] In the second specific implementation, based on the position information of the individual object, the corresponding region image is extracted from the surveillance image, and the region image is preprocessed, and then the preprocessed region image is input into a pre-trained convolutional neural network model to extract the apparent features of the individual object through the convolutional neural network model , represents the apparent features of the th individual object, and the apparent features represent the visual appearance features of the individual object. The apparent features include but are not limited to color, texture, etc. The preprocessing includes image size processing, image enhancement processing, etc. Compared with traditional feature extraction methods that rely on manually designed feature extraction rules, the convolutional neural network model can automatically learn the most effective apparent feature extraction method from the data through the layer-by-layer training of multiple neurons, greatly reducing the cost of manual intervention.
[0030] In the third specific implementation manner, based on the position information of the individual target, the corresponding regional image is extracted from the monitoring screen, and the optical flow information at the regional image is extracted through the optical flow estimation network, and then the motion characteristics of the individual target are determined according to the optical flow information at the regional image. , denotes the motion characteristics of the
[0031] th individual target. In the monitoring scenario, assuming that the brightness difference of the same object in adjacent frames is small, the motion characteristics of the individual target are extracted by the optical flow method using the temporal dimension information of the monitoring video sequence. This method does not require prior scene information and is applicable to the analysis of group motion in complex monitoring environments. , respectively represent the appearance characteristics, motion characteristics, confidence, and position information of the th individual target. The matrix is the node attribute of all nodes, is the number of nodes, is the dimension of the feature vector of the node, and the node set is represented by . Next, analyze how to obtain the edge set
[0032] based on the interaction intensity between any two individual targets.
[0033] Exemplarily, the calculation method of the weight value is as follows: ; The calculation method of the target matrix is as follows: ; is the edge weight attribute between the th individual target and the th individual target, is the spatial distance between the th individual target and the th individual target, and adjust the distribution and sparsity of the target matrix.
[0034] Finally, construct an actual group relationship feature graph based on the node set and edge set. .
[0035] It can be seen that the group relationship feature graph has powerful data understanding and cognitive capabilities. In the group relationship feature graph, nodes are connected to each other through edges, which can effectively express the relationships between different individual goals. By correlating each individual goal in the surveillance video through the group relationship feature graph, it helps to perform effective relationship reasoning, thereby enhancing the understanding ability of group anomalies and improving the accuracy of anomaly event detection.
[0036] Step S12: Input the group relationship feature graph sequence into a preset graph prediction network, and determine the predicted group relationship feature graph of the future frame through the preset graph prediction network.
[0037] In this embodiment, the preset graph prediction network includes a temporal convolutional block and a spatial convolutional block; among them, the temporal convolutional block processes the temporal dimension features of the group relationship feature graph sequence, and the spatial convolutional block processes the node topological features of the actual group relationship feature graph of a single frame in the group relationship feature graph sequence. Specifically, the preset graph prediction network is constructed around the collaborative processing of temporal and spatial features, and its operation logic Figure 3 is consistent with that of, and this network is based on an autoencoder framework, inputting a group relationship feature graph sequence containing M time steps and a target matrix , and the whole can be abstracted as , represents the input feature dimension of the th individual target, represents the real number field, the encoder part integrates the temporal convolutional block and the spatial convolutional block, and the two take the "spatiotemporal convolutional block" as the basic unit (each spatiotemporal convolutional block consists of two temporal convolutional blocks and one spatial convolutional block). The temporal convolutional block focuses on the temporal dimension features of the group relationship feature graph sequence, follows the time causality by means of causal convolution, and uses the historical time step features and relationship features , and predicts the relationship feature of the current time step according to the probability chain rule . The spatial convolutional block extracts node topological features within the graph structure of a single time step (not between time steps) to ensure that the spatial dimension analysis focuses on the current frame. After being encoded by the spatiotemporal convolutional network, the hidden layer feature is output, represents the output feature dimension. After obtaining the hidden layer feature, it is necessary to further adapt to the input requirements of the decoder, and complete the fine adjustment of the feature dimension or distribution through the hidden vector transformation (corresponding to in Figure 4). The adjusted flows into the decoder and uses the "inner product decoding" mechanism: first, is multiplied by its transpose Perform an inner product operation, and then pass through an activation function Map the result of the inner product operation to the interval [0,1] to generate a predicted group relationship feature map (corresponding to Figure 3 in ), realizing the deduction from historical spatio-temporal features to future group relationships. To constrain the accuracy of encoding and decoding, a loss function is introduced, with the real group relationship feature map as the supervision, so that the hidden layer features and decoding logic are continuously optimized to ensure that the predicted map approximates the evolution trend of the actual group relationship, and respectively represent the hidden layer feature vectors of the corresponding nodes after the encoder encodes the input sequence.
[0038] See Figure 4 as shown, Figure 4 is a schematic diagram of the spatio-temporal convolution module, which is the core component of the preset graph prediction network (corresponding to Figure 3 architecture). The left half of its process connects to Figure 3 encoding logic, Figure 4 the input l-th layer feature a l in Figure 3 is the intermediate feature of the encoding process in l . First, input a into the temporal gated convolution, extract the cross-time dynamics through the temporal gated convolution, and then input the result of the temporal gated convolution and the target matrix l+1 into the spatial graph convolution, extract the spatial topology of the current time step through the spatial graph convolution, and then fuse through the temporal gated convolution to output the l+1-th layer feature a Figure 4 The right half of l t-(M-1) , … a l t ), processed by two parallel causal convolutions. The output of one causal convolution passes through an activation function to generate the gating weight, and the output of the other causal convolution is fused with the information introduced by the residual connection. Then, adaptive feature selection is achieved through element-wise multiplication according to the gating weight. This module provides the core spatio-temporal feature extraction ability for the preset graph prediction network through the cooperation of temporal gated convolution and spatial graph convolution. The output a l+1 is gradually refined into the hidden layer feature Z in Figure 3 after multiple iterations, thus realizing the processing of the temporal dimension features and single-frame node topology features of the group relationship feature map sequence.
[0039] Step S13: Determine the peak signal-to-noise ratio between the predicted group relationship feature map of the future frame and the actual group relationship feature map of the future frame, and generate a target anomaly score according to the peak signal-to-noise ratio.
[0040] In this embodiment, the target peak signal-to-noise ratio between the predicted group relationship feature map of the future frame and the actual group relationship feature map of the future frame is determined; the target anomaly score is generated according to the target peak signal-to-noise ratio, the historical minimum peak signal-to-noise ratio, and the historical maximum peak signal-to-noise ratio. Specifically, the first difference between the target peak signal-to-noise ratio and the historical minimum peak signal-to-noise ratio is calculated; the second difference between the historical maximum peak signal-to-noise ratio and the historical minimum peak signal-to-noise ratio is calculated; the ratio of the first difference to the second difference is normalized to obtain the target anomaly score. The calculation process of the above process is as follows: ; ; in, represents the target peak signal-to-noise ratio, Indicates the minimum historical peak signal-to-noise ratio, Indicates the maximum value of the historical peak signal-to-noise ratio, Represents the target anomaly score.
[0041] Among them, the larger the peak signal-to-noise ratio, the higher the fit between the predicted group relationship feature map and the real feature map, the better the prediction effect, which means that the corresponding scene is more likely to be a normal event.
[0042] Step S14: judging whether a group abnormal event occurs according to the target abnormality score, so as to perform early warning processing when a group abnormal event occurs.
[0043] In this embodiment, judging whether a group abnormal event occurs according to the target abnormality score includes: judging whether the target abnormality score exceeds a preset score threshold, and if so, judging that a group abnormal event occurs.
[0044] See also Figure 5As shown in the figure, the figure reflects the complete process of the group anomaly detection scheme, including two major stages: offline learning and anomaly detection: (1) Offline learning: Taking the normal scene training sample set as the input, first through graph relationship construction, the video frames are transformed into a graph structure with node features (i.e., the aforementioned group relationship feature graph); then input into the network model (i.e., the aforementioned preset graph prediction network), and trained with the goal of minimizing the loss function, so that the network model learns the graph pattern of normal group relationships. After the training is completed, the network parameters are fixed, and a model that can accurately predict normal group relationships is output. (2) Anomaly detection: Input the test samples (including abnormal / normal scenes), and also perform graph relationship construction to generate a graph structure; input it into the trained model with fixed parameters, and output the predicted group relationship feature graph; compare the difference between the graph generated by the test sample and the predicted graph output by the model through similarity measurement (i.e., calculating the peak signal-to-noise ratio), and finally calculate the anomaly score to determine whether the group behavior is abnormal (the greater the difference, the higher the probability of anomaly). The advantages of the above scheme are: (1) Detection framework based on pre-learning of normal patterns + comparison of abnormal differences: Breaking through the limitations of traditional frame-by-frame detection, first offline learning the graph pattern of normal group relationships, and then online comparing the graph differences of test scenes, adapting to the scene requirements of "complex associations and occasional anomalies" in group behavior, making anomaly detection more in line with the logic of group interaction. (2) Modeling group relationships with graph structure: Using a graph (nodes + edges) to replace traditional feature vectors, accurately capturing the spatial associations and behavioral interactions of individuals in the group (such as people gathering and motion coordination), enabling the model to learn more fine-grained normal patterns and improving the accuracy and robustness of anomaly detection. In short, through "graph structure modeling + offline-online staged detection", the efficient recognition of group anomaly events in complex scenarios is realized, and the problems of difficult description of group associations and high anomaly omission rates in traditional methods are solved.
[0045] Furthermore, the present application also proposes a dynamic adaptive mechanism for adjusting the weights of group relationship graph features. This adjustment mechanism can design an attention mechanism to analyze the contribution degree of each feature to group relationship modeling in real time. For example, in a crowded scene, it automatically increases the weights of position features and motion features, and in a low-density scene, it enhances the influence of appearance features, thereby adaptively optimizing the representation ability of the graph structure. In addition, a cross-modal feature fusion strategy can be introduced to dynamically weight the optical flow information and appearance features through an attention gating mechanism, avoiding problems of feature redundancy or missing caused by traditional fixed weights, and further improving the generalization and accuracy of anomaly detection, especially suitable for group behavior analysis in complex environments.
[0046] It can be seen that the present application proposes a method for detecting group abnormal events, including: constructing an actual group relationship feature map corresponding to each frame based on the detection information of individual targets in each frame of the monitoring video, and determining a sequence of group relationship feature maps based on the actual group relationship feature maps of each frame; inputting the sequence of group relationship feature maps into a preset graph prediction network, and determining a predicted group relationship feature map of a future frame through the preset graph prediction network; determining the peak signal-to-noise ratio between the predicted group relationship feature map of the future frame and the actual group relationship feature map of the future frame, and generating a target abnormal score according to the peak signal-to-noise ratio; and judging whether a group abnormal event occurs according to the target abnormal score, so as to perform early warning processing when a group abnormal event occurs.
[0047] Beneficial effects: Based on the detection information of individual targets in each frame of the monitoring video, the present application constructs an actual group relationship feature map corresponding to each frame, and determines a sequence of group relationship feature maps based on the actual group relationship feature maps of each frame. In this way, the group relationship can be quantitatively characterized in the form of a graph structure. Compared with the traditional method that only detects abnormal events based on appearance information and motion information, the present application captures the key feature of group relationship more comprehensively, provides richer feature inputs for abnormal detection, and thus improves the accuracy of abnormal detection. Further, the present application inputs the sequence of group relationship feature maps into a preset graph prediction network, thereby determining the predicted group relationship feature map of a future frame through the preset graph prediction network, realizing the modeling and prediction of group behavior, and generating a target abnormal score by calculating the peak signal-to-noise ratio between the predicted group relationship feature map of the future frame and the actual group relationship feature map of the future frame. In this way, the difference degree between the predicted group relationship feature map and the actual group relationship feature map can be quantified, so as to accurately judge whether a group abnormal event occurs, avoiding the problem of inaccurate detection caused by the traditional method not fully considering the group relationship feature, and at the same time realizing the effective early warning processing of group abnormal events.
[0048] Correspondingly, the embodiment of the present application also discloses a device for detecting group abnormal events. Refer to Figure 6 as shown, the device includes: Group relationship feature map construction module 11: constructing an actual group relationship feature map corresponding to each frame based on the detection information of individual targets in each frame of the monitoring video, and determining a sequence of group relationship feature maps based on the actual group relationship feature maps of each frame; Graph prediction network processing module 12: inputting the sequence of group relationship feature maps into a preset graph prediction network, and determining the predicted group relationship feature map of a future frame through the preset graph prediction network; Abnormal score calculation module 13: determining the peak signal-to-noise ratio between the predicted group relationship feature map of the future frame and the actual group relationship feature map of the future frame, and generating a target abnormal score according to the peak signal-to-noise ratio; Abnormal event determination and early warning module 14: Determine whether a group abnormal event occurs according to the target abnormal score, so as to perform early warning processing when a group abnormal event occurs.
[0049] Among them, for the more specific working processes of the above-mentioned various modules, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.
[0050] It can be seen that the present application proposes a method for detecting group abnormal events, including: constructing an actual group relationship feature map corresponding to each frame based on the detection information of individual targets in each frame of the monitoring video, and determining a sequence of group relationship feature maps based on the actual group relationship feature maps of each frame; inputting the sequence of group relationship feature maps into a preset graph prediction network, and determining a predicted group relationship feature map for a future frame through the preset graph prediction network; determining the peak signal-to-noise ratio between the predicted group relationship feature map of the future frame and the actual group relationship feature map of the future frame, and generating a target abnormal score according to the peak signal-to-noise ratio; determining whether a group abnormal event occurs according to the target abnormal score, so as to perform early warning processing when a group abnormal event occurs.
[0051] Beneficial effects: Based on the detection information of individual targets in each frame of the monitoring video, the present application constructs an actual group relationship feature map corresponding to each frame, and determines a sequence of group relationship feature maps based on the actual group relationship feature maps of each frame. In this way, the group relationship can be quantitatively characterized in the form of a graph structure. Compared with the traditional method of detecting abnormal events only based on appearance information and motion information, the present application captures the key feature of group relationship more comprehensively, provides richer feature inputs for abnormal detection, and thus improves the accuracy of abnormal detection. Further, the present application inputs the sequence of group relationship feature maps into a preset graph prediction network, and thus determines a predicted group relationship feature map for a future frame through the preset graph prediction network, realizing the modeling and prediction of group behavior, and generating a target abnormal score by calculating the peak signal-to-noise ratio between the predicted group relationship feature map of the future frame and the actual group relationship feature map of the future frame. In this way, the difference degree between the predicted group relationship feature map and the actual group relationship feature map can be quantified, so as to accurately determine whether a group abnormal event occurs, avoiding the problem of inaccurate detection caused by the traditional method not fully considering the group relationship feature, and at the same time realizing the effective early warning processing of group abnormal events.
[0052] Further, an embodiment of the present application also provides a computer program product, including a computer program, which when executed by a processor, implements the steps of the foregoing method for detecting group abnormal events.
[0053] Further, an embodiment of the present application also provides an electronic device. Figure 7 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be regarded as any limitation on the scope of use of the present application.
[0054] Figure 7 FIG. 20 is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a display screen 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27. Among them, the memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the group anomaly event detection method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0055] In this embodiment, the power supply 26 is used to provide a working voltage for each hardware device on the electronic device 20; the communication interface 25 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed thereon here; the input / output interface 24 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0056] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc., and the resources stored thereon may include a computer program 221, and the storage method may be short-term storage or permanent storage. Among them, in addition to the computer program that can be used to complete the group anomaly event detection method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 221 may further include a computer program that can be used to complete other specific tasks.
[0057] Further, an embodiment of the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the group anomaly event detection method disclosed above is implemented.
[0058] For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.
[0059] The various embodiments in this application are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts between the various embodiments, reference may be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and reference may be made to the description in the method part for the relevant parts.
[0060] Those skilled in the art may further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0061] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0062] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0063] The above has introduced in detail a method for detecting group abnormal events, a computer program product, a device and a medium provided by this application. Specific examples are used in this document to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for detecting group abnormal events, characterized in that, Including: Based on the detection information of individual targets in each frame of the monitoring video, construct the actual group relationship feature map corresponding to each frame, and determine the sequence of group relationship feature maps based on the actual group relationship feature maps of each frame; Input the sequence of group relationship feature maps into a preset graph prediction network, and determine the predicted group relationship feature map of the future frame through the preset graph prediction network; Determine the peak signal-to-noise ratio between the predicted group relationship feature map of the future frame and the actual group relationship feature map of the future frame, and generate a target anomaly score according to the peak signal-to-noise ratio; Judge whether a group anomaly event occurs according to the target anomaly score, so as to perform early warning processing when the group anomaly event occurs.
2. The method for detecting group abnormal events according to claim 1, wherein The constructing the actual group relationship feature map corresponding to each frame based on the detection information of individual targets in each frame of the monitoring video includes: Identify individual targets in each frame of the monitoring video, and obtain the detection information of the individual targets; Take the individual targets as the nodes in the actual group relationship feature map, and take the detection information of the individual targets as the node attributes of the nodes to obtain a node set; Obtain an edge set based on the interaction intensity between any two of the individual targets; Construct the actual group relationship feature map based on the node set and the edge set.
3. The method for detecting group abnormal events according to claim 2, wherein, The detection information includes the confidence, position information, appearance feature and motion feature of the individual target.
4. The method for detecting group abnormal events according to claim 3, wherein, The identifying individual targets in each frame of the monitoring video and obtaining the detection information of the individual targets includes: Obtain the monitoring video from the surveillance video; Perform multi-scale scaling processing on the monitoring video to generate input images of multiple scales, and input the input images of multiple scales into a multi-scale object detection model in parallel; Identify individual targets in the monitoring video through the multi-scale object detection model, and obtain the confidence and position information of the individual targets.
5. The method for detecting group abnormal events according to claim 3, wherein The identifying individual targets in each frame of the monitoring video and obtaining the detection information of the individual targets includes: Extract the corresponding regional image from the monitoring video based on the position information of the individual target; Preprocess the regional image, and input the preprocessed regional image into a pre-trained convolutional neural network model, so as to extract the appearance feature of the individual target through the convolutional neural network model; wherein, the appearance feature represents the visual appearance feature of the individual target.
6. The method for detecting group abnormal events according to claim 3, wherein The identifying individual targets in each frame of the monitoring video and obtaining the detection information of the individual targets includes: Extract the corresponding regional image from the monitoring video based on the position information of the individual target; Extract the optical flow information at the regional image through an optical flow estimation network, and determine the motion feature of the individual target according to the optical flow information at the regional image.
7. The method for detecting group abnormal events according to claim 3, wherein Also including: Based on the position information of any two individual targets, determine the spatial distance between the any two individual targets, and calculate the edge weight attribute between the any two individual targets according to the spatial distance; wherein, the edge weight attribute represents the interaction intensity between any two of the individual targets.
8. The method for detecting group abnormal events according to claim 7, wherein Obtaining an edge set based on the interaction intensity between any two of the individual targets, including: Taking the edge weight attribute between any two individual targets as the weight value between any two individual targets, and constructing a target matrix according to the weight value between any two individual targets; Determining the edge set based on the target matrix.
9. The method for detecting group abnormal events according to claim 1, wherein Judging whether a group anomaly event occurs according to the target anomaly score, including: Judging whether the target anomaly score exceeds a preset score threshold; If the target anomaly score exceeds the preset score threshold, it is determined that a group anomaly event occurs.
10. The method for detecting group abnormal events according to claim 1, wherein The preset graph prediction network includes a time-domain convolution block and a spatial-domain convolution block; wherein, the time-domain convolution block processes the time-domain dimension features of the group relationship feature map sequence, and the spatial-domain convolution block processes the node topology features of the actual group relationship feature map of a single frame in the group relationship feature map sequence.
11. The method for detecting group abnormal events according to any one of claims 1 to 10, characterized in that, Determining the peak signal-to-noise ratio between the predicted group relationship feature map of the future frame and the actual group relationship feature map of the future frame, and generating a target anomaly score according to the peak signal-to-noise ratio, including: Determining the target peak signal-to-noise ratio between the predicted group relationship feature map of the future frame and the actual group relationship feature map of the future frame; Generating the target anomaly score according to the target peak signal-to-noise ratio, the minimum value of the historical peak signal-to-noise ratio, and the maximum value of the historical peak signal-to-noise ratio.
12. The method for detecting group abnormal events according to claim 11, wherein Generating the target anomaly score according to the target peak signal-to-noise ratio, the minimum value of the historical peak signal-to-noise ratio, and the maximum value of the historical peak signal-to-noise ratio, including: Calculating a first difference between the target peak signal-to-noise ratio and the minimum value of the historical peak signal-to-noise ratio; Calculating a second difference between the maximum value of the historical peak signal-to-noise ratio and the minimum value of the historical peak signal-to-noise ratio; Normalizing the ratio of the first difference to the second difference to obtain the target anomaly score.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the group anomaly event detection method according to any one of claims 1 to 12 are implemented.
14. An electronic device, characterized in that, Including: A memory for storing a computer program; A processor for executing the computer program to implement the group anomaly event detection method according to any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by a processor, the group anomaly event detection method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Abnormal group detection method, device and equipment
CN111770047A
Crowd abnormity detection method based on generative adversarial network
CN111881750A
Group abnormal behavior identification method and system, storage medium and equipment
CN113269104A
Crowd abnormal behavior detection method based on improved SSD
CN115273234A
Crowd safety abnormal event identification method
CN116229347A