Early warning method and device based on anti-terrorism riot abnormality detection
By comprehensively utilizing multi-source situational data and using neural network detection, the problem of insufficient threat information acquisition in complex environments in existing counter-terrorism and anti-riot early warning technologies has been solved. This has enabled unified analysis of multi-level threat information, improving the security and rapid response capabilities of counter-terrorism and anti-riot decision-making.
Patent Information
- Application Number
- CN202411620771.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2044-11-14
AI Technical Summary
Existing counter-terrorism and riot control early warning technologies are unable to accurately acquire threat information in complex environments and lack unified and holistic comprehensive analysis of multi-level threat information, making it difficult to meet the security and rapid response requirements of counter-terrorism and riot control emergency response decision-making.
Multi-source situational data is used to detect armed targets, suspicious behaviors, and target yaw trajectories. Threat level is determined by using a neural network with multi-layer feature extraction, pooling, upsampling, multi-scale feature fusion, and key feature focus. The residual BiLSTM neural network is combined to detect target yaw trajectories, thereby achieving a unified and holistic comprehensive analysis of multi-level threat information.
It enables comprehensive acquisition and real-time perception of information on the situation of terrorism and violence in complex environments, improves crisis response and handling capabilities in counter-terrorism and anti-riot decision-making, and enhances the accuracy and speed of detection.
Smart Images

Figure CN119559553B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of anti-terrorism and riot monitoring and early warning. Specifically, it relates to an early warning method and device based on anti-terrorism and riot anomaly detection, an electronic device and a computer readable storage medium. BACKGROUND
[0002] The peacekeeping anti-terrorism and riot task area usually covers an environment with complex terrain and climate, resulting in complex and diverse situation information with numerous types and quantities. Moreover, the terrorist targets usually disguise and hide, have unusual and variable movement types, and are interfered by background crowds, which seriously affects the anti-terrorism and riot emergency response decision-making.
[0003] The existing anti-terrorism and riot early warning technology is aimed at normal areas such as private residences, shopping malls, public areas, etc. The detection means used is not targeted at the anti-terrorism and riot scene, and cannot accurately obtain threat information in the anti-terrorism and riot scene. Moreover, the existing anti-terrorism and riot detection means is single, lacks complementary linkage and integration of situation information, and is difficult to form a unified and integrated comprehensive analysis of multi-level threat information, which cannot meet the safety and rapid response requirements of anti-terrorism and riot emergency response decision-making. SUMMARY
[0004] The purpose of the present application is to comprehensively obtain anti-terrorism and riot situation information and integrate and complement it, to realize a unified and integrated comprehensive analysis of multi-level threat information, and to improve the crisis response and disposal capability in anti-terrorism and riot decision-making, so as to solve the problems in the background art.
[0005] To achieve the above purpose, in a first aspect, the present application provides an early warning method based on anti-terrorism and riot anomaly detection, comprising:
[0006] obtaining multi-source situation data of a detection task area, the multi-source situation data being multi-modal situation data obtained by a plurality of detection and perception devices detecting the same detection task area, the multi-source situation data comprising video stream data, picture frame data and target space-time trajectory data;
[0007] performing weapon-carrying target detection using the video stream data in the multi-source situation data, screening weapon-carrying targets in the detection task area, and obtaining a weapon-carrying target screening result;
[0008] performing suspicious behavior detection using the picture frame data in the multi-source situation data, screening behavior-suspected targets in the detection task area, and obtaining a behavior-suspected target screening result;
[0009] performing target yaw trajectory detection using the target space-time trajectory data in the multi-source situation data, screening yaw targets in the detection task area, and obtaining a yaw target screening result;
[0010] Based on the screening result of the weapon-carrying target, the screening result of the suspicious behavior target and the screening result of the deviation target, a threat level of the detection task area is judged, and a warning is given based on the threat level.
[0011] In some embodiments of the present application, video stream data in multi-source situation data is input into a pre-trained weapon-carrying target detection neural network for weapon-carrying target detection to screen weapon-carrying targets in the detection task area; wherein the weapon-carrying target detection includes multi-layer feature extraction, pooling, up-sampling, multi-scale feature fusion, key feature attention and multi-scale feature detection.
[0012] In some embodiments of the present application, the weapon-carrying target detection neural network includes a feature extraction module, a pooling module, a feature fusion module and a target detection module; wherein the video stream data in the multi-source situation data is input into the feature extraction module for multi-layer feature extraction, one output of the feature extraction module is input into the feature fusion module, and the other output is input into the feature fusion module after being pooled by the pooling module, the feature fusion module performs up-sampling, multi-scale feature fusion and key feature attention on the two inputs, and the output of the feature fusion module is input into the target detection module, and the target detection module performs multi-scale feature detection to output the detection result of the weapon-carrying target.
[0013] In some embodiments of the present application, the pooling module includes a first convolution module, a second convolution module, a first strip pooling module, a second strip pooling module, a third strip pooling module, a first maximum pooling module, a second maximum pooling module, a third maximum pooling module, a first splicing module and a second splicing module; wherein the output of the feature extraction module is input into the first convolution module; the first output of the first convolution module is directly input into the first splicing module, the second output of the first convolution module is subjected to three times of strip pooling by the first strip pooling module, the second strip pooling module and the third strip pooling module, and the third output of the first convolution module is subjected to three times of maximum pooling by the first maximum pooling module, the second maximum pooling module and the third maximum pooling module; the outputs of the first strip pooling module and the second strip pooling module are input into the first splicing module after being multiplied by a first weight, the outputs of the first maximum pooling module and the second maximum pooling module are input into the first splicing module after being multiplied by a second weight, the first splicing module splices the five inputs to output a result to the second splicing module, the output of the third strip pooling module is input into the second splicing module after being multiplied by the first weight, the output of the third maximum pooling module is input into the second splicing module after being multiplied by the second weight, the second splicing module splices the three inputs to input into the second convolution module, and the second convolution module convolves to output a final pooling result.
[0014] In some embodiments of the present application, picture frame data in multi-source situation data is input into a pre-trained suspicious behavior detection neural network for suspicious behavior detection to screen suspicious behavior targets in a detection task area; wherein the suspicious behavior detection neural network predicts a prediction frame corresponding to an N+1th picture frame from N continuous picture frames, thereby generating prediction frames with time sequence, calculating errors between the prediction frames and the picture frames corresponding to the same time point, and determining the suspicious behavior targets and their threat intentions at the time point according to whether the errors exceed a preset threshold.
[0015] In some embodiments of the present application, the suspicious behavior detection neural network is a two-way constraint comparison detection network, which adopts optical flow frame constraint and diffusion frame constraint to jointly optimize network loss of the suspicious behavior detection neural network.
[0016] In some embodiments of the present application, target space-time trajectory data in multi-source situation data is input into a pre-trained target deviation trajectory detection neural network for target deviation trajectory detection to screen deviation targets in a detection task area; wherein the target deviation trajectory detection neural network is a residual BiLSTM neural network.
[0017] In a second aspect, the present application provides a warning device based on anti-terrorism and riot abnormality detection, comprising:
[0018] A data acquisition module is configured to acquire multi-source situation data of a detection task area, wherein the multi-source situation data is multi-modal data obtained by a plurality of detection and sensing devices detecting the same detection task area, and the multi-source situation data includes video stream data, picture frame data and target space-time trajectory data.
[0019] A weapon-carrying target screening module is configured to detect weapon-carrying targets in the detection task area using the video stream data in the multi-source situation data, and obtain a weapon-carrying target screening result.
[0020] A suspicious target screening module is configured to detect suspicious behavior in the detection task area using the picture frame data in the multi-source situation data, and obtain a suspicious behavior target screening result.
[0021] A deviation target screening module is configured to detect target deviation trajectories in the detection task area using the target space-time trajectory data in the multi-source situation data, and obtain a deviation target screening result.
[0022] A threat judgment and warning module is configured to judge a threat level of the detection task area based on the weapon-carrying target screening result, the suspicious behavior target screening result and the deviation target screening result, and give a warning based on the threat level.
[0023] In a third aspect, the present application provides an electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the early warning method based on anti-terrorism and riot abnormality detection.
[0024] In a fourth aspect, the present application provides a computer readable storage medium, storing a computer program, wherein the computer program is executed by a processor to implement the early warning method based on anti-terrorism and riot abnormality detection.
[0025] The present application has the following beneficial effects:
[0026] The present application utilizes multi-source situation data to form complementation, more comprehensively masters suspicious target trend, realizes real-time perception of on-site situation, and improves crisis response and disposal capability in anti-terrorism and riot decision. BRIEF DESCRIPTION OF DRAWINGS
[0027] Figure 1 A flow chart of the early warning method based on anti-terrorism and riot abnormality detection in an embodiment of the present application.
[0028] Figure 2 A flow chart of the early warning method based on anti-terrorism and riot abnormality detection in another embodiment of the present application.
[0029] Figure 3 A flow chart of the armed target detection in an embodiment of the present application.
[0030] Figure 4 A component schematic diagram of the armed target detection neural network in an embodiment of the present application.
[0031] Figure 5 A component schematic diagram of the pooling module in an embodiment of the present application.
[0032] Figure 6 A module connection schematic diagram of the armed target detection neural network in an embodiment of the present application.
[0033] Figure 7 A flow chart of the suspicious behavior detection in an embodiment of the present application.
[0034] Figure 8 A schematic diagram of the two-hop memory U-Net network structure in an embodiment of the present application.
[0035] Figure 9 A principle schematic diagram of the double-channel constraint contrast detection network DCC-UNet training in an embodiment of the present application.
[0036] Figure 10 Flowchart of the target yaw trajectory detection in an embodiment of the present application.
[0037] Figure 11 Schematic diagram of the LSTM basic structure in an embodiment of the present application.
[0038] Figure 12 Schematic diagram of the residual double BiLSTM network structure in an embodiment of the present application.
[0039] Figure 13 Schematic diagram of the error accumulation using LSTM multi-step prediction in an embodiment of the present application.
[0040] Figure 14 Schematic diagram of the dynamic time step prediction mechanism using LSTM in an embodiment of the present application.
[0041] Figure 15 Schematic diagram of the threat assessment and early warning principle in an embodiment of the present application.
[0042] Figure 16 Schematic diagram of the composition principle of the early warning device in an embodiment of the present application.
[0043] Figure 17 Schematic diagram of the composition principle of the electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0044] The present application is described herein with reference to specific embodiments thereof which are illustrated in the attached drawings. These embodiments are described in detail so as to enable practice of the present application in its various aspects. Other embodiments of the present application, however, will become apparent to those skilled in the art upon reading the description herein, which is illustrative of the principles of the present application. Therefore, other embodiments and modifications to the disclosed embodiments can be employed without departing from the spirit and scope of the present application. It should be noted that features illustrated in the following embodiments and in the claims can be combined with each other, unless specifically noted otherwise.
[0045] It should be noted that the drawings included in the following embodiments are only schematic and so are not to precise scale shown in the figures. In the description of embodiments of the present application, relative terms are used to describe their orientation, for example, near, distal, upper, lower, behind, in front of, between, on, under, and the like for the purposes of this description. These terms are used as terms of convenience to aid understanding of the application and are in no way intended to limit the scope of the application. Moreover, the drawings are not to scale unless specifically noted.
[0046] The particularity of the anti-terrorism and riot scene application scene lies in the accuracy and rapidity requirements of the disposal decision. The task area usually covers the environment with complex terrain and climate, the situation information is often complex and diverse, the target movement type is abnormal and changeable, the background crowd often interferes, and various factors lead to high safety requirements and rapid response requirements of the disposal decision of the emergency.
[0047] The large amount of multi-dimensional data obtained by the front-end detector of the situation awareness needs effective analysis and processing methods, requires accurately extracting potential threat factors of the target from multi-source data, conducting threat evaluation, making effective early warning, and providing reliable criteria for decision makers. The potential threat target has distinct characteristics: the potential threat individual may act at an abnormal time, often repeatedly lingers, refuses to cooperate and other behaviors, its mode is often inconsistent with the general behavior of the background crowd, the potential threat group may carry the same logo pattern on the skin or personal belongings, often illegally gathers, provokes conflicts and other behaviors, the target usually holds or tries to hide dangerous weapons such as 'knife', 'gun','stick', and the action trajectory, speed and the like often have sudden changes.
[0048] Therefore, the present application provides an early warning method, device, platform and computer readable storage medium based on anti-terrorism and riot anomaly detection. The purpose is to comprehensively obtain anti-terrorism and riot situation information, and integrate and complement each other, realize unified and integrated comprehensive research and judgment of multi-level threat information, and improve the crisis response and disposal capability in anti-terrorism and riot decision.
[0049] In one embodiment, the early warning method based on anti-terrorism and riot anomaly detection provided by the present application comprises the following steps: Figure 1
[0050] In step S100, multi-source situation data of a detection task area is obtained, the multi-source situation data being multi-modal situation data obtained by detecting the same detection task area by a plurality of detection and perception devices, and the multi-source situation data comprising video stream data, picture frame data and target space-time trajectory data;
[0051] In step S200, the video stream data in the multi-source situation data is used for weapon-carrying target detection, a weapon-carrying target in the detection task area is screened, and a weapon-carrying target screening result is obtained;
[0052] In step S300, the picture frame data in the multi-source situation data is used for suspicious behavior detection, a behavior-suspected target in the detection task area is screened, and a behavior-suspected target screening result is obtained;
[0053] In step S400, the target space-time trajectory data in the multi-source situation data is used for target deviation trajectory detection, a deviation target in the detection task area is screened, and a deviation target screening result is obtained;
[0054] Step S500: Based on the screening results of armed targets, suspicious behavior targets, and yaw targets, determine the threat level of the detection task area and issue an early warning based on the threat level.
[0055] In this embodiment, steps S200 and S400 are implemented in a serial processing manner. However, in other embodiments of the present invention, steps S200 and S400 can also be implemented in a parallel processing manner, such as... Figure 2 As shown.
[0056] Furthermore, in some embodiments of the present invention, in step S200, video stream data from multi-source situational data is input into a pre-trained armed target detection neural network to detect armed targets and filter armed targets in the detection task area; wherein, as Figure 3 As shown, the weapon target detection sequentially performs multi-layer feature extraction, pooling, upsampling, multi-scale feature fusion, key feature attention, and multi-scale feature detection.
[0057] Specifically, such as Figure 4 As shown, the weapon-carrying target detection neural network includes a feature extraction module, a pooling module, a feature fusion module, and a target detection module. Video stream data from multi-source situational data is input to the feature extraction module for multi-layer feature extraction. One output of the feature extraction module is input to the feature fusion module, and the other output is pooled by the pooling module and then input to the feature fusion module. The feature fusion module upsamples the two inputs, performs multi-scale feature fusion, and focuses on key features. The output of the feature fusion module is input to the target detection module, which performs multi-scale feature detection and outputs the detection result of the weapon-carrying target.
[0058] Pooling includes max pooling, average pooling, global average pooling, random pooling, etc. Preferably, in some embodiments of the present invention, in order to avoid false detections, missed detections, and mixed detections in counter-terrorism and riot control scenarios, distinguish between people and dangerous instruments, and specifically detect dangerous instruments targeting abnormal targets, the pooling module employs methods such as... Figure 5The shown symmetric weighted pooling structure includes a first convolution module, a second convolution module, a first strip pooling module, a second strip pooling module, a third strip pooling module, a first maximum pooling module, a second maximum pooling module, a third maximum pooling module, a first splicing module and a second splicing module; the output of the feature extraction module is input into the first convolution module; the first path output of the first convolution module is directly input into the first splicing module, the second path output of the first convolution module is subjected to three strip poolings through the first strip pooling module, the second strip pooling module and the third strip pooling module, and the third path output of the first convolution module is subjected to three maximum poolings through the first maximum pooling module, the second maximum pooling module and the third maximum pooling module; the outputs of the first strip pooling module and the second strip pooling module are multiplied by a weight w1 and then input into the first splicing module, the outputs of the first maximum pooling module and the second maximum pooling module are multiplied by a weight w2 and then input into the first splicing module, the first splicing module splices five path inputs and then inputs into the second splicing module, the output of the third strip pooling module is multiplied by the weight w1 and then input into the second splicing module, the output of the third maximum pooling module is multiplied by the weight w2 and then input into the second splicing module, the second splicing module splices three path inputs and then inputs into the second convolution module, and the second convolution module outputs a final pooling result.
[0059] For the anti-terrorism and anti-riot scene, the above-mentioned pooling module can detect targets carrying weapons (such as guns, knives, sticks, etc.), without missing reports or false positives, and with high accuracy.
[0060] Specifically, in some embodiments of the present application, the structure of the weapon-carrying target detection neural network is as shown in the figure. Figure 6 The meanings of the modules in the figure are as follows:
[0061] CBS: a basic convolution combination module composed of a convolution layer Conv, a batch normalization layer BN and an activation function Silu.
[0062] C3: a module combined by CBS, including three CBS modules, a bottleneck module and a Concat operation. Two CBS modules constitute a bottleneck layer.
[0063] Concat: a splicing operation that splices inputs in the channel dimension and then outputs.
[0064] SPWP: a new type of pooling layer used in the embodiments of the present application, Symmetric Pyramid Weighted Pooling, which realizes maximum pooling and strip pooling at the same time, and further includes a weighting coefficient and a splicing Concat operation.
[0065] Upsample: an up-sampling operation that expands the scale of the input.
[0066] GSConv: a lighter convolution module than CBS module, which divides the input into two parts, one part is operated by normal convolution, and the other part is operated by depth separable convolution.
[0067] SA: a channel shuffle attention mechanism, which groups the input and separately calculates the attention on the space and channel, and finally splices using Concat.
[0068] Detect: a detection module that simultaneously performs classification and bounding box regression, which can receive inputs of different scales.
[0069] Specifically, the input image is transmitted in the form of a feature map in the weapon-holding target detection neural network, and the transmission process is as follows:
[0070] The first column includes a feature extraction module and a pooling module. The input image size is [640, 640, 3], which represents the width x height = 640 x 640, and the channel number = 3. After the 0th layer CBS is converted into the initial feature map, the output size is [320, 320, 64], and after the 1st-8th layer feature extraction module, the output size is [20, 20, 1024]; then input to the 9th layer SPWP for pooling, and the output size is [20, 20, 1024].
[0071] The second column and the third column are feature fusion modules. The output feature map of the 9th layer passes through the 10th layer CBS, the 11th layer Upsample module, and the output size is [40, 40, 512], which is the same size as the output feature map of the 6th layer, and the two are spliced in the channel dimension by the 12th layer concat, and the spliced size is [40, 40, 1024]; then the spliced feature map passes through the 13th layer C3, the 14th layer CBS, and the 15th layer Upsample module, and the output size is [80, 80, 256], which is the same size as the output feature map of the 4th layer, and the two are spliced in the channel dimension by the 16th layer concat, and the spliced size is [80, 80, 512]; then the spliced feature map passes through the 17th layer C3, and the output size is [80, 80, 256]; then the spliced feature map passes through the 18th layer GSConv, and the output size is [40, 40, 256], which is the same size as the output feature map of the 14th layer, and the two are spliced in the channel dimension by the 19th layer concat, and the spliced size is [40, 40, 512]; then the spliced feature map passes through the 20th layer C3, and the output size is [40, 40, 512]; finally, the spliced feature map passes through the 21st layer GSConv, and the output size is [20, 20, 512], which is the same size as the output feature map of the 10th layer, and the two are spliced in the channel dimension by the 22nd layer concat, and the spliced size is [20, 20, 1024]; then the spliced feature map passes through the 23rd layer C3 and the 24th layer SA, and the output size is [20, 20, 1024].
[0072] The fourth column is a target detection module. The outputs of the 17th, 20th and 24th layers are received, which cover feature information of shallow, middle and high layers, and can predict small, medium and large scale targets. The 1*1 convolution block of the detection layer changes the channel number of the feature map to (category number + 5) x anchor number of each detection layer, wherein “5” represents the horizontal coordinate, vertical coordinate, width, height and confidence of the prediction box. Here, we set three categories of “stick”, “knife” and “gun”, and each detection layer has 3 anchors, so the reshaped channel number is 24, that is, the Detect module receives the outputs of the 17th, 20th and 24th layers as inputs, and the output sizes are [80, 80, 24], [40, 40, 24] and [20, 20, 24] respectively. Then, according to non-maximum suppression, redundant prediction boxes are removed.
[0073] In particular, the symmetric pyramid weighted pooling module SPWP proposed in the present application refers to Figure 5 The pooling structure simultaneously adopts strip pooling and maximum pooling to form a symmetric pyramid structure. Before splicing the two kinds of pooling layers, the spliced input is weighted by using the strip pooling weight w1 and the maximum pooling weight w2 respectively.
[0074] Specifically, the symmetric pyramid weighted pooling module SPWP is located at the 9th layer in the weapon-carrying target detection neural network, receives the output from the 8th layer as input, the input size is [20, 20, 1024], and after a CBS convolution, the output size is [20, 20, 512]. Then the output enters two pooling branches, the left branch performs three times of strip pooling in turn, the size does not change after pooling, and the output after each pooling is multiplied by the weight coefficient ω1; the right branch performs three times of maximum pooling in turn, the size does not change after pooling, and the output after each pooling is multiplied by the weight coefficient ω2; there are 6 outputs on both sides, the size is [20, 20, 512], and the original output with the same size is subjected to Concat operation, and after a CBS convolution, the output size is [20, 20, 1024].
[0075] For the training of the weapon-carrying target detection neural network, a training set containing three classes of "stick", "knife" and "gun" is used for training, after the training is completed, the monitoring video stream data is input into the model for prediction, after feature extraction, pooling is performed at the SPWP structure, the pooled feature map is up-sampled in the network fusion module to obtain a multi-scale feature map, the features of multiple scales are fused, the key features are focused through the channel shuffle attention module, and finally the multi-scale feature detection result is output, including the detection frame position, class and confidence information of the target. The detection frame position and class of the target are real-time outputs of the sub-module, and the confidence level information carrying the threat ability is transmitted to step S500 for threat assessment and early warning.
[0076] Further, in some embodiments of the present application, in step S300, as shown in Figure 7 The suspicious behavior detection process includes target frame segmentation, generating a predicted frame, jointly loss optimizing the generated predicted frame based on diffusion frame constraint and optical flow frame constraint, comparing the prediction error and setting an abnormal threshold based on the abnormal score curve, and performing behavior judgment based on the comparison result and the like.
[0077] Specifically, in some embodiments of the present application, in step S300, the picture frame data in the multi-source situation data is input into a pre-trained suspicious behavior detection neural network for suspicious behavior detection to screen the behavior suspicious targets in the detection task area; wherein the suspicious behavior detection neural network predicts the predicted frame corresponding to the N+1th picture frame from the continuous N picture frames, thereby generating the predicted frame with time sequence, calculating the error between the predicted frame and the picture frame corresponding to the same time point, and determining the behavior suspicious target and its threat intention corresponding to the time point according to whether the error exceeds the preset threshold.
[0078] In some embodiments of the present application, the suspicious behavior detection neural network is a dual constraint comparison detection network, which adopts optical flow frame constraint and diffusion frame constraint to jointly optimize the network loss of the suspicious behavior detection neural network.
[0079] Specifically, the construction of the suspicious behavior detection neural network is to improve the U-Net network structure into a two-hop memory U-Net, and then construct a dual constraint comparison detection network (DCC-UNet) based on the two-hop memory U-Net, which is used for identifying the suspicious behavior of the target in the task area.
[0080] The two-hop memory U-Net network structure is as shown in Figure 8 The size of the input data key picture frame is set to width x height = 256 x 256, and the initial channel number of the RGB image is 3, that is, the input size is [256, 256, 3]. The first convolution downsampling module downblock1 first performs two convolution operations to expand the input in the channel dimension, and the convolution output size is [256, 256, 64], and then one downsampling is performed, and the downsampling output size is [128, 128, 64]. The second convolution downsampling module downblock2 first performs two convolution operations, and the convolution output size is [128, 128, 128], and then one downsampling is performed, and the downsampling output size is [64, 64, 128]. The third convolution downsampling module downblock3 first performs two convolution operations, and the convolution output size is [64, 64, 256], and then one downsampling is performed, and the downsampling output size is [32, 32, 256].
[0081] Mem is a memory module that stores feature maps and is used for addressing. The output of the third convolution downsampling module downblock3 is first subjected to one convolution operation, and the size becomes [32, 32, 512], and then the addressed feature map is subjected to one convolution operation and output, and the size remains [32, 32, 512], which is used as the input of the third upsampling convolution module upblock3.
[0082] The third upsampling convolutional module, upblock3, receives a feature map with input size [32, 32, 512]. It first performs an upsampling operation, resulting in an upsampled output size of [64, 64, 512]. Then, the convolutional output of the third downsampling module, downblock3, with a size of [64, 64, 256], is concatenated using skip connections. Next, two convolutional operations are performed, resulting in a convolutional output size of [64, 64, 256]. The second upsampling convolutional module, upblock2, receives a feature map with input size [64, 64, 256]. It first performs an upsampling operation... The sampled output size is [128, 128, 256]. The convolutional output of downblock2 with size [128, 128, 128] is then concatenated via skip connections, followed by two convolutional operations, resulting in a convolutional output size of [128, 128, 128]. The first upsampling convolutional module upblock1 receives a feature map with input size [128, 128, 128]. It first performs an upsampling operation, resulting in an upsampling output size of [256, 256, 128]. Then, it performs two convolutional operations, resulting in a convolutional output size of [256, 256, 64]. Finally, the output of the first upsampling convolutional module upblock1 undergoes one more convolutional operation, reducing its size to [256, 256, 3].
[0083] The dual-path constraint contrast detection network DCC-UNet, built upon the aforementioned two-hop memory U-Net, is as follows: Figure 9 As shown. The pre-training process is as follows: The training data consists of image frames from a video segment divided in chronological order, called ground truth frames, denoted as I1, I2, ..., I... t ,I t+1 , I1,I2,…,I t Input the two-hop memory U-Net, and generate the predicted (t+1)th frame through encoding and decoding. These are called prediction frames. In this process, optical flow frame constraints and diffusion frame constraints are used to jointly optimize the network loss.
[0084] Among them, optical flow frame constraint refers to: predicting the frame And the real frame of the previous moment I t Optical flow is calculated using the Flownet module, and the corresponding real frame I is used. t+1 And the real frame from the previous moment I t Optical flow is calculated using the Flownet module, and its consistency is constrained using L1 regularization loss.
[0085] Diffusion frame constraint refers to: from the real frame I t+1 Latent variables are generated, and through the encoding and decoding process of a pre-trained diffusion U-Net model, diffusion frames with structural feature constraints are generated. Diffusion Frame With the predicted frame Both are generated images after the coding and decoding process, so they should be aligned in structural features to make the model prediction more accurate, and L2 regularization loss is used to constrain their consistency.
[0086] Finally, under the action of the dual-channel constraint, the predicted frame and the real frame I t+1 are compared by the discriminator to determine the authenticity, and the dual-channel constraint comparison detection network DCC-UNet is trained as a good model that can accurately predict the next frame.
[0087] After the training is completed, the process of suspicious behavior detection using the pre-trained suspicious behavior detection neural network includes:
[0088] The data input is obtained, and the data is preprocessed. The input data is usually the situation monitoring video of the task area (such as a terrorist battlefield), and when the video is detected and segmented, the frame strategy can be selected according to the application requirements, such as directly extracting each frame image, and in the case of low delay, the video can be frame extracted.
[0089] The above frame data is input into the trained dual-channel constraint comparison detection network DCC-UNet, and the next frame is predicted using 5 consecutive frames, and the predicted frame with time sequence is generated.
[0090] The error between the predicted frame and the real frame is calculated, and the PSNR value is used to represent it, which is normalized to the interval [0, 1], converted into an abnormal score curve and used as the output of the sub-module. At the same time, according to the experience value, the threshold value of the abnormal score is set, and the frame exceeding the threshold value is regarded as suspicious behavior, and the judgment of the sub-module is given. The normalized abnormal score is used as the confidence, and the hierarchical information of the threat intention is transmitted to step S500 for threat assessment and early warning.
[0091] Further, in some embodiments of the present application, step S400, the target space-time trajectory data in the multi-source situation data is input into the pre-trained target yaw trajectory detection neural network to detect the target yaw trajectory and screen the yaw target in the detection task area; wherein the target yaw trajectory detection neural network is a residual BiLSTM neural network.
[0092] Specifically, as Figure 10 As shown, the target yaw trajectory detection implementation process is: obtaining the trajectory data of the target, first preprocessing, that is, data segmentation, eliminating some repeated values, drift values and the like; constructing the residual BiLSTM model, and pre-training the model; using the trained model to predict the trajectory data, taking the prediction result and the error between the predicted value and the actual value as the output, performing abnormal trajectory determination, setting the error threshold value according to the average error value of the pre-trained model, and regarding the trajectory point as yaw if the threshold value is exceeded, and giving the trajectory abnormality determination. The error between the above predicted value and the actual value is normalized to [0, 1] as a confidence degree, and the layered information carrying threat opportunity is transmitted to step S500 for threat assessment and early warning.
[0093] The detection perception device can obtain the continuous trajectory of some targets in the current detection area, including coordinate information such as longitude and latitude, speed and distance information and the like. Using these space-time data, the trajectory of the target can be predicted, the BiLSTM model is trained, so that it can learn the rule of the first k trajectory points, and the k+1 trajectory point is predicted. When the target real trajectory and the predicted trajectory have a large deviation, or the target has a distance or speed mutation, it is considered that the target has a threat opportunity.
[0094] The traditional recurrent neural network is prone to "gradient disappearance" when processing time series tasks, and cannot solve the long-term dependence in the trajectory sequence. The long short-term memory (LSTM) network uses unique gating logic to process long-term memory in time series tasks, and can effectively alleviate the problem of gradient disappearance.
[0095] Referring to Figure 11 , the LSTM network uses "gates" to control the amount of information passing through. Among them, C t represents the current time LSTM cell state, x t represents the current input, h t represents the current output, f t , i t and o t are respectively the three most important "gates" in LSTM, the forgetting gate, the input gate and the output gate, represents the addition of two vectors, represents the vector cross product.
[0096] The Bi-LSTM is adopted, and the advantage is that both the forward features of the input sequence and the backward features can be captured. The influence of the trajectory on the future and the influence of the future on the present can be captured by the model at the same time, increasing the reliability of the extracted time features. The residual connection structure is adopted, so that the output result of the first layer Bi-LSTM is connected to the output of the second layer in the form of residual, and it is ensured that the multi-layer structure will not cause the prediction effect to decrease.
[0097] Specifically, in some embodiments of the present application, as Figure 12As shown, the target yaw trajectory detection neural network adopts a residual double BiLSTM network, including an input layer composed of LSTM modules, an intermediate layer composed of LSTM modules, a full connection layer, and an activation layer, the LSTM modules in the input layer and the intermediate layer are connected to each other, the output end of the LSTM module of the input layer and the input end of the LSTM module of the intermediate layer are connected one by one, the output end of the LSTM module of the input layer is also connected to the output end of the corresponding LSTM module of the intermediate layer, the intermediate layer is connected to the full connection layer, and the full connection layer is connected to the activation layer.
[0098] Further, the single-step prediction using LSTM is more accurate, but only one future value is predicted at a time. In some embodiments of the present application, multi-step prediction can be used to predict n future values, but since historical data and predicted data are used, the prediction error will increase over time, as shown in Figure 13 .
[0099] Considering the dynamic variability of application scenarios, preferably, in some embodiments of the present application, a dynamic time step prediction mechanism is used to predict trajectory data using a trained model, which alleviates the problems of LSTM error accumulation and short perceived time of single-step prediction.
[0100] Specifically, referring to Figure 14 , let the current time stamp be t, the historical data update time interval be T, and a variable k value be selected as the number of prediction data at the current time, which is dynamically changed. The variable k value is dynamically selected according to the time stamp t and whether an early warning information is issued within the time period T starting from the last time stamp t-T. When a trajectory anomaly is predicted within the last time period T, according to the fact correlation, it is indicated that the abnormality is likely to occur again in this period, and single-step prediction is used for updating and predicting until k tY data are predicted to ensure accuracy; when no trajectory anomaly is predicted within the last time period T, it is indicated that the situation is stable in this period, but the situation needs to be observed for a longer time, and multi-step prediction is used to predict k tN data at a time to search for abnormal conditions as much as possible. In the next T time period from time stamp t to time stamp t+T, the historical data is updated, and k tN data are predicted again.
[0101] Further, in some embodiments of the present application, a simple implementation of step S500 is as follows Figure 15As shown, specifically includes: using the simple results of step S200-step S400 to carry out threat assessment and early warning, that is, dividing the detection results according to threat factors. When it is detected that a person holds a weapon such as a knife, a gun, a stick, etc., it is considered that the target has threat ability, and 1 threat ability index is accumulated. When it is detected that the inter-frame abnormal score exceeds the set threshold, it is considered that suspicious behavior occurs, and the target has threat intention, and 1 threat intention index is accumulated. When it is detected that the target real trajectory and the predicted trajectory have a large deviation, or the target has a distance and speed mutation, it is considered that the target has threat opportunity, and 1 threat opportunity index is accumulated.
[0102] According to the cumulative result of the threat index, the warning level is determined, when the cumulative index = 3, it is determined as a severe threat, and the first-level warning is issued; when the cumulative index = 2, it is determined as a moderate threat, and the second-level warning is issued; when the cumulative index = 1, it is determined as a mild threat, and the third-level warning is issued.
[0103] Preferably, in other embodiments of the present application, in order to comprehensively and accurately carry out threat assessment and early warning, step S200-step S400 further includes:
[0104] The result of step S200 output multi-scale feature detection includes the detection frame position, category and confidence information of the target. The detection frame position and category of the target are the real-time output of the sub-module, and the confidence carries the hierarchical information of the threat ability and is transmitted to step S500 for threat assessment and early warning.
[0105] In step S300, the error between the predicted frame and the real frame is calculated, here the PSNR value is used to represent, the PSNR value is normalized to [0, 1] interval, converted into an abnormal score curve and used as the output of the sub-module. The normalized abnormal score is used as the confidence, carries the hierarchical information of the threat intention and is transmitted to S500 for threat assessment and early warning.
[0106] In step S400, the predicted future result is used as the output of the sub-module, and the error between the predicted value and the actual value is calculated. The error between the predicted value and the actual value is normalized to [0, 1] and used as the confidence to carry the hierarchical information of the threat opportunity and is transmitted to step S500 for threat assessment and early warning.
[0107]
[0108] In order to fully utilize the hierarchical information carried by the confidence transmitted by step S200-step S400, it is stipulated that when the confidence of each step falls into the confidence interval corresponding to the first column in the table, it is revalued with an additive factor, such as when the confidence of step S200 falls into [0, α1), it is valued with x1 factor, if the confidence of step S200 falls into [α1, α2), it is valued with x2 factor, and so on.
[0109] According to the threat level regulations of this application scenario, y1 < z1 < x1 < y2 < z2 < x2 < y3 < z3 < x3. Based on the confidence levels detected by different sub-modules, these factors are combined additively to classify the warning levels.
[0110] Set the confidence levels α1 = 0.5 and α2 = 0.8. A feasible matrix is as follows:
[0111]
[0112] The warning thresholds are: [0, 9] for low risk, mild threat, level III warning; (9, 15] for medium risk, moderate threat, level II warning; (15, 24] for high risk, severe threat, level I warning.
[0113] The present invention uses the data obtained from multi-source detectors for situation awareness and anomaly detection, can timely discover the potential threats of targets, analyze the threats of abnormal activities and issue accurate warnings, comprehensively master the movements of suspicious targets, realizes the real-time awareness of the on-site situation by decision-makers, and improves the crisis response and emergency disposal capabilities in anti-terrorism and anti-riot decision-making.
[0114] The present invention can, in view of the complex terrain in the detection task area, the variety and quantity of targets, and the diverse movement types of targets, etc., adopt target perception data such as surveillance video streams, picture frames, and coordinate point trace matrices for anomaly detection, form a complementary situation, and more comprehensively master the movements of suspicious targets.
[0115] The present invention establishes and realizes a multi-level threat assessment according to real-time analysis and detection, can real-time master the key information of the threat level of targets, and thus effectively assist decision-makers in allocating disposal resources.
[0116] In the second aspect, as Figure 16 shown, the warning device based on anti-terrorism and anti-riot anomaly detection in an embodiment of the present invention includes:
[0117] A data acquisition module, configured to acquire multi-source situation data of the detection task area. The multi-source situation data is multi-modal data obtained by multiple detection and perception devices detecting the same detection task area, and the multi-source situation data includes video stream data, picture frame data, and target spatio-temporal trajectory data;
[0118] An armed target screening module, configured to use the video stream data in the multi-source situation data to detect armed targets, screen the armed targets in the detection task area, and obtain the armed target screening result;
[0119] A suspicious target screening module, configured to use the picture frame data in the multi-source situation data to detect suspicious behaviors, screen the behavior-suspicious targets in the detection task area, and obtain the behavior-suspicious target screening result;
[0120] The yaw target screening module is configured to detect a yaw track of a target by using the target space-time track data in the multi-source situation data, screen a yaw target in the detection task area, and obtain a yaw target screening result.
[0121] The threat judgment and early warning module is configured to judge a threat level of the detection task area based on the armed target screening result, the suspicious behavior target screening result, and the yaw target screening result, and perform early warning based on the threat level.
[0122] In a third aspect, an electronic device is provided in an embodiment of the present application, which comprises at least one processor and a memory connected to the at least one processor in communication, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the early warning method based on the anti-terrorism and anti-violence anomaly detection. Figure 17
[0123] The memory and the processor are connected in a bus mode, the bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripheral devices, voltage stabilizers, and power management circuits, etc. together, which are well known in the art. The interface provides an interface between the bus and the transceiver, such as a communication interface, a user interface. The transceiver can be one element or multiple elements, such as multiple receivers and transmitters, which provide a unit for communicating with various other devices on a transmission medium. The data processed by the processor is transmitted on a wireless medium through an antenna, further, the antenna also receives data and transmits the data to the processor.
[0124] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory can be used to store data used by the processor during operation.
[0125] In a fourth aspect, a computer readable storage medium is provided in an embodiment of the present application, which stores a computer program, and the computer program is executed by a processor to implement the early warning method based on the anti-terrorism and anti-violence anomaly detection.
[0126] The existing abnormality detection technology is often directed to normal monitoring areas such as private houses, shopping malls, public facilities and the like, and the detection means adopted is not targeted to emergency events, and cannot accurately obtain various threat factors in a violent and terrorist emergency scene; the detection means is single, lacks complementation and linkage, cannot form a unified and integrated high-level criterion for multi-level threat factors, cannot integrate situation information for system decision, and is highly dependent on manual inspection and special monitoring personnel, and has slow response speed and high resource consumption.
[0127] Therefore, the present application utilizes multi-source situation data to form mutual corroboration or complementation, solves efficient and accurate abnormal target detection in a violent and terrorist scene, more comprehensively masters suspicious target trends, realizes real-time perception of the scene situation, and improves crisis response and disposal capability in violent and terrorist decision-making.
[0128] The present application coordinates multiple neural network functions, establishes a multi-level threat evaluation system architecture, performs threat evaluation and early warning through confidence, masters key information of target threat degree, and can effectively assist decision-making personnel in allocating disposal resources.
[0129] Those skilled in the art can understand from the above description that all or part of the steps in the above-mentioned embodiment methods can be completed by programs instructing related hardware, the programs are stored in a storage medium, and include a plurality of instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes but is not limited to a U disk, a mobile hard disk, a magnetic memory, an optical memory and various program code storage media.
[0130] In several embodiments provided in the present application, it should be understood that the disclosed system, device or method can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the modules / units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of modules or units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed each other can be indirect coupling or communication connection through some interfaces, devices or modules or units, and can be electrical, mechanical or other forms.
[0131] The modules / units described as separated parts can or can not be physically separated, and the parts shown as modules / units can or can not be physical modules, that is, can be located in one place, or can be distributed to multiple network units. Part or all of the modules / units can be selected according to actual needs to achieve the purpose of the embodiments of the present application. For example, the functional modules / units in each embodiment of the present application can be integrated in one processing module, or each module / unit can be physically present alone, or two or more modules / units can be integrated in one module / unit.
[0132] Those of ordinary skill in the art should further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software, or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0133] The above description of the flow or structure of each figure has its own emphasis, and the parts not described in detail in a certain flow or structure can refer to the related description of other flows or structures.
[0134] The above embodiments only exemplarily illustrate the principles and effects of the present application, and are not used to limit the present application. Any person skilled in the art can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes completed by those skilled in the art without departing from the spirit and technical idea of the present application should be covered by the claims of the present application.
Claims
1. An early warning method based on anomaly detection in counter-terrorism and riot control, characterized in that, include: Acquire multi-source situational data of the detection task area. The multi-source situational data is multimodal situational data obtained by multiple detection and sensing devices detecting the same detection task area. The multi-source situational data includes video stream data, image frame data, and target spatiotemporal trajectory data. Weaponized targets are detected by using video stream data from multi-source situational data, and weaponized targets in the detection task area are screened to obtain weaponized target screening results. Suspicious behavior detection is performed using image frame data from multi-source situational data. Suspicious targets in the detection task area are then filtered to obtain the results of the suspicious behavior target filtering. Target yaw trajectory detection is performed using target spatiotemporal trajectory data from multi-source situational data. Yaw targets in the detection mission area are then screened to obtain yaw target screening results. Based on the results of the screening of armed targets, the results of the screening of suspicious targets, and the results of the screening of yaw targets, the threat level of the detection task area is determined, and an early warning is issued based on the threat level. Video stream data from multi-source situational data is input into a pre-trained neural network for armed target detection to screen armed targets in the detection task area; wherein, the armed target detection includes multi-layer feature extraction, pooling, upsampling, multi-scale feature fusion, key feature attention and multi-scale feature detection; The weapon-carrying target detection neural network includes a feature extraction module, a pooling module, a feature fusion module, and a target detection module. Video stream data from multi-source situational data is input to the feature extraction module for multi-layer feature extraction. One output of the feature extraction module is input to the feature fusion module, and the other output is pooled by the pooling module and then input to the feature fusion module. The feature fusion module upsamples the two inputs, performs multi-scale feature fusion, and focuses on key features. The output of the feature fusion module is input to the target detection module, which performs multi-scale feature detection and outputs the detection result of the weapon-carrying target.
2. The method as described in claim 1, characterized in that, The pooling methods include max pooling, average pooling, global average pooling, and random pooling.
3. The method as described in claim 1, characterized in that, The pooling module includes a first convolution module, a second convolution module, a first strip pooling module, a second strip pooling module, a third strip pooling module, a first max pooling module, a second max pooling module, a third max pooling module, a first concatenation module, and a second concatenation module; wherein, The output of the feature extraction module is input into the first convolution module; the first output of the first convolution module is directly input into the first concatenation module; the second output of the first convolution module is subjected to three strip pooling operations through the first strip pooling module, the second strip pooling module, and the third strip pooling module; and the third output of the first convolution module is subjected to three max pooling operations through the first max pooling module, the second max pooling module, and the third max pooling module. The outputs of the first and second strip pooling modules are multiplied by the first weight and then input into the first concatenation module. The outputs of the first and second max pooling modules are multiplied by the second weight and then input into the first concatenation module. The first concatenation module concatenates the five inputs and then inputs the result into the second concatenation module. The output of the third strip pooling module is multiplied by the first weight and then input into the second concatenation module. The output of the third max pooling module is multiplied by the second weight and then input into the second concatenation module. The second concatenation module concatenates the three inputs and then inputs them into the second convolution module. The second convolution module performs convolution and outputs the final pooling result.
4. The method as described in claim 1, characterized in that, Image frame data from multi-source situational data is input into a pre-trained suspicious behavior detection neural network to detect suspicious behavior and filter suspicious targets in the detection task area. The suspicious behavior detection neural network predicts the corresponding N+1th image frame for N consecutive image frame inputs, thereby generating prediction frames with time order. The error between the prediction frame and the image frame corresponding to the same time point is calculated, and the suspicious target and its threat intent corresponding to the time point are determined based on whether the error exceeds a preset threshold.
5. The method as described in claim 4, characterized in that, The suspicious behavior detection neural network is a dual-path constraint comparison detection network, which uses optical flow frame constraints and diffusion frame constraints to jointly optimize the network loss of the suspicious behavior detection neural network.
6. The method as described in claim 5, characterized in that, The construction of the suspicious behavior detection neural network involves first improving the U-Net network structure to a two-hop memory U-Net, and then constructing a dual-path constraint comparison detection network based on the two-hop memory U-Net.
7. The method as described in claim 1, characterized in that, The target spatiotemporal trajectory data from multi-source situational data is input into a pre-trained target yaw trajectory detection neural network to detect target yaw trajectories and filter out yaw targets in the detection task area; wherein, the target yaw trajectory detection neural network is a residual BiLSTM neural network.
8. An early warning device based on anti-terrorism and anti-riot anomaly detection, comprising: The data acquisition module is used to acquire multi-source situational data of the detection task area. The multi-source situational data is multimodal data obtained by multiple detection and sensing devices detecting the same detection task area. The multi-source situational data includes video stream data, image frame data and target spatiotemporal trajectory data. The weapon-carrying target screening module is used to detect weapon-carrying targets using video stream data from multi-source situational data, screen weapon-carrying targets in the detection task area, and obtain weapon-carrying target screening results. The suspicious target screening module is used to detect suspicious behavior by using image frame data from multi-source situational data, and to screen suspicious targets in the detection task area to obtain the suspicious target screening results. The yaw target screening module is used to detect target yaw trajectories by using target spatiotemporal trajectory data from multi-source situational data, and to screen yaw targets in the detection task area to obtain yaw target screening results. The threat assessment and early warning module is used to determine the threat level of the detection task area based on the screening results of armed targets, the screening results of suspicious targets, and the screening results of yaw targets, and to issue an early warning based on the threat level. Video stream data from multi-source situational data is input into a pre-trained neural network for armed target detection to screen armed targets in the detection task area; wherein, the armed target detection includes multi-layer feature extraction, pooling, upsampling, multi-scale feature fusion, key feature attention and multi-scale feature detection; The weapon-carrying target detection neural network includes a feature extraction module, a pooling module, a feature fusion module, and a target detection module. Video stream data from multi-source situational data is input to the feature extraction module for multi-layer feature extraction. One output of the feature extraction module is input to the feature fusion module, and the other output is pooled by the pooling module and then input to the feature fusion module. The feature fusion module upsamples the two inputs, performs multi-scale feature fusion, and focuses on key features. The output of the feature fusion module is input to the target detection module, which performs multi-scale feature detection and outputs the detection result of the weapon-carrying target.
9. An electronic device, comprising: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the early warning method based on anti-terrorism and anti-riot anomaly detection as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, wherein, When the computer program is executed by the processor, it implements the early warning method based on anti-terrorism and anti-riot anomaly detection as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Power grid safety early warning method and device, computer equipment and storage medium
CN115620208A