A training method, system and computer storage medium for a violent terrorist knife and gun detection model

By adding feature map pyramid network to the swin-transformer feature extraction network, the problem of low detection accuracy of violent terrorist knife and gun pictures in the existing technology is solved, and high-precision detection of pictures on the network social platform is achieved, with the detection accuracy reaching 90%.

CN114972283BActive Publication Date: 2025-06-20CHENGDU BENYING INTERACTIVE ENTERTAINMENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210642623.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-08
Publication Date
2025-06-20
Estimated Expiration
2042-06-08

AI Technical Summary

Technical Problem

In the prior art, the accuracy of detecting and positioning of violent terrorist swords and guns through the swin-transformer model is not high enough, and this problem cannot be effectively solved.

Method used

A feature map pyramid network is added to the swin-transformer feature extraction network, and the training set is model trained through the updated network, and verified through the verification set and the test set until the preset accuracy and recall thresholds are reached, and the final terrorist knife and gun detection model is obtained.

Benefits of technology

Through this method, the images on the online social platform can be accurately detected by violent terrorist swords and guns, with a detection accuracy of 90%, and specific position coordinates are given, effectively solving the problems of insufficient manual review efficiency and increased image volume.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972283B_ABST
    Figure CN114972283B_ABST
Patent Text Reader

Abstract

The present invention discloses a training method, system and computer storage medium for a violent terrorist knife and gun detection model. Among them, the method includes: obtaining a set of violent terrorist knife and gun pictures, and batch-labeling their position and category information; after labeling, dividing them into a training set, a validation set and a test set; performing data augmentation on the training set; adding a feature pyramid network to the hierarchical vision (swin-transformer) feature extraction network to obtain an updated swin-transformer feature extraction network, training the model with the data-augmented training set on the updated network and testing it through the validation set and the test set to obtain the accuracy rate and recall rate of the test set. When both are greater than their respective preset thresholds, the final violent terrorist knife and gun detection model is obtained. On the contrary, a large number of violent terrorist knife and gun pictures are detected, and the pictures with good detection effects are selected and put into the training set for repeated training. Through this method, pictures on network social platforms can be accurately detected for violent terrorist knives and guns, and the detection accuracy reaches 90%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and in particular, to a method and system for training a violent terrorist knife and gun detection model and a computer storage medium. Background Art

[0002] Object detection is an important computer vision task for a certain category (such as people, animals or cars) in digital images.

[0003] Around 2000, most of the proposed methods were based on sliding windows and artificial feature extraction for object detection, which had the defects of high computational complexity and poor robustness in complex scenarios. In 2014, the R-CNN algorithm was proposed, which uses deep learning technology to automatically extract hidden features in the input image and perform higher-precision classification and prediction on samples. After R-CNN, many deep learning-based image object detection algorithms such as Fast R-CNN, Faster R-CNN, SPPNet, and YOLO series emerged. Currently, CNN has reached a bottleneck, and it is difficult to greatly improve the performance of the model through new structures.

[0004] In 2020, Vision Transformer based on the self-attention mechanism successfully applied the Transformer model used in the NLP field to image classification in the CV field and achieved an accuracy of 88.55% on the ImageNet dataset. Compared with the convolutional network CNN, it has achieved a great improvement in image classification accuracy. Since then, the Transformer network has officially entered the era of dominating the CV field.

[0005] Now there are many violent terrorist knife and gun pictures on the network platform. In order to prevent violent terrorist knife and gun pictures from spreading on the social network and thus maintain the green and healthy content of the Internet platform, the accuracy of detecting and locating violent terrorist knife and gun pictures only through the swin-transformer model is not high enough.

[0006] Aiming at the problem that the accuracy of detecting and locating violent terrorist knife and gun pictures only through the swin-transformer model in the prior art is not high enough, no effective solution has been proposed yet. Summary of the Invention

[0007] An embodiment of the present invention provides a method and system for training a violent terrorist knife and gun detection model and a computer storage medium to solve the problem that the accuracy of detecting and locating violent terrorist knife and gun pictures only through the swin-transformer model in the prior art is not high enough.

[0008] To achieve the above object, on the one hand, the present invention provides a method for training a violent terrorist knife and gun detection model, the method comprising: Step S101, obtaining a target violent terrorist knife and gun picture set, randomly extracting a preset number of pictures from the target violent terrorist knife and gun picture set as the first target violent terrorist knife and gun picture set; taking the remaining pictures as the second target violent terrorist knife and gun picture set; performing position annotation and category annotation on the first target violent terrorist knife and gun picture set; Step S102, dividing the annotated first target violent terrorist knife and gun picture set into a training set, a test set and a validation set; Step S103, adding a feature pyramid network to the hierarchical vision (swin-transformer) feature extraction network to obtain an updated swin-transformer feature extraction network; inputting the training set into the updated swin-transformer feature extraction network for model training, and verifying through the validation set to obtain a primary violent terrorist knife and gun detection model; testing the test set according to the primary violent terrorist knife and gun detection model to obtain the accuracy rate and recall rate of the test set detection; Step S104, when it is determined that both the accuracy rate and the recall rate are greater than their respective preset thresholds, obtaining a final violent terrorist knife and gun detection model.

[0009] Optionally, when it is determined that the accuracy rate or the recall rate does not reach their respective preset thresholds, go to Step S105; Step S105, detecting the second target violent terrorist knife and gun picture set according to the primary violent terrorist knife and gun detection model to obtain correct pictures, incorrect pictures, and the position information and category information of the correct pictures; inputting the correct pictures into the first target violent terrorist knife and gun picture set and dividing it into an updated training set, an updated test set and an updated validation set; training the model according to the primary violent terrorist knife and gun detection model on the updated training set and verifying through the updated validation set to obtain an updated violent terrorist knife and gun detection model; Step S106, taking the updated violent terrorist knife and gun detection model as the primary violent terrorist detection model, taking the incorrect pictures as the second target violent terrorist knife and gun picture set, repeating the above Step S105 until both the accuracy rate and the recall rate of the test set detection are greater than their respective preset thresholds, obtaining a final violent terrorist knife and gun detection model.

[0010] Optionally, the adding a feature pyramid network to the hierarchical vision (swin-transformer) feature extraction network to obtain an updated swin-transformer feature extraction network includes: extracting features of a sample picture through the backbone network of the swin-transformer feature extraction network to obtain four feature layers; merging and connecting the four feature layers to obtain three feature maps; respectively sending the three feature maps into a convolutional layer for convolution to obtain the updated swin-transformer feature extraction network.

[0011] Optionally, the calculation formula for merging and connecting the four feature layers to obtain three feature maps is as follows:

[0012] F1 = C(P2, P3, P4, P5) = P2 || Upx2(P3) || Upx4(P4) || Upx8(P5);

[0013] F2 = C(P3, P4, P5) = Upx2(P3) || Upx4(P4) || Upx8(P5);

[0014] F3 = C(P4, P5) = Upx4(P4) || Upx8(P5);

[0015] Among them, P2 is the first feature layer; P3 is the second feature layer; P4 is the third feature layer; P5 is the fourth feature layer; F1 is the first feature map; F2 is the second feature map; F3 is the third feature map; Upx2 is upsampling by a factor of 2; Upx4 is upsampling by a factor of 4; Upx8 is upsampling by a factor of 8.

[0016] Optionally, the obtaining of the target terrorist knife and gun picture set includes: obtaining the original terrorist knife and gun picture set; screening the original terrorist knife and gun picture set; and manually reviewing and selecting the screened terrorist knife and gun picture set to obtain the target terrorist knife and gun picture set.

[0017] Optionally, after the step S102, it includes: performing data augmentation on the divided training set; the data augmentation includes: homogeneous augmentation and heterogeneous augmentation; the heterogeneous augmentation includes: mixup data augmentation and mosaic data augmentation.

[0018] On the other hand, the present invention provides a training system for a violent terrorist knife and gun detection model. The system includes: an acquisition unit, configured to acquire a target violent terrorist knife and gun picture set, randomly extract a preset number of pictures from the target violent terrorist knife and gun picture set as the first target violent terrorist knife and gun picture set; use the remaining pictures as the second target violent terrorist knife and gun picture set; perform position annotation and category annotation on the first target violent terrorist knife and gun picture set; a division unit, configured to divide the labeled first target violent terrorist knife and gun picture set into a training set, a test set, and a validation set; a model training unit, configured to add a feature pyramid network to a hierarchical vision (swin-transformer) feature extraction network to obtain an updated swin-transformer feature extraction network; input the training set into the updated swin-transformer feature extraction network for model training, and perform validation through the validation set to obtain a primary violent terrorist knife and gun detection model; test the test set according to the primary violent terrorist knife and gun detection model to obtain the accuracy rate and recall rate of the test set detection; a judgment unit, configured to obtain a final violent terrorist knife and gun detection model when it is determined that both the accuracy rate and the recall rate are greater than their respective preset thresholds.

[0019] Optionally, when it is determined that the accuracy rate or the recall rate does not reach their respective preset thresholds, enter a model update unit; the model update unit is configured to detect the second target violent terrorist knife and gun picture set according to the primary violent terrorist knife and gun detection model to obtain correct pictures, incorrect pictures, and the position information and category information of the correct pictures; input the correct pictures into the first target violent terrorist knife and gun picture set and divide it into an updated training set, an updated test set, and an updated validation set; perform model training on the updated training set according to the primary violent terrorist knife and gun detection model and perform validation through the updated validation set to obtain an updated violent terrorist knife and gun detection model; a repeated training unit, configured to use the updated violent terrorist knife and gun detection model as the primary violent terrorist detection model, use the incorrect pictures as the second target violent terrorist knife and gun picture set, and repeat the model update unit until both the accuracy rate and the recall rate of the test set detection are greater than their respective preset thresholds to obtain a final violent terrorist knife and gun detection model.

[0020] Optionally, it further includes: a data augmentation unit, configured to perform data augmentation on the divided training set; the data augmentation includes: homogeneous augmentation and heterogeneous augmentation; the heterogeneous augmentation includes: mixup data augmentation and mosaic data augmentation.

[0021] On the other hand, the present invention also provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the above-mentioned violent terrorist knife and gun detection model training method.

[0022] Advantages of the present invention:

[0023] The present invention provides a method, a system and a computer storage medium for training a terrorist knife and gun detection model. Among them, the method provides an improved version of the swin-transformer for terrorist knife and gun picture detection. Specifically, a feature pyramid network is added to the hierarchical vision (swin-transformer) feature extraction network to obtain an updated swin-transformer feature extraction network. Through this method, it is possible to accurately detect terrorist knives and guns in pictures on the network social platform, with a detection accuracy of 90%, and give the specific position coordinates of the terrorist knife and gun pictures. It can determine whether the picture is green and healthy content, effectively solving the problems of insufficient efficiency in manual picture review and the consumption of manpower and financial resources caused by the increasing amount of pictures. Brief Description of the Drawings

[0024] Figure 1 is a flowchart of a method for training a terrorist knife and gun detection model provided by an embodiment of the present invention;

[0025] Figure 2 is a schematic structural diagram of a system for training a terrorist knife and gun detection model provided by an embodiment of the present invention. Detailed Embodiments

[0026] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0027] Now there are many terrorist knife and gun pictures on the network platform. In order to prevent terrorist knife and gun pictures from spreading on the social network and thus maintain the green and healthy content of the Internet platform, the accuracy of detecting and positioning terrorist knife and gun pictures only by the swin-transformer model is not high enough.

[0028] Therefore, the present invention provides a method for training a terrorist knife and gun detection model, Figure 1 is a flowchart of a method for training a terrorist knife and gun detection model provided by an embodiment of the present invention. As Figure 1 shown, the method includes:

[0029] Step S101, obtain a target terrorist knife and gun picture set, randomly select a preset number of pictures from the target terrorist knife and gun picture set as the first target terrorist knife and gun picture set; use the remaining pictures as the second target terrorist knife and gun picture set; perform position annotation and category annotation on the first target terrorist knife and gun picture set;

[0030] In an alternative embodiment, the obtaining of the target terrorist knife and gun picture set includes:

[0031] S1011. Obtain the original terrorist knife and gun picture set;

[0032] Specifically, since the number of terrorist picture sets is small, it is necessary to crawl the original terrorist knife and gun picture set from multiple channels (such as Google, Baidu) on the Internet.

[0033] S1012. Screen the original terrorist knife and gun picture set;

[0034] Specifically, there are many types of picture sets crawled. At this time, pictures in gif format and single-channel (8bit) pictures need to be removed, and at the same time, pictures with too small resolution crawled are removed.

[0035] S1013. Manually review and select the screened terrorist knife and gun picture set to obtain the target terrorist knife and gun picture set.

[0036] After screening, the picture data format is relatively clean. Then, the target terrorist knife and gun picture set is selected through manual review (that is, the original terrorist knife and gun picture set crawled from the Internet will be mixed with pictures that do not belong to terrorist knives and guns. Therefore, it is necessary to further select through manual review to ensure that each picture is a terrorist knife and gun picture, and all terrorist knife and gun pictures are combined to obtain the target terrorist knife and gun picture set).

[0037] After obtaining the target terrorist knife and gun picture set, randomly select a preset number of pictures from the target terrorist knife and gun picture set as the first target terrorist knife and gun picture set; use the remaining pictures as the second target terrorist knife and gun picture set; perform position annotation and category annotation on the first target terrorist knife and gun picture set;

[0038] Specifically, in the present invention, after obtaining the target terrorist knife and gun picture set, it is necessary to extract a preset number (a small batch) of pictures from the target terrorist knife and gun picture set as the first target terrorist knife and gun picture set, and then use the lableimg annotation software to manually annotate the first target terrorist knife and gun picture set to obtain the position information and category information of each picture in the first target terrorist knife and gun picture set. Since manual annotation requires a lot of time, manpower and energy, the number of the first target terrorist knife and gun picture sets extracted is very small. For example: if the target terrorist knife and gun picture set is 10,000 pictures, 1,000 pictures are extracted as the first target terrorist knife and gun picture set, and the remaining 9,000 pictures are used as the second target terrorist knife and gun picture set.

[0039] Step S102, divide the labeled first target terrorist knife and gun picture set into a training set, a test set and a validation set;

[0040] In an alternative embodiment, the labeled first set of target terrorist knife and gun pictures is divided into a training set, a test set, and a validation set according to a quantity ratio (8:1:1); that is, among the 1000 pictures mentioned above, 800 are the training set, 100 are the test set, and 100 are the validation set.

[0041] In an alternative embodiment, after the step S102, it includes:

[0042] Performing data augmentation on the divided training set;

[0043] The data augmentation includes: augmentation of the same class and augmentation of mixed classes; the augmentation of mixed classes includes: mixup data augmentation and mosaic data augmentation.

[0044] In an alternative embodiment, to obtain a well-performing terrorist knife and gun detection model, a large amount of data is often required for support. The better the model trained, the stronger the generalization ability of the model. However, in reality, the number of samples of terrorist knives and guns is insufficient or the sample quality is not good enough. Moreover, the work of obtaining new terrorist knife and gun data often requires a large amount of time and labor costs. This requires performing data augmentation on the samples to improve the sample quality. By using data augmentation techniques, the computer can be fully utilized to generate data and increase the amount of terrorist knife and gun data. For example, methods such as scaling, translation, rotation, and color transformation are used to augment the data. The advantage of data augmentation is that it can increase the number of training samples and add appropriate noise data at the same time. In this way, training with the pictures after data augmentation can improve the generalization ability of the model, and the robustness of the model is also improved by increasing the noise data.

[0045] The data augmentation adopted by the present invention includes: augmentation of the same class and augmentation of mixed classes;

[0046] The augmentation of the same class is to horizontally flip a picture (a picture in the training set) with a certain random probability, so that the picture retains the original information.

[0047] The augmentation of mixed classes includes: mixup data augmentation and mosaic data augmentation.

[0048] The mixup data augmentation is specifically: generating new sample-label data by adding two sample-label data in proportion, that is, fusing two terrorist knife and gun pictures together according to a certain transparency. The data distribution after fusion can be expressed by the following formula:

[0049]

[0050]

[0051] Among them, (x i , y i ), (x j , y j ) respectively represent the labels (knife or gun) of two pictures of terrorist knives and guns. The value range of λ is [0, 1], is the label of the fused picture.

[0052] Furthermore, judge whether the sizes of the two pictures are equal. When they are equal, place the two pictures overlappingly, and judge whether the labels in the two pictures overlap. If they overlap, there is no need to perform a fusion operation on these two pictures, and only one of the two pictures needs to be selected; when the sizes of the two pictures are not equal, perform a scaling operation on the two pictures to make their sizes equal.

[0053] The mosaic data augmentation is specifically as follows:

[0054] Use four pictures of terrorist knives and guns, splice the four pictures. Each picture has its corresponding frame. After splicing the four pictures, a new picture is obtained, and at the same time, the frame corresponding to this picture is also obtained. Then, such a new picture is input into the neural network for learning, which is equivalent to inputting four pictures for learning at once.

[0055] The specific process is as follows: Randomly select four pictures of terrorist knives and guns, create a new mosaic drawing board (the drawing board is square, and the side length of the drawing board is s_mosaic), and randomly generate a point (x_c, y_c) on the mosaic drawing board, and place 4 pictures of terrorist knives and guns around the random point (x_c, y_c).

[0056] Place the picture in the upper left corner of the drawing board. Assume the width and height of the picture are h and w. The position of the picture after placement is (x1a, y1a, x2a, y2a). When the width and height of the picture exceed the drawing board, the exceeded part is discarded. Therefore, the coordinate of the picture placement position is (0, 0, x_c, y_c); when the width and height of the picture do not exceed the drawing board, the picture placement position is (x_c - w, y_c - h, x_c, y_c). Then the coordinates after placement are as follows:

[0057] x1a, y1a, x2a, y2a = max(x_c - w, 0), max(y_c - h, 0), x_c, y_c

[0058] Place an image in the upper right corner of the drawing board. Assume the width and height of the image are h and w. The position of the image after placement is (x1a, y1a, x2a, y2a). When the width and height of the image exceed the drawing board, the exceeded part is discarded, and the coordinate of the image placement position is (x_c, 0, s_mosaic, y_c); when the width and height of the image do not exceed the drawing board, the image placement position is (x_c, y_c - h, x_c + w, y_c). Then the coordinates after placement are as follows:

[0059] x1a, y1a, x2a, y2a = x_c, max(y_c - h, 0), min(x_c + w, s_mosaic), y_c

[0060] Similarly, the coordinates for placing the image in the lower left corner of the drawing board and the coordinates for placing the image in the lower right corner of the drawing board can be obtained.

[0061] The above mosaic data augmentation method can enrich the number of samples.

[0062] Furthermore, in the present invention, a part of the training set is randomly augmented with the same class, a part of the training set is randomly augmented with mixup data, and a part of the training set is randomly augmented with mosaic data, rather than performing only one type of data augmentation on all the training sets.

[0063] Step S103, add a feature map pyramid network to the hierarchical vision (swin-transformer) feature extraction network to obtain an updated swin-transformer feature extraction network; input the training set into the updated swin-transformer feature extraction network for model training, and verify through the validation set to obtain a primary violent terrorist knife and gun detection model; test the test set according to the primary violent terrorist knife and gun detection model to obtain the accuracy and recall rate of the test set detection;

[0064] In an optional implementation manner, the adding a feature map pyramid network to the hierarchical vision (swin-transformer) feature extraction network to obtain an updated swin-transformer feature extraction network includes:

[0065] S1031. Extract features from the sample image through the backbone network of the swin-transformer feature extraction network to obtain four feature layers;

[0066] Specifically, they are four feature layers P2, P3, P4, and P5.

[0067] S1032. Combine and connect (concat) the four feature layers to obtain three feature maps;

[0068] Its calculation formula is as follows:

[0069] F1 = C(P2, P3, P4, P5) = P2 || Upx2(P3) || Upx4(P4) || Upx8(P5);

[0070] F2 = C(P3, P4, P5) = Upx2(P3) || Upx4(P4) || Upx8(P5);

[0071] F3 = C(P4, P5) = Upx4(P4) || Upx8(P5);

[0072] Among them, P2 is the first feature layer; P3 is the second feature layer; P4 is the third feature layer; P5 is the fourth feature layer; F1 is the first feature map; F2 is the second feature map; F3 is the third feature map; Upx2 is upsampling by 2 times; Upx4 is upsampling by 4 times; Upx8 is upsampling by 8 times.

[0073] S1033. Send the three feature maps into a convolutional layer respectively for convolution to obtain the updated swin-transformer feature extraction network.

[0074] Specifically, send the three obtained feature maps into the Conv(3,3)-BN-ReLU layer respectively for convolution, and finally obtain the updated swin-transformer feature extraction network.

[0075] Input the training set (800 violent terrorist knife and gun pictures) into the updated swin-transformer feature extraction network for model training. When all the pictures in the training set are trained, it is the first round of model training, that is, the first-round violent terrorist knife and gun detection model is obtained. At this time, verify the first-round violent terrorist knife and gun detection model through the validation set (100 violent terrorist knife and gun pictures). If the accuracy rate and recall rate detected by the validation set do not reach their respective preset thresholds, conduct the second round of model training, that is, retrain all the pictures in the training set through the first-round violent terrorist knife and gun detection model to obtain the second-round violent terrorist knife and gun detection model. At this time, verify the second-round violent terrorist knife and gun detection model through the validation set (100 violent terrorist knife and gun pictures). If the accuracy rate and recall rate detected by the validation set do not reach their respective preset thresholds, conduct the third round of model training. Repeat the above operations until after n rounds, obtain the primary violent terrorist knife and gun detection model; test the test set (100 violent terrorist knife and gun pictures) according to the primary violent terrorist knife and gun detection model to obtain the accuracy rate and recall rate detected by the test set;

[0076] Step S104, when it is determined that both the accuracy rate and the recall rate are greater than their respective preset thresholds, obtain the final violent terrorist knife and gun detection model.

[0077] In the present invention, when it is determined that the accuracy rate is greater than 90% and the recall rate is greater than 80%, a final violent terrorist knife and gun detection model is obtained.

[0078] In an alternative embodiment, when it is determined that the accuracy rate or the recall rate does not reach their respective preset thresholds, step S105 is entered;

[0079] Step S105: Detect the second target violent terrorist knife and gun picture set according to the primary violent terrorist knife and gun detection model to obtain correct pictures, incorrect pictures, and the position information and category information of the correct pictures; input the correct pictures into the first target violent terrorist knife and gun picture set and divide it into an updated training set, an updated test set, and an updated verification set; perform model training on the updated training set according to the primary violent terrorist knife and gun detection model and verify it through the updated verification set to obtain an updated violent terrorist knife and gun detection model;

[0080] Specifically, detect the second target violent terrorist knife and gun picture set (9000 pictures) according to the primary violent terrorist knife and gun detection model to obtain correct pictures (2000 pictures), incorrect pictures (7000 pictures), and according to this primary violent terrorist knife and gun detection model, the position information and category information of the labels (knives and / or guns) in the violent terrorist knife and gun pictures can be directly detected.

[0081] Input the correct pictures (2000 pictures) into the first target violent terrorist knife and gun picture set (1000 pictures) and divide it into an updated training set, an updated test set, and an updated verification set;

[0082] Divide them according to the quantity ratio of 8:1:1 to obtain an updated training set (2400 pictures), an updated test set (300 pictures), and an updated verification set (300 pictures).

[0083] Perform model training on the updated training set (2400 pictures) according to the primary violent terrorist knife and gun detection model and verify it through the updated verification set (300 pictures) to obtain an updated violent terrorist knife and gun detection model.

[0084] Step S106: Use the updated violent terrorist knife and gun detection model as the primary violent terrorist detection model, and use the incorrect pictures as the second target violent terrorist knife and gun picture set, and repeat step 105 until the accuracy rate and recall rate detected by the test set are both greater than their respective preset thresholds, then a final violent terrorist knife and gun detection model is obtained.

[0085] At this time, the second target violent terrorist knife and gun picture set is (7,000 pictures), and the first target violent terrorist knife and gun picture set is 3,000 pictures. The second target violent terrorist knife and gun picture set (7,000 pictures) is detected according to the updated violent terrorist knife and gun detection model, and correct pictures (3,000 pictures) and incorrect pictures (4,000 pictures) are obtained. And according to this primary violent terrorist knife and gun detection model, the position information and category information of the labels (knives and / or guns) in the violent terrorist knife and gun pictures can be directly detected.

[0086] Input the correct pictures (3,000 pictures) into the first target violent terrorist knife and gun picture set (3,000 pictures) and divide them into an updated training set, an updated test set, and an updated validation set; divide them according to the quantity ratio of 8:1:1 to obtain an updated training set (4,800 pictures), an updated test set (600 pictures), and an updated validation set (600 pictures). Repeat the above steps until the accuracy rate and recall rate detected by the test set are both greater than their respective preset thresholds, and the final violent terrorist knife and gun detection model is obtained.

[0087] Figure 2 It is a schematic structural diagram of a violent terrorist knife and gun detection model training system provided by an embodiment of the present invention, as Figure 2 shown. This system includes:

[0088] An acquisition unit 201, configured to acquire a target violent terrorist knife and gun picture set, randomly extract a preset number of pictures from the target violent terrorist knife and gun picture set as the first target violent terrorist knife and gun picture set; use the remaining pictures as the second target violent terrorist knife and gun picture set; perform position annotation and category annotation on the first target violent terrorist knife and gun picture set;

[0089] In an optional implementation manner, the acquisition of the target violent terrorist knife and gun picture set includes:

[0090] An acquisition subunit 2011, configured to acquire an original violent terrorist knife and gun picture set;

[0091] Specifically, because the number of violent terrorist picture sets is small, it is necessary to crawl the original violent terrorist knife and gun picture set from multiple channels (such as Google, Baidu) on the Internet.

[0092] A screening subunit 2012, configured to screen the original violent terrorist knife and gun picture set;

[0093] Specifically, the types of the crawled picture sets are relatively many. At this time, pictures in gif format and single-channel (8-bit) pictures need to be excluded, and at the same time, pictures with too small resolution crawled are excluded.

[0094] A selection subunit 2013, configured to manually review and select the screened violent terrorist knife and gun picture set to obtain the target violent terrorist knife and gun picture set.

[0095] The image data after screening is in a relatively clean format. Then, through manual review, a target set of terrorist knife and gun images is selected (that is, the original set of terrorist knife and gun images crawled from the network may contain images that do not belong to terrorist knives and guns. Therefore, further selection through manual review is required to ensure that each image is a terrorist knife and gun image. All the terrorist knife and gun images are combined to obtain the target set of terrorist knife and gun images).

[0096] After obtaining the target set of terrorist knife and gun images, a preset number of images are randomly selected from the target set of terrorist knife and gun images as the first target set of terrorist knife and gun images; the remaining images are used as the second target set of terrorist knife and gun images; the first target set of terrorist knife and gun images is subjected to position annotation and category annotation;

[0097] Specifically, in the present invention, after obtaining the target set of terrorist knife and gun images, a preset number (a small batch) of images need to be extracted from the target set of terrorist knife and gun images as the first target set of terrorist knife and gun images. Then, the first target set of terrorist knife and gun images is manually annotated using the lableimg annotation software to obtain the position information and category information of each image in the first target set of terrorist knife and gun images. Since manual annotation requires a large amount of time, manpower, and energy, the number of the first target set of terrorist knife and gun images extracted is very small. For example: if the target set of terrorist knife and gun images is 10,000, 1000 images are extracted as the first target set of terrorist knife and gun images, and the remaining 9000 images are used as the second target set of terrorist knife and gun images.

[0098] The partitioning unit 202 is used to partition the annotated first target set of terrorist knife and gun images into a training set, a test set, and a validation set;

[0099] In an alternative embodiment, the annotated first target set of terrorist knife and gun images is partitioned into a training set, a test set, and a validation set according to a quantity ratio (8:1:1); that is, among the above 1000 images, 800 are for the training set, 100 are for the test set, and 100 are for the validation set.

[0100] In an alternative embodiment, the terrorist knife and gun detection model training system further includes:

[0101] The data augmentation unit is used to perform data augmentation on the partitioned training set;

[0102] The data augmentation includes: same-class augmentation and mixed-class augmentation; the mixed-class augmentation includes: mixup data augmentation and mosaic data augmentation.

[0103] In an alternative embodiment, to obtain a well-performing terrorist knife and gun detection model, a large amount of data is often required as support. The better the trained model, the stronger its generalization ability. However, in practice, the number of terrorist knife and gun samples is insufficient or the sample quality is not good enough. Moreover, the work of obtaining new terrorist knife and gun data often requires a large amount of time and labor costs. Therefore, data augmentation for terrorist knife and gun samples is necessary to improve the sample quality. By using data augmentation techniques, the computer can be fully utilized to generate data and increase the amount of terrorist knife and gun data. For example, methods such as scaling, translation, rotation, and color transformation can be used to augment the data. The advantage of data augmentation is that it can increase the number of training samples and add appropriate noise data. Thus, training with the augmented images can improve the generalization ability of the model, and adding noise data also enhances the robustness of the model.

[0104] The data augmentation adopted by the present invention includes: homogeneous augmentation and heterogeneous augmentation;

[0105] The homogeneous augmentation is to horizontally flip a picture (a picture in the training set) with a certain random probability, so that the picture retains the original information.

[0106] The heterogeneous augmentation includes: mixup data augmentation and mosaic data augmentation.

[0107] The mixup data augmentation is specifically as follows: new sample-label data is generated by adding two sample-label data in proportion, that is, two terrorist knife and gun pictures are fused together according to a certain transparency. The data distribution after fusion can be expressed by the following formula:

[0108]

[0109]

[0110] where (x i , y i ), (x j , y j ) respectively represent the labels (knife or gun) of two terrorist knife and gun pictures, and the value range of λ is [0, 1], is the label of the fused picture.

[0111] Furthermore, it is judged whether the sizes of the two pictures are equal. When they are equal, the two pictures are placed overlapping, and it is judged whether the labels in the two pictures overlap. If they overlap, there is no need to fuse these two pictures, and only one of the two pictures is selected; when the sizes of the two pictures are not equal, the two pictures are scaled to make their sizes equal.

[0112] The mosaic data augmentation is specifically as follows:

[0113] Using four pictures of terrorist knives and guns, splice the four pictures. Each picture has its corresponding frame. After splicing the four pictures, a new picture is obtained, and at the same time, the frame corresponding to this picture is also obtained. Then, such a new picture is input into the neural network for learning, which is equivalent to inputting four pictures at once for learning.

[0114] The specific process is as follows: Randomly select four pictures of terrorist knives and guns, create a mosaic drawing board (the drawing board is square, and the side length of the drawing board is s_mosaic), and randomly generate a point (x_c, y_c) on the mosaic drawing board. Place 4 pictures of terrorist knives and guns around the random point (x_c, y_c).

[0115] Place the picture in the upper left corner of the drawing board. Assume the width and height of the picture are h and w. The position of the picture after placement is (x1a, y1a, x2a, y2a). When the width and height of the picture exceed the drawing board, the excess part is discarded. Therefore, the coordinate of the picture placement position is (0, 0, x_c, y_c); when the width and height of the picture do not exceed the drawing board, the picture placement position is (x_c - w, y_c - h, x_c, y_c). Then the coordinates after placement are as follows:

[0116] x1a, y1a, x2a, y2a = max(x_c - w, 0), max(y_c - h, 0), x_c, y_c

[0117] Place the picture in the upper right corner of the drawing board. Assume the width and height of the picture are h and w. The position of the picture after placement is (x1a, y1a, x2a, y2a). When the width and height of the picture exceed the drawing board, the excess part is discarded, and the picture placement position coordinate is (x_c, 0, s_mosaic, y_c); when the width and height of the picture do not exceed the drawing board, the picture placement position is (x_c, y_c - h, x_c + w, y_c). Then the coordinates after placement are as follows:

[0118] x1a, y1a, x2a, y2a = x_c, max(y_c - h, 0), min(x_c + w, s_mosaic), y_c

[0119] Similarly, the coordinates for placing the picture in the lower left corner of the drawing board and the coordinates for placing the picture in the lower right corner of the drawing board can be obtained.

[0120] Through the above mosaic data augmentation method, the number of samples can be enriched.

[0121] Further, in the present invention, a part of the training set is randomly enhanced with the same class, a part of the training set is randomly enhanced with mixup data, and a part of the training set is randomly enhanced with mosaic data, rather than only performing the same type of data enhancement on all the training sets.

[0122] The model training unit 203 is configured to add a feature pyramid network to the hierarchical vision (swin-transformer) feature extraction network to obtain an updated swin-transformer feature extraction network; input the training set into the updated swin-transformer feature extraction network for model training, and verify it through the validation set to obtain a primary terrorist knife and gun detection model; test the test set according to the primary terrorist knife and gun detection model to obtain the accuracy and recall rate of the test set detection.

[0123] In an optional embodiment, adding a feature pyramid network to the hierarchical vision (swin-transformer) feature extraction network to obtain an updated swin-transformer feature extraction network includes:

[0124] The extraction subunit 2031 is configured to extract features from the sample image through the backbone network of the swin-transformer feature extraction network to obtain four feature layers.

[0125] Specifically, they are four feature layers P2, P3, P4, and P5.

[0126] The fusion unit 2032 is configured to concatenate the four feature layers to obtain three feature maps.

[0127] The calculation formula is as follows:

[0128] F1 = C(P2, P3, P4, P5) = P2 || Upx2(P3) || Upx4(P4) || Upx8(P5);

[0129] F2 = C(P3, P4, P5) = Upx2(P3) || Upx4(P4) || Upx8(P5);

[0130] F3 = C(P4, P5) = Upx4(P4) || Upx8(P5);

[0131] Wherein, P2 is the first feature layer; P3 is the second feature layer; P4 is the third feature layer; P5 is the fourth feature layer; F1 is the first feature map; F2 is the second feature map; F3 is the third feature map; Upx2 is upsampling by 2 times; Upx4 is upsampling by 4 times; Upx8 is upsampling by 8 times.

[0132] The convolutional unit 2033 is configured to separately input the three feature maps into a convolutional layer for convolution to obtain the updated swin-transformer feature extraction network.

[0133] Specifically, the three obtained feature maps are separately input into a Conv(3,3)-BN-ReLU layer for convolution, and finally the updated swin-transformer feature extraction network is obtained.

[0134] Input the training set (800 terrorist knife and gun pictures) into the updated swin-transformer feature extraction network for model training. After all the pictures in the training set are trained, it is the first round of model training, that is, the first-round terrorist knife and gun detection model is obtained. At this time, use the validation set (100 terrorist knife and gun pictures) to verify the first-round terrorist knife and gun detection model. If the accuracy rate and recall rate detected by the validation set do not reach their respective preset thresholds, then perform the second round of model training, that is, retrain all the pictures in the training set through the first-round terrorist knife and gun detection model to obtain the second-round terrorist knife and gun detection model. At this time, use the validation set (100 terrorist knife and gun pictures) to verify the second-round terrorist knife and gun detection model. If the accuracy rate and recall rate detected by the validation set do not reach their respective preset thresholds, then perform the third round of model training. Repeat the above operations until after n rounds, the primary terrorist knife and gun detection model is obtained; test the test set (100 terrorist knife and gun pictures) according to the primary terrorist knife and gun detection model to obtain the accuracy rate and recall rate of the test set detection;

[0135] The judgment unit 204 is configured to obtain the final terrorist knife and gun detection model when it is determined that both the accuracy rate and the recall rate are greater than their respective preset thresholds.

[0136] In the present invention, when it is determined that the accuracy rate is greater than 90% and the recall rate is greater than 80%, the final terrorist knife and gun detection model is obtained.

[0137] When it is determined that the accuracy rate or the recall rate does not reach their respective preset thresholds, enter the model update unit 205;

[0138] The model update unit 205 is configured to detect the second target terrorist knife and gun picture set according to the primary terrorist knife and gun detection model to obtain correct pictures, wrong pictures, and the position information and category information of the correct pictures; input the correct pictures into the first target terrorist knife and gun picture set and divide them into an updated training set, an updated test set, and an updated validation set; perform model training on the updated training set according to the primary terrorist knife and gun detection model and verify it through the updated validation set to obtain an updated terrorist knife and gun detection model;

[0139] Specifically, the second target set of terrorist knife and gun pictures (9,000 pictures) is detected by the primary terrorist knife and gun detection model, obtaining correct pictures (2,000 pictures) and incorrect pictures (7,000 pictures), and the position information and category information of the labels (knives and / or guns) in the terrorist knife and gun pictures can be directly detected according to the primary terrorist knife and gun detection model.

[0140] The correct pictures (2,000 pictures) are input into the first target set of terrorist knife and gun pictures (1,000 pictures) and divided into an updated training set, an updated test set, and an updated validation set.

[0141] They are divided according to the quantity ratio of 8:1:1 to obtain an updated training set (2,400 pictures), an updated test set (300 pictures), and an updated validation set (300 pictures).

[0142] The updated training set (2,400 pictures) is used for model training according to the primary terrorist knife and gun detection model and verified through the updated validation set (300 pictures) to obtain an updated terrorist knife and gun detection model.

[0143] The training unit 206 is repeated, which is used to use the updated terrorist knife and gun detection model as the primary terrorist detection model and the incorrect pictures as the second target set of terrorist knife and gun pictures, and repeat the model update unit until the accuracy rate and recall rate of the test set detection are both greater than their respective preset thresholds, and then the final terrorist knife and gun detection model is obtained.

[0144] At this time, the second target set of terrorist knife and gun pictures is (7,000 pictures), and the first target set of terrorist knife and gun pictures is 3,000 pictures. The second target set of terrorist knife and gun pictures (7,000 pictures) is detected according to the updated terrorist knife and gun detection model, obtaining correct pictures (3,000 pictures) and incorrect pictures (4,000 pictures), and the position information and category information of the labels (knives and / or guns) in the terrorist knife and gun pictures can be directly detected according to the primary terrorist knife and gun detection model.

[0145] The correct pictures (3,000 pictures) are input into the first target set of terrorist knife and gun pictures (3,000 pictures) and divided into an updated training set, an updated test set, and an updated validation set; they are divided according to the quantity ratio of 8:1:1 to obtain an updated training set (4,800 pictures), an updated test set (600 pictures), and an updated validation set (600 pictures). Repeat the above steps until the accuracy rate and recall rate of the test set detection are both greater than their respective preset thresholds, and then the final terrorist knife and gun detection model is obtained.

[0146] Through the method of the present invention, the accuracy of detecting violent and terrorist knives and guns in pictures on the online social platform can reach 90%, ensuring the accuracy of detection.

[0147] Advantages of the present invention:

[0148] The present invention provides a method, a system and a computer storage medium for training a violent and terrorist knife and gun detection model. Among them, the method provides an improved version of swin-transformer for detecting violent and terrorist knife and gun pictures. Specifically, a feature pyramid network is added to the hierarchical vision (swin-transformer) feature extraction network to obtain an updated swin-transformer feature extraction network; through this method, pictures on the online social platform can be accurately detected for violent and terrorist knives and guns, with a detection accuracy of 90%, and the specific position coordinates of the violent and terrorist knife and gun pictures are given. It is judged whether the picture is green and healthy content, effectively solving the problems of insufficient efficiency in manual picture review and the consumption of human and financial resources brought by the increasing amount of pictures.

[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for training a violent terrorist knife and gun detection model, characterized in that, Including: Step S101: Obtain a target set of terrorist knife and gun pictures, and randomly select a preset number of pictures from the target set of terrorist knife and gun pictures as the first target set of terrorist knife and gun pictures; Use the remaining pictures as the second target set of terrorist knife and gun pictures; perform position annotation and category annotation on the first target set of terrorist knife and gun pictures; Step S102: Divide the annotated first target set of terrorist knife and gun pictures into a training set, a test set, and a validation set; Step S103: Add a feature pyramid network to the hierarchical visual swin-transformer feature extraction network to obtain an updated swin-transformer feature extraction network; input the training set into the updated swin-transformer feature extraction network for model training, and verify through the validation set to obtain a primary terrorist knife and gun detection model; test the test set according to the primary terrorist knife and gun detection model to obtain the accuracy rate and recall rate of the test set detection; Step S104: When it is determined that both the accuracy rate and the recall rate are greater than their respective preset thresholds, obtain a final terrorist knife and gun detection model; When it is determined that the accuracy rate or the recall rate does not reach their respective preset thresholds, go to step S105; Step S105: Detect the second target set of terrorist knife and gun pictures according to the primary terrorist knife and gun detection model to obtain correct pictures, incorrect pictures, and the position information and category information of the correct pictures; input the correct pictures into the first target set of terrorist knife and gun pictures and divide them into an updated training set, an updated test set, and an updated validation set; perform model training on the updated training set according to the primary terrorist knife and gun detection model and verify through the updated validation set to obtain an updated terrorist knife and gun detection model; Step S106: Use the updated terrorist knife and gun detection model as the primary terrorist detection model, and use the incorrect pictures as the second target set of terrorist knife and gun pictures, repeat step S105 until both the accuracy rate and the recall rate of the test set detection are greater than their respective preset thresholds, and obtain a final terrorist knife and gun detection model; The adding a feature pyramid network to the hierarchical visual swin-transformer feature extraction network to obtain an updated swin-transformer feature extraction network includes: Extract features from the sample pictures through the backbone network of the swin-transformer feature extraction network to obtain four feature layers; Merge and connect the four feature layers to obtain three feature maps; Send the three feature maps into convolutional layers for convolution respectively to obtain the updated swin-transformer feature extraction network; The calculation formula for merging and connecting the four feature layers to obtain three feature maps is: F1 = C(P2, P3, P4, P5) = P2 || Upx2(P3) || Upx4(P4) || Upx8(P5); F2 = C(P3, P4, P5) = Upx2(P3) || Upx4(P4) || Upx8(P5); F3 = C(P4, P5) = Upx4(P4) || Upx8(P5); Among them, P2 is the first feature layer; P3 is the second feature layer; P4 is the third feature layer; P5 is the fourth feature layer; F1 is the first feature map; F2 is the second feature map; F3 is the third feature map; Upx2 is upsampling by a factor of 2; Upx4 is upsampling by a factor of 4; Upx8 is upsampling by a factor of 8.

2. The method according to claim 1, characterized in that, The obtaining of the target violent terrorist knife and gun picture set includes: Obtaining the original violent terrorist knife and gun picture set; Screening the original violent terrorist knife and gun picture set; Manually reviewing and selecting the screened violent terrorist knife and gun picture set to obtain the target violent terrorist knife and gun picture set.

3. The method according to claim 1, characterized in that, After the step S102, it includes: Performing data augmentation on the divided training set; The data augmentation includes: augmentation of the same class and augmentation of mixed classes; the augmentation of mixed classes includes: mixup data augmentation and mosaic data augmentation.

4. A violent terrorist knife and gun detection model training system, characterized in that, It includes: An obtaining unit, configured to obtain a target violent terrorist knife and gun picture set, and randomly extract a preset number of pictures from the target violent terrorist knife and gun picture set as the first target violent terrorist knife and gun picture set; Taking the remaining pictures as the second target violent terrorist knife and gun picture set; Performing position annotation and class annotation on the first target violent terrorist knife and gun picture set; A dividing unit, configured to divide the labeled first target violent terrorist knife and gun picture set into a training set, a test set, and a validation set; A model training unit, configured to add a feature map pyramid network to the hierarchical visual swin-transformer feature extraction network to obtain an updated swin-transformer feature extraction network; input the training set into the updated swin-transformer feature extraction network for model training, and verify through the validation set to obtain a primary violent terrorist knife and gun detection model; test the test set according to the primary violent terrorist knife and gun detection model to obtain the accuracy rate and recall rate of the test set detection; A judging unit, configured to obtain a final violent terrorist knife and gun detection model when it is determined that both the accuracy rate and the recall rate are greater than their respective preset thresholds; When it is determined that the accuracy rate or the recall rate does not reach their respective preset thresholds, enter the model updating unit; A model updating unit, configured to detect the second target violent terrorist knife and gun picture set according to the primary violent terrorist knife and gun detection model to obtain correct pictures, wrong pictures, and the position information and class information of the correct pictures; input the correct pictures into the first target violent terrorist knife and gun picture set and divide it into an updated training set, an updated test set, and an updated validation set; perform model training on the updated training set according to the primary violent terrorist knife and gun detection model and verify through the updated validation set to obtain an updated violent terrorist knife and gun detection model; A repeated training unit, configured to use the updated terrorist knife and gun detection model as the primary terrorist detection model, use the error pictures as the second target terrorist knife and gun picture set, and repeat the model update unit until the accuracy rate and recall rate of the test set detection are both greater than their respective preset thresholds, so as to obtain a final terrorist knife and gun detection model; The step of adding a feature pyramid network to the hierarchical vision swin-transformer feature extraction network to obtain an updated swin-transformer feature extraction network includes: Performing feature extraction on the sample pictures through the backbone network of the swin-transformer feature extraction network to obtain four feature layers; Combining and connecting the four feature layers to obtain three feature maps; Feeding the three feature maps into convolutional layers respectively for convolution to obtain the updated swin-transformer feature extraction network; The calculation formula for combining and connecting the four feature layers to obtain three feature maps is: F1 = C(P2, P3, P4, P5) = P2 || Upx2(P3) || Upx4(P4) || Upx8(P5); F2 = C(P3, P4, P5) = Upx2(P3) || Upx4(P4) || Upx8(P5); F3 = C(P4, P5) = Upx4(P4) || Upx8(P5); Wherein, P2 is the first feature layer; P3 is the second feature layer; P4 is the third feature layer; P5 is the fourth feature layer; F1 is the first feature map; F2 is the second feature map; F3 is the third feature map; Upx2 is upsampling by 2 times; Upx4 is upsampling by 4 times; Upx8 is upsampling by 8 times.

5. The system according to claim 4, wherein, It further includes: A data augmentation unit, configured to perform data augmentation on the divided training set; The data augmentation includes: homogeneous augmentation and heterogeneous augmentation; the heterogeneous augmentation includes: mixed mixup data augmentation and mosaic data augmentation.

6. A computer storage medium having a computer program stored thereon, wherein, When the program is executed by a processor, it implements the terrorist knife and gun detection model training method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Violence and terrorism picture safety detection system based on deep learning

    CN112906588A

  • Target detection method and picture detection model training method

    CN114529792A