Bullet hole detection method based on deep neural network RT-detr
Through the bullet hole detection method based on the deep neural network RT-detr, the existing automatic target reporting system is solved inefficient in target surface recognition and performance statistics, and high-precision and robust bullet hole detection is achieved, which is suitable for outdoor areas and frequent site changes.
Patent Information
- Application Number
- CN202510099770.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-13
AI Technical Summary
The existing automatic target reporting system is inefficient in target surface recognition and performance statistics, and is prone to misreporting and misreporting, and is not suitable for use in the field or frequently changing venues.
The bullet hole detection method based on the deep neural network RT-detr is adopted to collect target surface images through the front-end information acquisition device, preprocess and label, and an improved RT-detr network structure is constructed, including backbone, encoder, decoder, prediction heads, and auxiliary reversible branches of the programmable gradient information module are added to the backbone part. The Hilo attention mechanism is used in the encoder part to achieve efficient bullet hole detection.
It improves the accuracy and robustness of bullet hole detection, reduces information transmission loss, enhances adaptability to field and frequent site replacement, and provides higher detection accuracy and mobility.
Smart Images

Figure CN119992058A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection by computer vision, and in particular relates to a bullet hole detection method based on deep neural network RT-detr. Background Art
[0002] The traditional method of reporting targets is generally manual. The traditional manual reporting method is simple, but it also has many problems. For example, manual reporting takes a long time. From putting on the target, reading the target, changing the target and recording the score, all need to be done manually. Especially for the target surface with multiple bullet holes, manual reporting is inefficient in identifying bullet holes and counting scores, and it is easy to miss or report incorrectly. This error will directly affect the shooting training results. Moreover, manual reporting poses a certain threat to the life safety of the reporter.
[0003] The automatic target reporting system is a new type of target reporting method that has emerged in recent years. It has developed very rapidly and there are already many systems that can be put into practical use. From the perspective of technical implementation, whether at home or abroad, the mainstream automatic target reporting systems for live-fire shooting in recent years are mainly divided into electronic target surface type, photoelectric target surface type, acoustic and electric type, optical fiber coding type and image processing type. However, there are still many defects in the existing target reporting system: the photoelectric target reporting system has many components, serious power consumption and high heat generation; in general, cable transmission is used, resulting in poor overall mobility; the characteristics of sensitivity to light changes, complex circuit structure and precise characteristics of sensors determine that this type of target reporting system is not suitable for use in the field or in places where training grounds are frequently changed. Acoustic and electric positioning automatic target reporting system: Since sound waves are continuously propagated in the air, the sensor is easily interfered by the sound waves of the target near the target position. When continuous shooting is carried out, the sound waves of two consecutive bullets hitting the target will also interfere with each other. In addition, the noise in the environment around the shooting range will also cause great interference to the accuracy of the target reporting.
[0004] Therefore, researching an automatic target reporting system with high target reporting accuracy, low cost, strong stability and timeliness is a key factor in promoting and popularizing the automatic target reporting system. Summary of the invention
[0005] In view of the above problems existing in the prior art, the present invention proposes a bullet hole detection method based on deep neural network RT-detr, which has a reasonable design, solves the shortcomings of the prior art and has good effects.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] A bullet hole detection method based on deep neural network RT-detr, comprising the following steps:
[0008] Step 1: Collect the target surface image through the front-end information acquisition device;
[0009] Step 2: Preprocess and annotate the bullet hole images to generate a bullet hole dataset for deep learning;
[0010] Step 3: Construct a bullet hole detection model based on the improved RT-detr network structure;
[0011] Step 4: Use the bullet hole data set to train and test the bullet hole detection model to obtain a trained bullet hole detection model;
[0012] Step 5: Input the bullet hole image collected in real time into the trained bullet hole detection model to obtain the bullet hole detection result.
[0013] Step 6: Pass the bullet hole detection results to the back-end device for image processing and score statistics.
[0014] Furthermore, in step 2, the preprocessing is to remove the background through Maskrcnn to obtain the target surface image.
[0015] Furthermore, in step 3, the improved RT-detr network structure includes four parts: backbone, encoder, decoder, and prediction heads; an auxiliary reversible branch in the programmable gradient information module is added to the backbone part to generate reliable gradients and update network parameters.
[0016] Furthermore, the feature map output by the backbone is input into the encoder part, and the encoder includes an AIFI module and a CCFM module. The AIFI module introduces the Hilo attention mechanism, which captures the local details of the object at high frequency and encodes the global structure at low frequency.
[0017] The Hilo attention mechanism adopts two effective attentions to decouple the high and low frequencies in the feature map. In one path, multiple heads are assigned to high-frequency attention to capture fine-grained high frequencies; in the other path, average pooling is first applied to each window to obtain low-frequency signals, and then the remaining heads are assigned to low-frequency attention to model the relationship between each query position in the input feature map and the average pooled low-frequency key of each window.
[0018] Furthermore, the refined high-frequency features and low-frequency features are connected and input into the decoder and prediction heads in sequence to obtain the bullet hole detection results.
[0019] Beneficial technical effects brought by the present invention:
[0020] 1. Add an auxiliary supervision framework called programmable gradient information to the backbone network to solve the problem of information transmission loss. The main branch deep features that lose important information due to information bottlenecks will be able to receive reliable gradient information from the auxiliary reversible branches. This gradient information will drive parameter learning to help extract correct and important information, and the above actions can enable the main branch to obtain more effective features for the target task.
[0021] 2. The Hilo attention mechanism is used in the encoder part. The high frequency captures the local details of the object (bullet hole shape), while the low frequency encodes the global structure (texture and color). Hi-Fi is designed to capture fine-grained high frequencies with local window self-attention (such as 2×2 windows), which can save a lot of computational complexity; low-frequency attention (Lo-Fi) first applies average pooling to each window to obtain the low-frequency signal. Then, the remaining Head is assigned to Lo-Fi to model the relationship between each query position in the input feature map and the average pooled low-frequency key of each window. Benefiting from the reduction in key and value lengths, the complexity of Lo-Fi is significantly reduced;
[0022] 3. Since the algorithm is based on a deep learning neural network, it provides higher robustness than traditional target reporting methods under various conditions such as changing weather, light, seasons, viewpoints, and occlusion caused by the presence of moving objects, and provides higher detection accuracy than the current mainstream target detection network. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is the RT-detr network structure diagram of the present invention;
[0024] Figure 2 This is a flow chart of the bullet hole detection method of the present invention;
[0025] Figure 3 It is a bullet hole data diagram under different situations collected in the present invention;
[0026] Figure 4 It is the programmable gradient information and related network structure diagram in the present invention;
[0027] Among them, (a) is the path aggregation network; (b) is the reversible column; (c) is the traditional deep supervision; (d) is the programmable gradient information;
[0028] Figure 5 This is the structural diagram of the Hilo attention mechanism in the present invention;
[0029] Figure 6 This is a diagram of the bullet hole detection result in the present invention;
[0030] Figure 7This is a comparison chart of bullet hole detection results on different networks;
[0031] Figure 8 This is a screenshot of the program interface displayed for the target reporting system; DETAILED DESCRIPTION
[0032] The specific implementation of the present invention is further described below in conjunction with specific embodiments:
[0033] like Figure 1 and Figure 2 As shown, a bullet hole detection method based on deep neural network RT-detr includes the following steps:
[0034] Step 1: Collect the target surface image through the front-end information acquisition device;
[0035] Step 2: Preprocess and annotate the bullet hole images to generate a bullet hole dataset for deep learning;
[0036] The preprocessing is to remove the background through Maskrcnn and obtain the target surface image.
[0037] Step 3: Construct a bullet hole detection model based on the improved RT-detr network structure;
[0038] The improved RT-detr network structure consists of four parts: backbone, encoder, decoder, and prediction heads. An auxiliary reversible branch in the programmable gradient information module is added to the backbone part to generate reliable gradients and reduce the data loss problem caused by the propagation and transformation of bullet hole data.
[0039] To solve the problem of information loss, the present invention adopts a new auxiliary supervision framework called Programmable Gradient Information (PGI), such as Figure 4 As shown in (d), it is used to solve the problems of information loss and gradient decay caused by the increase in the depth of the neural network. PGI mainly consists of three components, namely the main branch, the auxiliary reversible branch and the multi-level auxiliary information. Among them, the auxiliary reversible branch is designed to deal with the problems caused by the deepening of the neural network. The deepening of the network will lead to information bottlenecks, making it impossible for the loss function to generate reliable gradients. As for the multi-level auxiliary information, it is designed to deal with the error accumulation problem caused by deep supervision, especially for the architecture of multiple prediction branches and lightweight models. Next, these two components will be introduced step by step.
[0040] In the programmable gradient information module, an auxiliary reversible branch is proposed to generate reliable gradients and update network parameters. By providing information about the mapping from data to the target, the loss function can provide guidance and avoid the possibility of discovering spurious correlations from incomplete feedforward features that are not very relevant to the target. The complete information is maintained by introducing a reversible structure, but adding a main branch to the reversible structure will consume a lot of inference cost. By analyzing Figure 4 (b) We found that when adding additional connections from deep layers to shallow layers, the inference time increased by 20%. When repeatedly adding input data to the high-resolution computation layer of the network, the inference time even exceeded twice the original time.
[0041] Since the goal of the present invention is to use a reversible structure to obtain reliable gradients, “reversibility” is not the only necessary condition for the inference phase. In view of this, the reversible branch is regarded as an extension of the deep supervision branch, and then an auxiliary reversible branch is designed, such as Figure 4 As shown in (d). As for the deep features of the main branch that lose important information due to the information bottleneck, they will be able to receive reliable gradient information from the auxiliary reversible branch. These gradient information will drive parameter learning to help extract correct and important information, and the above actions can enable the main branch to obtain more effective features for the target task. In addition, the performance of reversible architectures on shallow networks is worse than on general networks because complex tasks require transformations in deeper networks. The method proposed in the present invention does not force the main branch to retain the complete original information, but generates useful gradients through an auxiliary supervision mechanism to update it. The advantage of this design is that the proposed method can also be applied to shallower networks.
[0042] It includes deep supervision of multiple prediction branches, and its architecture is as follows Figure 4 (c). For object detection, different feature pyramids can be used to perform different tasks, for example, they can be used together to detect objects of different sizes. Therefore, after connecting to the deep supervision branch, the shallow features will be guided to learn the features required for small object detection, and the system will treat the locations of objects of other sizes as background. However, the above behavior will cause the deep feature pyramid to lose a lot of information required to predict the target object. Each feature pyramid needs to receive information of all target objects so that the subsequent main branch can retain complete information to learn to predict various targets.
[0043] The concept of multi-level auxiliary information is to insert an integrated network between the feature pyramid hierarchy and the main branch for auxiliary supervision, and then use it to combine the returned gradients from different prediction heads, such as Figure 4As shown in (d). The multi-level auxiliary information aggregates the gradient information of all target objects, passes it to the main branch, and then updates the parameters. At this point, the features of the main branch feature pyramid hierarchy will no longer be dominated by certain specific object information. Therefore, the information corruption problem in deep supervision can be alleviated. In addition, any comprehensive network can be used for multi-level auxiliary information. In addition, the required semantic level can be planned to guide the learning of network architectures of different scales.
[0044] The feature map output by the backbone is input into the encoder part. The encoder includes the AIFI module and the CCFM module. The AIFI module introduces the Hilo attention mechanism, which captures the local details of the object at high frequencies, such as the traces and shapes of bullet holes, and encodes the global structure at low frequencies, such as the blank space of the entire target paper and the human-shaped target image.
[0045] The Hilo attention mechanism uses two effective attentions to decouple the high and low frequencies in the feature map, such as Figure 5 As shown in Figure 1, in one of the paths, a simple solution for head allocation is to allocate the same number of heads to Hi-Fi and Lo-Fi as in the standard MSA layer. However, this will incur more computational cost. To achieve better efficiency, HiLo divides the same number of heads in MSA into 2 groups, with the number of heads being N. h , the distribution ratio is α, where (1-α)N h Head for Hi-Fi, the other αN h Heads are used for Lo-Fi. In this way, a certain number of heads are allocated to high-frequency attention Hi-Fi to capture fine-grained high frequencies through LocalWindow Self-Attention (such as 2×2 windows); in another path, low-frequency attention Lo-Fi is implemented, and average pooling is first applied to each window to obtain low-frequency signals, and then the remaining heads are assigned to low-frequency attention to model the relationship between each query position in the input feature map and the average pooled low-frequency key of each window. Benefiting from the reduction in key and value lengths, the complexity of Lo-Fi is significantly reduced.
[0046] The refined high-frequency features and low-frequency features are connected and input into the decoder and prediction heads in turn to obtain the bullet hole detection results. For the image of bullet holes, the bullet perforation marks are regarded as the high-frequency part of the image, while the other portrait target rings and parts outside the area are regarded as the low-frequency part of the image. Since neither Hi-Fi nor Lo-Fi are equipped with time-consuming operations such as dilation windows and recursion, the overall framework of HiLo is fast on the GPU. Comprehensive benchmark tests show that HiLo outperforms existing attention mechanisms in terms of performance, FLOPs, throughput, and memory consumption.
[0047] Step 4: Use Figure 3 The bullet hole data set shown is annotated and the bullet hole detection model is trained and tested to obtain a trained bullet hole detection model;
[0048] After the data set is trained with the improved RT-detr network, the following is obtained: Figure 6 The bullet hole detection result shown in the figure shows that for target images with fewer bullet holes, the bullet hole detection rate is close to 100%, and for target images with dense bullet holes, the bullet hole detection results are also very good. At the same time, the experiment compares the current mainstream detection networks Yolov8 and Yolov9 and the original version of the RT-detr network. The comparison results are shown in the figure below. Figure 7 As shown in the figure, it can be found that the average accuracy of the improved RT-detr network reaches 79.3%, which is much higher than the 71.2% of the yolov9 network, and is 2% higher than the 77.1% of the original RT-detr network. The highest accuracy is 84.1%, indicating that the algorithm has good recognition ability and robustness.
[0049] Step 5: Input the bullet hole image collected in real time into the trained bullet hole detection model to obtain the bullet hole detection result.
[0050] Step 6: Pass the bullet hole detection results to the back-end device for image processing and score statistics.
[0051] The summary software receives the detection results of all terminals and displays and saves them. After entering the target reporting system, it needs to be initialized first. The mobile terminal will obtain the target surface image at the current moment and extract the effective target surface area as a template. When the shooting is completed, click the "Start Detection" button, the mobile terminal will obtain the latest target surface image, determine the target reporting result according to the established algorithm, and save the current user number, shooting time and screenshot of the current program interface. Figure 8 When you click "Change Target", you will enter the next set of target practice. You can view these data in the historical results function.
[0052] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by technicians in this technical field within the essential scope of the present invention should also fall within the protection scope of the present invention.
Claims
1. A bullet hole detection method based on deep neural network RT-detr, characterized in that: The following steps are included: Step 1: Collect the target surface image through the front-end information acquisition device; Step 2: Preprocess and annotate the bullet hole images to generate a bullet hole dataset for deep learning; Step 3: Construct a bullet hole detection model based on the improved RT-detr network structure; Step 4: Use the bullet hole data set to train and test the bullet hole detection model to obtain a trained bullet hole detection model; Step 5: Input the bullet hole image collected in real time into the trained bullet hole detection model to obtain the bullet hole detection result; Step 6: Pass the bullet hole detection results to the back-end device for image processing and score statistics.
2. The bullet hole detection method based on deep neural network RT-detr according to claim 1 is characterized in that: In step 2, the preprocessing is to remove the background through Maskrcnn to obtain the target surface image.
3. The bullet hole detection method based on deep neural network RT-detr according to claim 2 is characterized in that: In step 3, the improved RT-detr network structure includes four parts: backbone, encoder, decoder, and prediction heads; an auxiliary reversible branch in the programmable gradient information module is added to the backbone part to generate reliable gradients and update network parameters.
4. The bullet hole detection method based on deep neural network RT-detr according to claim 3 is characterized in that: The feature map output by the backbone is input into the encoder part. The encoder includes an AIFI module and a CCFM module. The AIFI module introduces the Hilo attention mechanism, which captures the local details of the object at high frequency and encodes the global structure at low frequency. The Hilo attention mechanism adopts two effective attentions to decouple the high and low frequencies in the feature map. In one path, multiple heads are assigned to high-frequency attention to capture fine-grained high frequencies; in the other path, average pooling is first applied to each window to obtain low-frequency signals, and then the remaining heads are assigned to low-frequency attention to model the relationship between each query position in the input feature map and the average pooled low-frequency key of each window.
5. The bullet hole detection method based on deep neural network RT-detr according to claim 4 is characterized in that: The refined high-frequency features and low-frequency features are connected and input into the decoder and prediction heads in turn to obtain the bullet hole detection results.