Underground small target adaptive identification method and model based on explicit feature reconstruction anchor-free frame
By using the method of explicit feature reconstruction of the anchor-free framework, combined with the small target framework network structure, SERT module, adaptive network and feature reconstruction module, the problem of small target feature information loss in the traditional network is solved, and efficient recognition of small targets in underground space is achieved.
Patent Information
- Application Number
- CN202511050496.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional networks are prone to losing or confusing feature information when identifying small targets in visible light images, resulting in recognition difficulties. Especially in underground spaces, small targets have low resolution and a small pixel ratio, making it difficult to achieve ideal recognition results.
A method based on explicit feature reconstruction of the anchor-free framework is adopted, including the small target framework network structure CRT-AWWR, the SERT module, the adaptive network Adaptive_gty and the small target resampling module SJHD. Through multi-level feature fusion and reconstruction, the small target feature information is retained and enhanced, and the loss function module is optimized to improve the recognition accuracy.
It effectively solves the problem of small target feature information loss, improves the accuracy and efficiency of small target recognition, simplifies the prediction process, and enhances detection performance, especially the ability to recognize small targets in underground spaces.
Smart Images

Figure CN120707841A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an adaptive recognition method and model for underground small targets based on explicit feature reconstruction and anchor-free framework, belonging to computer vision technology. Background Art
[0002] In the field of computer vision, small object recognition has always been a very challenging task. With the development of technology, researchers have proposed many methods to deal with this challenge. For example, Huang et al. in " Seg2Sonar: A full - class sample synthesis method applied to underwater sonar image target detection, recognition, and segmentation tasks ” IEEE Trans. Geosci.Remote Sens Proposed SEG2SONAR Network, through Skip - Layer Channelwise The excitation module improves the feature extraction capability, reduces the network's demand for sonar data and improves model accuracy; Wang exist" Maritime radar target detection in sea clutter based on CNN with dual - perspective attention ” IEEE Geosci. Remote Sens.Lett The convolutional neural network ( CNN ) and dual attention ( DPA ) radar target detection method, using dynamic CNN The structure realizes adaptive function coding to reduce the influence of chaotic interference; Wang Waiting in Two - stage collaborative sea clutter suppression based on subregion gazing and reconstruction ” IEEE Geosci.Remote Sens. Lett A two-stage collaborative chaos suppression algorithm is proposed to improve radar detection accuracy; Tian et al. in " Vision transformer with enhanced self - attention for few - shot ship target recognition in complex environments ” IEEE Trans. Instrum. Meas Through enhanced self-movement ( SAVT ) The model's visual transformer effectively extracts discriminative features and improves the recognition rate of small objects in complex environments; Xu et al. proposed a spatial scale adaptive real-time object detection neural network and introduced a new drone detection box filter ( UDBF ) module to eliminate false positives.
[0003] However, visible light image object recognition algorithms still face challenges. Small objects have low resolution and a small pixel count, which limits their feature information. Traditional networks are prone to losing functional information about small objects during feature extraction, or confusing it with background information, making it difficult to achieve ideal recognition results. Summary of the Invention
[0004] Purpose of the invention: The present invention provides an adaptive recognition method and model for underground small targets based on an anchor-free framework for explicit feature reconstruction, which solves the problem that small targets in visible light images are difficult to recognize and traditional networks easily lose or confuse small target feature information, thereby improving the accuracy and efficiency of small target recognition.
[0005] Technical solution: To achieve the above purpose, the technical solution adopted by the present invention is: An adaptive recognition method for underground small targets based on explicit feature reconstruction and anchor-free framework is applied to the target recognition model, which includes a small target framework network structure. CRT-AWWR 、 SERT Module, adaptive network Adaptive_gty and the small target resampling module, which includes an explicit feature reconstruction module SJHD The method for adaptively identifying small targets in underground space includes the following steps: Step 1: Obtain the image to be processed and input it into the target recognition model; Step 2: Through the small target framework network structure CRT-AWWR Performing a first processing on the image to be processed to obtain a first feature map; Step 3, pass SERT The module performs a second processing on the first feature map to obtain a second feature map; the second feature map is used to represent the attention result of the key information of the small target; Step 4: Adaptive Network Adaptive_gty Perform multi-level function fusion on the second feature map to obtain the third feature map; Step 5: Reconstruct the module through explicit features SJHD Resample the third feature map to obtain resampled data, and then obtain the reconstructed features and the corresponding fourth feature map; Step 6: After performing classic sampling, fusion and feature enhancement on the first, second, third and fourth feature maps, the output features are sent to the prediction head (target recognition model). softmax The function calculates the probability that the image to be processed belongs to each category; Step 7: Determine the category of the image to be processed based on the probability that the image to be processed belongs to each category.
[0006] Specifically, the target recognition model also includes a EFR The optimization loss function module of the underground space small target adaptive recognition method further includes the following steps: Step 8. Input the reconstruction features into the optimization loss function module to obtain the reconstruction feature loss; Step 9: Optimize the target recognition model based on the reconstruction feature loss.
[0007] Specifically, the small target framework network structure CRT-AWWR Including backbone network and feature pyramid network, the backbone network includes three classic VGGNet The feature pyramid network consists of five series-connected feature extraction layers. The feature extraction layer is based on Resnet The feature extractor of the convolutional neural network uses the output features of the feature pyramid network as the first feature map.
[0008] Specifically, the SERT module includes a first branch unit, a second branch unit, a third branch unit and a comprehensive calculation unit; The execution process of the first branch unit is: first, the tensor of the first feature map is converted from C × H × W Rotate to W × H × C , and then the rotated tensor W × H × C Execute in sequence Z-Pool Operation and parallel convolution operation, followed by activation function Sigmoid Process the results of parallel convolution to obtain attention weights, and finally rotate the attention weights to C × H × W The output of the first branch unit is obtained in the form of a first branch unit, and the output is used as the second feature map of the first level; The execution process of the second branch is: first, the tensor of the first feature map is converted from C × H × W Rotate to H × C × W , and then the rotated tensor H × C × W Execute in sequence Z-Pool Operation and parallel convolution operation, followed by activation function Sigmoid Process the results of parallel convolution to obtain attention weights, and finally rotate the attention weights to C × H × W The output of the second branch unit is obtained in the form of a second feature map of the second level; The execution process of the third branch unit is: first calculate the spatial attention of the first feature map to obtain the attention weight, then apply the attention weight to the first feature map and adjust the result after application to C × H ×W In this form, the output of the third branch unit is obtained and used as the second feature map of the third level; The comprehensive calculation unit averages the output of the first branch unit, the output of the second branch unit, and the output of the third branch unit to obtain a final second feature map; in: C Indicates the number of channels, H Indicates height, W Indicates width.
[0009] Specifically, the adaptive network Adaptive_gty introduces an adaptive weighting mechanism to perform multi-level function fusion on the second feature map, and introduces learnable weights at the fusion point. w i ; The process of the adaptive weighting mechanism is: , in, represents the third feature map obtained by adaptively weighting the second feature maps of levels 1, 2, and 3, P 1. P 2. P 3 represents the second feature map of the 1st, 2nd and 3rd levels respectively, w 1. w 2 and w 3 respectively represent P 1. P 2. P A weight of 3, ε Represents a constant parameter.
[0010] Specifically, through the adaptive network Adaptive_gty Before performing multi-level feature fusion on the second feature map, the sizes of the second feature maps at different levels are unified: , in: resize Represents a resizing operation, size Indicates the size acquisition operation. D Pj ( P i ) means P i Upsample to P j The size, Conv ( P i ) indicates P i Perform convolution operation.
[0011] Specifically, in Step 5, the explicit feature reconstruction module SJHDResample the third feature map to a small target and use the explicit feature reconstruction module SJHD Point resampling is used and offset is introduced. The characterization formulas include: , in: X represents the feature map of the third feature map, X’ Represents the restored feature map after dynamic upsampling, which is used to represent the resampled data. dy sample is the sampling function, S represents a set of sampling points, g represents the original grid of the third feature map, O dynamic Indicates the offset produced by the dynamic range factor; based on the resampled data X’ , obtain the reconstructed features and the corresponding fourth feature map.
[0012] Specifically, in Step 8, the process of obtaining feature reconstruction loss is as follows: EFR Loss function, add Wasserstein Distance metric, quantified using explicit feature reconstruction loss EFR The difference between the predicted object and the foreground classification of the true labeled object, using trans-entropy loss CEL The calculation is performed, and the characterization formula includes: , in: L EFR represents the feature reconstruction loss, H’ 、 W’ Represent the height and width of the fourth feature map respectively, i Indicates the fourth feature map i pixels, L represents the labeling confidence, p Confidence in model predictions.
[0013] An adaptive recognition model for small targets in underground space based on explicit feature reconstruction and anchor-free framework, comprising an image acquisition module, a first feature acquisition module, a second feature acquisition module, a third feature acquisition module, a resampling module, an output module and an optimization module; The image acquisition module is used to acquire the image to be processed; The first feature acquisition module is based on the small target framework network structure CRT-AWWR Performing a first processing on the image to be processed to obtain a first feature map; The second feature acquisition module is based on SERT The module performs a second processing on the first feature map to obtain a second feature map; The third feature acquisition module is based on an adaptive network Adaptive_gty Perform multi-level function fusion on the second feature map to obtain the third feature map; The resampling module is based on the explicit feature reconstruction module SJHD Resample the third feature map to obtain resampled data, and then obtain the reconstructed features and the corresponding fourth feature map; The output module performs classic sampling, fusion and feature enhancement on the first feature map, the second feature map, the third feature map and the fourth feature map, and sends them to the prediction head (target recognition model). softmax The function calculates the probability that the image to be processed belongs to each category and completes target recognition; The optimization module calculates the reconstruction feature loss based on the optimization loss function module, and optimizes the target recognition model based on the reconstruction feature loss.
[0014] Beneficial effects: The method and model for adaptive recognition of underground small targets based on explicit feature reconstruction without anchor framework provided by the present invention are compared with the existing technology: through the small target framework network structure CRT-AWWR By processing the image to be processed and obtaining the first feature map, the position and category of the object can be directly predicted at each position of the function map, avoiding the problem of sample imbalance, reducing the increase in overhead caused by anchor box-related calculations, simplifying the prediction process and improving detection performance; SERT The module processes the first feature map and obtains the second feature map used to characterize the key information of the small target. This process fully considers the multi-dimensional attention and the interaction between dimensions, pays more attention to the impact of the attention mechanism itself on the small target, and solves the problem of insufficient functional information of the small target caused by information loss in the subsequent feature extraction process, so as to retain more functional information for fusion; through the adaptive network Adaptive_gty Perform multi-level feature fusion on the second feature map to obtain the third feature map; through the explicit feature reconstruction module SJHD Resampling small targets in the third feature map eliminates the need to rely on surrounding pixels to calculate new pixel values, which can solve the problem of over-smoothing caused by the loss of small target information. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a structural diagram of the model of the present invention; Figure 2 Schematic diagram of the process of the present invention; Figure 3 Schematic diagram of the algorithm architecture of the method of the present invention; Figure 4 Frame network structure for small targets CRT-AWWR Schematic diagram of the structure; Figure 5 forSERT Schematic diagram of the module structure; Figure 6 Adaptive Network Adaptive_gty Schematic diagram of the structure; Figure 7 Reconstructing modules for explicit features SJHD Schematic diagram of the structure. DETAILED DESCRIPTION
[0016] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0017] like Figure 1 The figure shows a structural block diagram of an adaptive recognition model for small targets in underground space based on explicit feature reconstruction and anchor-free framework, which includes an image acquisition module, a first feature acquisition module, a second feature acquisition module, a third feature acquisition module, a resampling module, an output module and an optimization module; the image acquisition module is used to acquire the image to be processed; the first feature acquisition module is based on the small target framework network structure. CRT-AWWR The image to be processed is first processed to obtain a first feature map; the second feature acquisition module is based on SERT The module performs a second processing on the first feature map to obtain a second feature map; the third feature acquisition module is based on the adaptive network Adaptive_gty Perform multi-level function fusion on the second feature map to obtain the third feature map; the resampling module is based on the explicit feature reconstruction module SJHD The third feature map is resampled to obtain resampled data, and then the reconstructed features and the corresponding fourth feature map are obtained; the output module performs classical sampling, fusion and feature enhancement on the first feature map, the second feature map, the third feature map and the fourth feature map. softmax The function calculates the probability that the image to be processed belongs to each category to complete target recognition; the optimization module calculates the reconstruction feature loss based on the optimization loss function module, and optimizes the target recognition model based on the reconstruction feature loss.
[0018] like Figure 2 The figure shows the implementation process of an adaptive recognition method for underground small targets based on an explicit feature reconstruction anchor-free framework based on the above model. Figure 3 This is a schematic diagram of the algorithm architecture of the above model, which is based on an explicit feature reconstruction anchor-free framework for the adaptive recognition of underground small targets. CRT-AWWR 、 SERT Module, adaptive network Adaptive_gty And the small target resampling module realizes target recognition; the following is a further explanation of the present invention in combination with specific implementation steps.
[0019] Step 1. Get the image to be processed
[0020] The image to be processed is obtained through the image acquisition module and input into the underground small target adaptive recognition model based on explicit feature reconstruction and anchor-free framework.
[0021] Step 2: Get the first processing feature map
[0022] Through Figure 4 The small target framework network structure shown CRT-AWWR Perform a first processing on the image to be processed to obtain a first feature map.
[0023] The small target framework network structure CRT-AWWR Including backbone network and feature pyramid network, the backbone network includes three classic VGGNet The feature pyramid network consists of five series-connected feature extraction layers. The feature extraction layer is based on The feature extractor of the convolutional neural network uses the output features of the feature pyramid network as the first feature map.
[0024] Different from the traditional anchor-based method, the small target frame network structure adopted in this case Resnet It does not use predefined anchors, but it can directly predict the location and category of the object at each position in the feature map; this feature avoids the problem of sample imbalance and the increased computational overhead associated with anchor boxes. At the same time, the small target framework network structure adopted in this case CRT-AWWR As with anchor-based methods, a single object is predicted at each location, rather than multiple possible objects; this can simplify the prediction process and help improve detection performance.
[0025] Step 3. Obtain the second processing feature map
[0026] Since the target size is small, using the convolution with kernel size 7 in the conventional triplet attention module may result in a large receptive field, which may cause the detailed information of the small target to be lost. Therefore, this paper improves the conventional triplet attention module and obtains the following CRT-AWWR Figure 5 shown SERT Module. SERT The module performs a second processing on the first feature map to obtain a second feature map; the second feature map is used to characterize the attention result of the key information of the small target. SERT The module includes a first branch unit, a second branch unit, a third branch unit and a comprehensive calculation unit; using C Indicates the number of channels, H Indicates height, W Indicates width.
[0027] 3.1 The first branch
[0028] The first branch unit is used to capture the channelC and height H The interaction between Permute Filter module along H Axis will be the tensor of the first feature map C × H × W Rotate 90° counterclockwise to get a new tensor W × H × C ; After this, first in H Dimension execution Z-Pool Operation, then through C_Block The module performs parallel convolution operations and uses activation functions Sigmoid to be processed.
[0029] Z-Pool The operations focus on max pooling and averaging input sets, and their characterization formulas include: , in: x express Z-pool The input of the operation, in this branch, x Representing a tensor W × H × C ; y 1 express Z-pool Output of the operation; 0 d Represents the maximum pooling operation MaxPool and average pooling operation AvgPool Position 0.
[0030] C_Block The module combines parallel convolutions with kernel sizes of 1, 3, and 7 to perform convolution, which is an activation function. Sigmoid Provide inputs of different sizes; C_Block The execution process of the module can be expressed as: , in: y 1 express C_Block The input of the module, that is, Z-pool The output of the operation; y 2 express C_Block Output of the module; 2 that represents kernel sizes of 1, 3, and 7 respectively D Convolution, three 2 D The stride of convolution is 1.
[0031] C_Block The output of the module is passed through the activation function SigmoidGenerate attention weights and apply attention weights to Z- pool On the input of the operation, use Permute The filter module rotates to C × H × W Finally, the output of the first branch unit is obtained and used as the second feature map of the first level. The process is expressed as: , in: y 2 Represents the activation function Sigmoid Input, Permute Indicates rotating the input tensor 90° counterclockwise. y 3 Indicates the output of this branch.
[0032] 3.2 Second Branch
[0033] The second branch unit is used to capture the channel C and width W The interaction between Permute Filter module along W Axis will be the tensor of the first feature map C × H × W Rotate 90° counterclockwise to get a new tensor H × C × W ; After that, similar to the processing of the first branch, first W Dimension execution Z-pool Operation (in this branch, Z-pool Operation Input x Representing a tensor H × C × W ), and then through C_Block The module performs parallel convolution operations and uses activation functions Sigmoid Processing, generating attention weights, and applying attention weights to Z-pool On the input of the operation, use Permute The filter module rotates to C × H × W Finally, the output of the second branch unit is obtained and used as the second feature map of the second level.
[0034] 3.3 The third branch
[0035] In the third branch unit, the tensor of the first feature map C × H ×W Directly Z-pool The input of the operation, after which, similar to the processing of the first branch, is passed C_Block The module performs parallel convolution operations and uses activation functions Sigmoid Processing, generating attention weights, and applying attention weights to Z-pool On the input of the operation, no need to use Permute The filter module rotates, and the output of this branch after the attention weight is applied is consistent with C × H × W Therefore, the output of the attention weight of this branch is directly used as the output of the third branch unit, and the output is used as the second feature map of the third level.
[0036] 3.4 Comprehensive computing unit
[0037] The output of the first branch unit, the output of the second branch unit, and the output of the third branch unit are averaged to obtain a final second feature map.
[0038] The case adopted SERT Module: First, the scale characteristics of small objects are taken into account and the attention mechanism is improved to avoid the use of huge convolution kernels; second, the attention of triplets is selected as the basis for improvement, which fully considers multi-dimensional attention and the interaction between dimensions; finally, the polymorphic embedding is integrated into SERT This module focuses on the key information of small objects during feature extraction and suppresses irrelevant information, thereby improving detection accuracy. Before applying the channel attention mechanism to enhance model representation, this case introduces dilated convolutions to capture contextual information. This polymorphism emphasizes the impact of the attention mechanism itself on small objects and addresses information loss during subsequent feature extraction, making it more practical.
[0039] Step 4. Get the third feature map
[0040] go through SERT After the processing of the module, small targets will occupy fewer pixels in the deep feature map, which may lead to the loss of feature information; in contrast, shallow features without multiple sampling retain more texture feature information of small targets. Therefore, balancing the information from different second feature maps is crucial for accurate detection of small targets. Figure 6 The adaptive network shown Adaptive_gty ,Using the idea mechanism of optimizing queen fusion, the second feature maps of each level are selected and integrated with the local feature maps to obtain the third feature map.
[0041] The adaptive network Adaptive_gtyIntroduce an adaptive weighting mechanism to perform multi-level function fusion on the second feature map, and introduce learnable weights at the fusion point w i ; The process of the adaptive weighting mechanism is: , in, represents the third feature map obtained by adaptively weighting the second feature maps of levels 1, 2, and 3, P 1. P 2. P 3 represents the second feature map of the 1st, 2nd and 3rd levels respectively, w 1. w 2 and w 3 respectively represent P 1. P 2. P The weights of 3 are usually 0.2, 0.3, and 0.5 respectively. ε Represents a constant parameter.
[0042] Through adaptive network Adaptive_gty Before performing multi-level feature fusion on the second feature map, the sizes of the second feature maps at different levels are unified: , in: resize Represents a resizing operation, size Indicates the size acquisition operation. D Pj ( P i ) means P i Upsample to P j The size, Conv ( P i ) indicates P i Perform convolution operation.
[0043] Step 5. Obtain the reconstructed features and the corresponding fourth feature map
[0044] Since small objects lack sufficient appearance information, such as color, texture, shape, etc., it is difficult for detection algorithms to distinguish small objects with limited features. In addition, the tiny size and modest features of small objects make them easily confused with complex backgrounds. RSI This hinders the detection model from accurately identifying these objects, resulting in missed detection instances. Therefore, this case designs an explicit feature reconstruction module, namely the explicit feature reconstruction module SJHDTo improve small object detection performance, this module is designed as an auxiliary branch during training. Its main purpose is to resample small objects in the third feature map, helping the object recognition model better learn and retain the characteristic information of small objects during training. The resampled data is then used to obtain reconstructed features and the corresponding fourth feature map. By introducing this branch during training, the network is able to learn how to extract features of small objects from the third feature map.
[0045] like Figure 7 As shown, the explicit feature reconstruction module SJHD Through concise convolutional layers and dynamic upsamplers UPSMPLING To achieve effective feature reconstruction; this design is not based on the traditional kernel method, but adopts a method of focused sampling and introducing an offset. The offset is used to adjust the position of the sampling point, which helps to better reflect the relationship between pixels in the input feature map and prevent the loss of restored feature map information. Explicit Feature Reconstruction Module SJHD The execution process is: C × H 1× W 1's third feature map feature map X , after a dynamic upsampler UPSMPLING Then, a set of sampling points is generated S , sampling point set S The size is 2× H 2× W 2. Sampling function dy sample Using a sampling point set S Sampling is performed at the sampling position in to obtain the restored feature map after dynamic upsampling X’ (The structure of this process is from the literature W. Liu, H. Lu, H. Fu, and Z. Cao, "Learning to upsample by learning to sample ” IEEE / CVF Int. Conf. Comput. Vis. ( ICCV ) , Oct. 2023,pp. 6027–6037. );like Figure 7 As shown, the process can be summarized using the following formula: , in, X represents the feature map of the third feature map, X’ Represents the restored feature map after dynamic upsampling, which is used to represent the resampled data. dy sample is the sampling function, S represents a set of sampling points, g represents the original grid of the third feature map, O dynamic Indicates the offset produced by the dynamic range factor.
[0046] Throughout the process, the dynamic upsampler UPSMPLING There are two methods to generate sampling points, one is static range factor and the other is dynamic range factor.
[0047] Static range factors are created by combining linear layers and pixel flushing techniques with fixed range factors, which are then added to the original mesh. g To obtain a set of sampling points S , thus generating an offset. The offset generated by the static range factor is as follows: , in: O static represents the offset produced by the static range factor, pixel represents pixel shuffling, linear represents pixel transformation, X represents the feature map of the third feature map, β Represents the static factor.
[0048] The dynamic range factor adjusts the original sampling point by introducing a fixed offset. The dynamic range factor can adjust the offset more flexibly so that the sampling point position can be dynamically adjusted according to the content of the function map. The process of generating the offset using the dynamic range factor is: , in: O dynamic represents the offset produced by the dynamic range factor, linear 1 and linear 2 represent linear transformation 1 and linear transformation 2 respectively.
[0049] Based on resampled data X’ , the reconstructed features and the corresponding fourth feature map can be obtained.
[0050] In this case, the dynamic range factor is used to generate the sampling point set considering the impact of upsampling on the feature information of small targets during the fusion process. S The dynamic range factor adaptively adjusts the range of sampling points based on the different characteristics of the input feature map, thereby better learning the relationship between feature map pixels and then recovering the feature map. Unlike most networks that use bilinear interpolation for upsampling, this approach relies on surrounding pixels to calculate the new pixel value, reducing the loss of small object information and avoiding the over-smoothing that can result from this loss.
[0051] Step 6: Complete target recognition
[0052] After the first feature map, the second feature map, the third feature map and the fourth feature map are sampled, fused and enhanced, the output features are sent to the prediction head. softmax The function calculates the probability that the image to be processed belongs to each category; based on the probability that the image to be processed belongs to each category, the category of the image to be processed is determined.
[0053] Step 7: Model Optimization
[0054] The reconstructed features are input into the optimization loss function module to obtain the reconstruction feature loss; the target recognition model is optimized based on the reconstruction feature loss.
[0055] Small objects have limited location information and occupy part of the image; therefore, during training, small objects can be easily ignored or masked by large objects; this can easily lead to incorrect regression and affect model performance.
[0056] To overcome this, the baseline network usage will be CIOU and distribution focus loss ( DFL ) is combined with the loss function to guide the model to learn the location information of the target. However, IOU The sensitivity to various scale objects varies significantly. For small targets, small position deviations can greatly reduce IOU , which results in incorrect label assignment. However, for common targets, IOU It can better reflect the difference between the predicted box and the real box. Therefore, this case is based on EFR loss function, and added a [[ID=7 Distance metric, quantified using explicit feature reconstruction loss The difference between the predicted object and the foreground classification of the true labeled object, using trans-entropy loss The calculation is performed, and the characterization formula includes: , in: L EFR represents the feature reconstruction loss, H’ 、 W’ Represent the height and width of the fourth feature map respectively, i Indicates the fourth feature map i pixels, L represents the labeling confidence, p Confidence in model predictions.
[0057] This method can better focus on the boundary regression problem of small objects, optimize the small object bounding box regression task and improve the performance of the model.
[0058] The loss function is crucial in the training process of object detection models by guiding the model on how to correctly classify detected objects, pinpoint the location of objects, and estimate the confidence level of objects in the object detection process. The loss function constitutes the overall loss of the model and is used to continuously optimize the model's parameters during training iterations.
[0059] Step 8. Experimental Verification
[0060] The data set and experimental settings can verify the efficacy of the model. 、 As well as a self-constructed small target dataset, experiments on small targets were conducted in underground space.
[0061] (1) :This dataset was collected by Tianjin University The team collected and produced a dataset of images from drone photography. The dataset is divided into ten categories, including workers, boxes, pressure gauges, flames, shovels, road cones, construction fences, warning posts, and safety lights. The dataset contains 6,471 images in the training set and 548 images in the validation set.
[0062] (2) Self-built dataset: The images of the self-built dataset are from publicly available datasets. Object Detection V2 ( ) and on-site photos taken in the underground garage. The dataset contains 14 227 images captured by drones from different heights and angles. Image, ranging from 0.11 to 0.92 m , from 0° to 90°.
[0063] The dataset of workers, boxes, pressure gauges, flames, shovels, road cones, construction fences, warning posts, and safety lights was annotated to produce the final dataset, which contains 2,180 training images and 432 validation images.
[0064] The experimental training is carried out with 12.6 Memory The test was conducted on the 3060 server. E 5-2680 V 4 and 3060 The experimental software environment uses 11 operating systems, Deep learning frameworks, 3.10, 2.0 and other related toolkits.
[0065] Training parameter settings: the training epochs are set to 100, the batch size is set to 4, the initial learning rate is 0.001, and the Unless otherwise stated, the above experimental setup was used in subsequent experiments.
[0066] The same training parameters will be used to conduct ablation experiments on the improved modules. Each improved module is gradually merged and then 50, 50:95, , Floating point operation ( )( )and The indicators verified the effectiveness of the improved method.
[0067] According to experimental analysis, the robot can identify nine target categories: workers, boxes, pressure gauges, flames, shovels, road cones, construction fences, warning posts, and safety indicators. Threat targets are defined as flames and warning posts. The robot's target recognition accuracy is ≥ 91.6% (for a sample of ≥ 9 targets, including different types). The robot's target annotation offset is ≤ 0.5 meters, with a coverage rate of ≥ 95%.
[0068] Step 9. Performance Analysis
[0069] The proposed method for adaptive recognition of small targets in underground space based on explicit feature reconstruction and anchor-free framework is an optimized object detection model that can significantly enhance the baseline 8, especially when detecting small objects. By incorporating advanced data augmentation techniques, feature fusion strategies, and multi-scale training, compared to the original model, 8 achieved higher accuracy and reduced false positives. Experimental results show that the proposed method The improvement was 1% (from 0.78 to 0.79) and 6% (from 0.94 to 1.0) in accuracy, highlighting its improved accuracy in small object localization and its ability to identify objects in challenging scenes. These improvements are attributed to architectural modifications, including the integration of compressed excitation ( ) modules, advanced data augmentation techniques, and multi-scale training, which together enhance the robustness and effectiveness of the model in addressing the limitations of existing models.
[0070] These improvements are particularly valuable for applications requiring high accuracy, such as autonomous driving, early cancer detection, surveillance, and other real-time monitoring systems. Sensitivity analysis further validates that, while high parameter tuning contributes to the performance of the method, the improvements remain consistent across a range of reasonable variations. Ablation studies confirm that Blocks, data augmentation techniques and high parameter tuning contribute to 8 performance improvements, while The patches have the greatest impact on small object detection accuracy. This robustness reveals the importance of the proposed architectural modifications, e.g. Modular and multi-scale training, which is crucial for the enhanced performance of the model. This method establishes a solid foundation for further research and provides a practical and robust solution for real-world situations where accurate small object detection is crucial.
[0071] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form, and any technical solutions obtained by equivalent replacement or equivalent transformation fall within the scope of protection of the present invention.
Claims
1. A method for adaptive recognition of small underground targets based on explicit feature reconstruction and anchor-free framework, characterized by: Applied to target recognition model, the target recognition model includes a small target framework network structure CRT-AWWR 、 SERT Module, adaptive network Adaptive_gty and the small target resampling module, which includes an explicit feature reconstruction module SJHD The method for adaptively identifying small targets in underground space includes the following steps: Step 1: Obtain the image to be processed and input it into the target recognition model; Step 2: Through the small target framework network structure CRT-AWWR Performing a first processing on the image to be processed to obtain a first feature map; Step 3, pass SERT The module performs a second processing on the first feature map to obtain a second feature map; the second feature map is used to represent the attention result of the key information of the small target; Step 4: Adaptive Network Adaptive_gty Perform multi-level function fusion on the second feature map to obtain the third feature map; Step 5: Reconstruct the module through explicit features SJHD Resample the third feature map to obtain resampled data, and then obtain the reconstructed features and the corresponding fourth feature map; Step 6: After performing classic sampling, fusion and feature enhancement on the first, second, third and fourth feature maps, the output features are sent to the prediction head. softmax The function calculates the probability that the image to be processed belongs to each category; Step 7: Determine the category of the image to be processed based on the probability that the image to be processed belongs to each category.
2. The method for adaptively identifying small targets in underground spaces based on an anchor-free framework using explicit feature reconstruction according to claim 1 is characterized by: The target recognition model also includes a EFR The optimization loss function module of the underground space small target adaptive recognition method further includes the following steps: Step 8. Input the reconstruction features into the optimization loss function module to obtain the reconstruction feature loss; Step 9: Optimize the target recognition model based on the reconstruction feature loss.
3. The method for adaptively identifying small targets in underground spaces based on an anchor-free framework using explicit feature reconstruction according to claim 1 is characterized by: The small target framework network structure CRT-AWWR Including backbone network and feature pyramid network, the backbone network includes three classic VGGNet The feature pyramid network consists of five series-connected feature extraction layers. The feature extraction layer is based on Resnet The feature extractor of the convolutional neural network uses the output features of the feature pyramid network as the first feature map.
4. The method for adaptively identifying small targets in underground spaces based on an anchor-free framework using explicit feature reconstruction according to claim 1 is characterized by: described SERT The module includes a first branch unit, a second branch unit, a third branch unit and a comprehensive calculation unit; The execution process of the first branch unit is: first, the tensor of the first feature map is converted from C × H × W Rotate to W × H × C , and then the rotated tensor W × H × C Execute in sequence Z-Pool Operation and parallel convolution operation, followed by activation function Sigmoid Process the results of parallel convolution to obtain attention weights, and finally rotate the attention weights to C × H × W The output of the first branch unit is obtained in the form of a first branch unit, and the output is used as the second feature map of the first level; The execution process of the second branch is: first, the tensor of the first feature map is converted from C × H × W Rotate to H × C × W , and then the rotated tensor H × C × W Execute in sequence Z-Pool Operation and parallel convolution operation, followed by activation function Sigmoid Process the results of parallel convolution to obtain attention weights, and finally rotate the attention weights to C × H × W The output of the second branch unit is obtained in the form of a second feature map of the second level; The execution process of the third branch unit is: first calculate the spatial attention of the first feature map to obtain the attention weight, then apply the attention weight to the first feature map and adjust the result after application to C × H × W In this form, the output of the third branch unit is obtained and used as the second feature map of the third level; The comprehensive calculation unit averages the output of the first branch unit, the output of the second branch unit, and the output of the third branch unit to obtain a final second feature map; in: C Indicates the number of channels, H represents the height, and represents the width.
5. The method for adaptively identifying small targets in underground spaces based on an anchor-free framework using explicit feature reconstruction according to claim 4 is characterized in that: The adaptive network Adaptive_gty An adaptive weighting mechanism is introduced to perform multi-level function fusion on the second feature map, and a learnable weight is introduced at the fusion point. w i ; The process of the adaptive weighting mechanism is: , in: represents the third feature map obtained by adaptively weighting the second feature maps of levels 1, 2, and 3, P 1. P 2. P 3 represents the second feature map of the 1st, 2nd and 3rd levels respectively, w 1. w 2 and w 3 respectively represent P 1. P 2. P A weight of 3, ε Represents a constant parameter.
6. The method for adaptively identifying small targets in underground spaces based on an anchor-free framework using explicit feature reconstruction according to claim 5 is characterized by: Through adaptive network Adaptive_gty Before performing multi-level feature fusion on the second feature map, the sizes of the second feature maps at different levels are unified: , in: resize Represents a resizing operation, size Indicates the size acquisition operation. D Pj ( P i ) means P i Upsample to P j The size, Conv ( P i ) indicates P i Perform convolution operation.
7. The method for adaptively identifying small targets in underground spaces based on an anchor-free framework using explicit feature reconstruction according to claim 1 is characterized by: In Step 5, the explicit feature reconstruction module SJHD Resample the third feature map to a small target and explicitly reconstruct the feature map SJHD Point resampling is used and offset is introduced. The characterization formulas include: , in: X represents the feature map of the third feature map, X’ Represents the restored feature map after dynamic upsampling, which is used to represent the resampled data. dy sample is the sampling function, S represents a set of sampling points, g represents the original grid of the third feature map, O dynamic Indicates the offset produced by the dynamic range factor; based on the resampled data X’ , obtain the reconstructed features and the corresponding fourth feature map.
8. The method for adaptively identifying small targets in underground spaces based on an anchor-free framework using explicit feature reconstruction according to claim 2 is characterized by: In Step 8, the process of obtaining feature reconstruction loss is as follows: EFR Loss function, add Wasserstein Distance metric, quantified using explicit feature reconstruction loss EFR The difference between the predicted object and the foreground classification of the true labeled object, using trans-entropy loss CEL The calculation is performed, and the characterization formula includes: , in: L EFR represents the feature reconstruction loss, H’ 、 W’ Represent the height and width of the fourth feature map respectively, i Indicates the fourth feature map i pixels, L represents the labeling confidence, p Confidence in model predictions.
9. An adaptive recognition model for small targets in underground spaces based on an anchor-free framework using explicit feature reconstruction, characterized by: It includes an image acquisition module, a first feature acquisition module, a second feature acquisition module, a third feature acquisition module, a resampling module, an output module and an optimization module; The image acquisition module is used to acquire the image to be processed; The first feature acquisition module is based on the small target framework network structure CRT-AWWR Performing a first processing on the image to be processed to obtain a first feature map; The second feature acquisition module is based on SERT The module performs a second processing on the first feature map to obtain a second feature map; The third feature acquisition module is based on an adaptive network Adaptive_gty Perform multi-level function fusion on the second feature map to obtain the third feature map; The resampling module is based on the explicit feature reconstruction module SJHD Resample the third feature map to obtain resampled data, and then obtain the reconstructed features and the corresponding fourth feature map; The output module performs classical sampling, fusion and feature enhancement on the first feature map, the second feature map, the third feature map and the fourth feature map. softmax The function calculates the probability that the image to be processed belongs to each category and completes target recognition; The optimization module calculates the reconstruction feature loss based on the optimization loss function module, and optimizes the target recognition model based on the reconstruction feature loss.