A Small Target Detection Method Based on Hierarchical Focus Feature Pyramid
By adopting a hierarchical focus feature pyramid network in the field of computer vision, combining the hierarchical feature subtraction module and feature fusion guidance attention mechanism, the problems of small size, low resolution and background noise in small object detection are solved, and significant detection performance improvement and robustness enhancement are achieved.
Patent Information
- Application Number
- CN202311122506.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-01
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-09-01
AI Technical Summary
In the field of computer vision, small object detection faces problems such as small size, low resolution and background noise, which leads to increased detection difficulty and lacks rich labeled data sets, which affects the generalization ability of the algorithm.
A small object detection method based on a hierarchical focus feature pyramid is proposed. Through the hierarchical feature subtraction module and feature fusion guidance attention mechanism, the image detail features are enhanced, the small object detection performance is improved, and the global fusion feature guide focuses on effective information and suppress noise information.
It significantly improves the accuracy and robustness of small object detection, improves the performance and detection capabilities of the model, and can more effectively discover small object characteristics in the absence of rich data sets.
Smart Images

Figure CN117173396B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a small target detection method based on a hierarchical focused feature pyramid. Background Art
[0002] The current field of computer vision is in a stage of rapid development and wide application. Computer vision aims to enable computer systems to perceive, understand, and interpret image and video data, similar to the human visual system. It combines multiple fields such as computer science, machine learning, artificial intelligence, and image processing, aiming to build intelligent systems that can simulate and understand human vision. In recent years, with the popularization of digital cameras, smartphones, and video acquisition devices, a large amount of image and video data has been generated and recorded. The availability of this large-scale data has promoted the development and training of computer vision algorithms. At the same time, the rise of deep learning has enabled neural networks to learn high-level feature representations from raw data, and these features can be used for tasks such as image classification, object detection, image segmentation, and pose estimation. The rapid development of computer vision has also prompted it to play an important role in various fields, such as medical image diagnosis, video surveillance and security, object recognition, image search, autonomous driving, augmented reality and virtual reality, industrial automation, etc.
[0003] In the field of computer vision, object detection is an important task, aiming to automatically identify and locate specific target objects in images or videos through computer algorithms. Object detection plays a key role in many practical applications, including autonomous driving, face detection, abnormal behavior detection, video surveillance, industrial defect detection, robot navigation, and image search. With the rapid development of computer hardware and the rise of deep learning methods, object detection has made significant progress in the past few years. Traditional object detection methods usually rely on hand-designed features and classifiers, such as Haar features and support vector machines. However, these methods often require a large amount of manual effort and domain knowledge, and their performance is limited in complex scenarios. The rise of deep learning technology has brought new breakthroughs to object detection. In particular, the emergence of Convolutional Neural Networks (CNNs) has enabled computers to learn to extract meaningful features from raw pixel data. By training on large-scale labeled datasets, deep learning models can learn complex image feature representations and accurately detect and locate targets of different categories.
[0004] As an important branch of object detection, small object detection refers to the task of detecting and locating small-sized target objects in the field of computer vision. In the definition of the COCO dataset (Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: ECCV. pp. 740-755. Springer, 2014), small objects refer to objects with a size smaller than 32×32, which is also a widely recognized small object size standard. Objects of this size are common in remote sensing images, autonomous driving, video surveillance, and security, and have very important research significance in many practical scenarios such as remote sensing detection, disaster rescue, intelligent transportation, and environmental monitoring.
[0005] However, small objects usually have small sizes and low resolutions, which makes it difficult to clearly distinguish and accurately locate the objects in the image. Small objects may only occupy a very small part of the image, and their features and textures may not be obvious, so they are easily submerged by background noise or misjudged as other objects. Moreover, small objects are often blocked by other objects or the background, which increases the difficulty of object recognition and localization. In addition, there may be multiple small objects in the image, and the mutual interference between them also increases the complexity of the detection task. More importantly, relatively few large-scale labeled datasets for small objects are available, resulting in a lack of rich data resources for training and evaluation. Datasets lacking diversity and coverage make it more challenging to evaluate the generalization ability and performance of algorithms.
[0006] To overcome the difficulties and challenges in small object detection tasks, researchers have proposed many innovative methods and technologies, which can be summarized as various methods such as data augmentation, multi-scale strategies, feature fusion, super-resolution, and context information modeling. Among them, feature fusion is a common technique to improve the performance of small object detection, which is mainly reflected in the Feature Pyramid Network (FPN) commonly used in detectors. In traditional Convolutional Neural Networks (CNNs), the feature maps at the bottom layer often have high resolutions and rich detailed information, which are suitable for detecting small objects; while the feature maps at the top layer have low resolutions and semantic information, which are suitable for detecting large objects. By fusing features at different levels through FPN, a set of feature maps with multi-scale information can be obtained, improving the detection accuracy of small objects and providing more comprehensive semantic information, thereby enhancing the performance and robustness of the model.
[0007] To further improve the feature fusion ability of FPN, a series of improved feature pyramid networks have been continuously proposed. Liu et al. (Liu, S., Qi, L., Qin, H., Shi, J., Jia, J.: Path aggregation network for instance segmentation. In: CVPR. pp. 8759 - 8768, 2018) first proposed PAFPN, which introduced a bottom - up fusion branch on the basis of the top - down branch of FPN to help the bottom - layer information better transmit to the top layer. Ghiasi et al. (Ghiasi, G., Lin, T.Y., Le, Q.V.: Nas - fpn: Learning scalable feature pyramid architecture for object detection. In: CVPR. pp. 7036 - 7045, 2019) proposed NASFPN, which automatically searches for the optimal feature fusion method through neural architecture search (NAS) to improve the performance of object detection. In addition, Wang et al. (Wang, J., Sun, K., Cheng, T., Jiang, B., Deng, C., Zhao, Y., Liu, D., Mu, Y., Tan, M., Wang, X., et al.: Deep high - resolution representation learning for visual recognition. TPAMI 43(10), 3349 - 3364, 2020) proposed HRFPN, which achieves the purpose of exchanging multi - resolution representation information by repeatedly fusing multi - resolution features. Recently, Park et al. (Park, H.J., Kang, J.W., & Kim, B.G.: ssFPN: Scale Sequence(S2) Feature - Based Feature Pyramid Network for Object Detection. Sensors, 23(9), 4432, 2023) proposed ssFPN, a method specifically for small object detection, which finds the scale - invariant features of FPN by introducing the scale sequence space to improve the detection performance of small objects. However, the fusion strategies adopted by these methods are all based on the addition of corresponding elements of different - level features. This simple fusion method combines features of different resolutions. However, small object information often only exists in the bottom - layer feature maps. When they are fused with high - level information, they are easily overwhelmed by the main information. To extract detailed features beneficial to small object detection, a more reasonable fusion strategy can be considered.The introduction of feature subtraction and attention mechanism can help FPN strengthen and focus on small object features without compromising the detection performance of other scale objects. Summary of the Invention
[0008] The purpose of the present invention is to provide a small object detection method based on a hierarchical focused feature pyramid, abbreviated as HFFPN. The hierarchical focused feature pyramid enhances the detailed features of the image while retaining the main features of the image to improve the small object detection performance. Moreover, each level of features is guided by the globally fused features and focused on the effective information, helping to discover the information beneficial to small object detection and suppressing the invalid noise information.
[0009] The present invention includes the following steps:
[0010] Step 1, preprocess the image to be detected, and send the preprocessed image to be detected and its corresponding image-level label into the neural network;
[0011] Step 2, the neural network extracts features from the image and sends the features into the hierarchical focused feature pyramid for fusion;
[0012] Step 3, in the training stage, the model uses the features obtained after fusion to output the position and category of the target in the image to be detected;
[0013] Step 4, in the testing stage, the image to be detected also goes through feature extraction and then enters the hierarchical focused feature pyramid, and uses the fused features to output the coordinates, category, and score of the predicted box of the image to be detected.
[0014] In step 1, for the preprocessing, the image can be first normalized, then scaled to a size of 256×256, and finally randomly cropped to a size of 224×224;
[0015] In step 2, the hierarchical focused feature pyramid includes: a top-down feature fusion branch, a Hierarchical Feature Subtraction Module (HFSM), and a Feature Fusion Guidance Attention (FFGA); the Hierarchical Feature Subtraction Module (HFSM) introduces a feature subtraction operation to obtain detailed information and enhances this part of the information, thereby improving the detection ability for small targets; at the same time, to avoid the loss of main body information caused by the feature subtraction operation in high-level semantic information, the Hierarchical Feature Subtraction Module (HFSM) introduces a hierarchical strategy to control the feature subtraction operation between effective feature layers. The Feature Fusion Guidance Attention FFGA is a generalized self-attention mechanism that uses the global information during feature fusion to guide the features of this layer to focus on effective information, suppress noise information, and guide the discovery of small target features.
[0016] The neural network extracts features from the picture and sends the features into the hierarchical focused feature pyramid for fusion, including the following steps:
[0017] Step a1, given a dataset set with image-level labels, divide the set into a training picture sample set and a test picture sample set;
[0018] Step a2, arbitrarily select an image I from the training picture sample set, and input the image I and its corresponding image-level label y into the feature extraction backbone network of the neural network to obtain the image features C of each layer extracted by the network. The specific operation is as follows:
[0019]
[0020] Among them, is a convolution module in the feature extraction backbone network, and t is the number of feature layers;
[0021] Step a3, the image feature C passes through the Hierarchical Feature Subtraction Module HFSM. During the top-down fusion process of C, a hierarchical feature subtraction operation is adopted to obtain the intermediate feature M. The specific operation is as follows:
[0022]
[0023] Among them, σ(·) is a 1×1 convolutional layer, UP(·) is an upsampling operation with a ratio of 2. ⊕ and respectively represent element-wise addition and element-wise subtraction. |·| represents the absolute value operation, and l is a hyperparameter that implements the hierarchical strategy, which determines between which feature layers the hierarchical operation is performed;
[0024] Step a4, the intermediate feature M undergoes further feature fusion through a 3×3 convolutional layer to obtain the fused feature P. The specific operation is as follows:
[0025] P i = conv 3×3 (M i )
[0026] Step a5, the fused feature P is fed into the feature fusion guiding attention. For a certain level of feature take its upper-level feature after upsampling it, and concatenate it with P in the channel dimension i to obtain the fused guiding feature The specific operation is as follows:
[0027] F g = concat(P i , UP(P i+1 ))
[0028] where concat(·) represents the operation of concatenating two feature maps in the channel dimension.
[0029] Step a6, the fused guiding feature F g successively passes through the channel attention module and the spatial attention module to obtain the attention feature The specific operation is as follows:
[0030]
[0031]
[0032] where F ca is the channel attention feature obtained by F a passing through the channel attention module, and it is immediately fed into the spatial attention module to obtain the attention feature F a . avgpool1(·) and avgpool2(·) respectively represent the average pooling operations in the spatial dimension and in the channel dimension. CR(·) and CS(·) are respectively 1×1 convolutions with ReLU activation layers and 1×1 convolutions with sigmoid activation layers. δ(·) represents the operation of extending in the dimension;
[0033] Step a7, F a passes through a 1×1 convolution to obtain an attention weighted map with a dimension of 1
[0034] W a = conv 1×1 (F a )
[0035] Step a8, attention weighted graph W a is multiplied by the feature at this level to obtain the final output feature O. The specific operation is as follows:
[0036]
[0037] In step 3, during the training phase, the model uses the features obtained after fusion to output the position and category of the target in the image to be detected. That is, the model uses the output feature O to guide the regression of the detection box, predicts the target category, and compares the results of regression and classification with the label y, so as to continuously update the model parameters for learning.
[0038] In step 4, during the testing phase, the image to be detected is input into the network. After feature extraction, it enters the hierarchical focused feature pyramid. After the hierarchical focused feature pyramid obtains features containing more small target information, the model uses these features to output the coordinates, category, and score of the predicted box.
[0039] Compared with the prior art, the present invention has the following outstanding advantages:
[0040] 1. The present invention proposes a hierarchical focused feature pyramid network (HFFPN) for the characteristics of small targets. While achieving multi-scale detection, it pays more attention to the utilization and exploration of small target features. HFFPN can be easily applied to various detectors to improve the performance of the detectors.
[0041] 2. The present invention designs a hierarchical feature subtraction module (HFSM), which adopts a feature subtraction operation completely different from the previous feature fusion design. This makes full use of the difference information between feature layers and helps to improve the performance of small target detection. At the same time, the introduction of the hierarchical strategy improves the robustness of the model.
[0042] 3. The present invention introduces a feature fusion guided attention (FFGA), which uses more global information after fusion to guide the refinement of the feature map. The attention mechanism used weights its own features, focuses on useful information, suppresses invalid information, and helps to explore potential small target information.
[0043] 4. Extensive experiments conducted on the DOTA and COCO datasets show that the use of HFFPN has greatly improved the performance of the baseline detector. Compared with other competing methods, the HFFPN proposed by the present invention has achieved significant and consistent performance improvements, and the detection ability of small targets has been significantly improved. Description of the Drawings
[0044] Figure 1 is a schematic diagram of the network structure of the hierarchical focused feature pyramid (HFFPN) of the present invention.
[0045] Figure 2 It is a schematic diagram of the network structure of Feature Fusion Guided Attention (FFGA) of the present invention. Specific Embodiment
[0046] The following embodiments will describe in detail the technical solutions and beneficial effects of the present invention in conjunction with the accompanying drawings.
[0047] An embodiment of the present invention is a small target detection method based on a hierarchical focused feature pyramid, abbreviated as HFFPN, which enhances the information of small targets during the feature fusion process to improve the detection performance. First, the designed hierarchical feature subtraction module HFSM introduces a feature subtraction operation to obtain detailed information and strengthens this part of the information, thereby improving the detection ability for small targets. At the same time, in order to avoid the loss of main body information caused by the feature subtraction operation in the high-level semantic information, HFSM introduces a hierarchical strategy to control the feature subtraction operation between effective feature layers. Secondly, the introduced Feature Fusion Guided Attention FFGA is a generalized self-attention mechanism, which uses the global information during feature fusion to guide the features of this layer to focus on effective information, suppress noise information, and guide the discovery of small target features.
[0048] The embodiment of the present invention specifically includes the following steps:
[0049] Step 1, preprocess the image to be detected, and send the preprocessed image to be detected and its corresponding image-level label into the neural network;
[0050] Step 2, the neural network extracts features from the image and sends the features into the hierarchical focused feature pyramid for fusion;
[0051] Step 3, in the training stage, the model uses the features obtained after fusion to output the position and category of the target in the image to be detected;
[0052] Step 4, in the test stage, the image to be detected also undergoes feature extraction and then enters the hierarchical focused feature pyramid, and uses the fused features to output the coordinates, category, and score of the predicted box of the image to be detected.
[0053] In step 1, for the preprocessing, the image can be first normalized, then scaled to a size of 256×256, and finally randomly cropped to a size of 224×224;
[0054] In step 2, as Figure 1As shown, the hierarchical focused feature pyramid includes: a top-down feature fusion branch, a Hierarchical Feature Subtraction Module (HFSM), and a Feature Fusion Guidance Attention (FFGA); the Hierarchical Feature Subtraction Module (HFSM) introduces a feature subtraction operation to obtain detailed information and enhances this part of the information, thereby improving the detection ability for small targets; at the same time, to avoid the loss of main body information caused by the feature subtraction operation in high-level semantic information, HFSM introduces a hierarchical strategy to control the feature subtraction operation to be carried out between effective feature layers; the Feature Fusion Guidance Attention FFGA is a generalized self-attention mechanism that utilizes the global information during feature fusion to guide the features of this layer to focus on effective information, suppress noise information, and guide the discovery of small target features.
[0055] The feature fusion of the feature extraction and the hierarchical focused feature pyramid includes the following steps:
[0056] Step a1, given a dataset set with image-level labels, divide the set into a training image sample set and a test image sample set;
[0057] Step a2, arbitrarily select an image I from the training image sample set, and input the image I and its corresponding image-level label y into the feature extraction backbone network of the neural network to obtain the image features C of each layer extracted by the network. The specific operation is as follows:
[0058]
[0059] Among them, is a convolutional module in the feature extraction backbone network, and t is the number of feature layers;
[0060] Step a3, the image feature C passes through the hierarchical feature subtraction module HFSM. During the top-down fusion of C, a hierarchical feature subtraction operation is performed to obtain the intermediate feature M. The specific operation is as follows:
[0061]
[0062] Among them, σ(·) is a 1×1 convolutional layer, and UP(·) is an upsampling operation with a ratio of 2; and respectively represent element-wise addition and element-wise subtraction; |·| represents the operation of taking the absolute value, and l is a hyperparameter that implements the hierarchical strategy, which determines between which feature layers the hierarchical operation is carried out;
[0063] Step a4, the intermediate feature M undergoes further feature fusion through a 3×3 convolutional layer to obtain the fused feature P. The specific operation is as follows:
[0064] P i = conv 3×3 (M i )
[0065] Step a5, as Figure 2 shown, the fused feature P is fed into the feature fusion guiding attention. For a certain level of feature take its upper-level feature after upsampling it, and concatenate it with P i in the channel dimension to obtain the fused guiding feature The specific operation is as follows:
[0066] F g = concat(P i , UP(P i+1 ))
[0067] where concat(·) represents the operation of concatenating two feature maps in the channel dimension;
[0068] Step a6, as Figure 2 shown, the fused guiding feature F g successively passes through the channel attention module and the spatial attention module to obtain the attention feature The specific operation is as follows:
[0069]
[0070]
[0071] where F ca is the channel attention feature obtained by F a passing through the channel attention module, which is then immediately fed into the spatial attention module to obtain the attention feature F a ; avgpool1(·) and avgpool2(·) respectively represent average pooling operations in the spatial dimension and in the channel dimension; CR(·) and CS(·) are respectively 1×1 convolutions with ReLU activation layers and 1×1 convolutions with sigmoid activation layers; δ(·) represents the operation of extending in the dimension;
[0072] Step a7, F a passes through a 1×1 convolution to obtain an attention weighted map with a dimension of 1
[0073] W a = conv 1×1 (F a)
[0074] Step a8, attention weighted graph W a is multiplied by the feature at this level to obtain the final output feature O. The specific operation is as follows:
[0075]
[0076] In step 3, during the training phase, the model uses the feature obtained after fusion to output the position and category of the target in the image to be detected. That is, the model uses the output feature O to guide the regression of the detection box, predicts the target category, and compares the results of regression and classification with the label y, so as to continuously update the model parameters for learning;
[0077] In step 4, during the testing phase, the image to be detected is input into the network. After feature extraction, it enters the hierarchical focused feature pyramid. After the hierarchical focused feature pyramid obtains the features containing more small target information, the model uses these features to output the coordinates, category, and score of the predicted box.
[0078] The effect of the present invention is further illustrated by the following simulation experiments.
[0079] 1) Simulation conditions
[0080] The present invention is developed on the Ubuntu platform, and the developed deep learning framework is based on Pytorch. The main language used in the present invention is Python.
[0081] 2) Simulation content
[0082] The small target dataset DOTA and the COCO dataset are taken. The network is trained according to the above steps and tested using the test set. Tables 1 and 2 show the detection results of the present invention and other methods on the two datasets respectively. It can be seen that the application of the present invention has consistently improved the detection performance of the baseline algorithm. On the DOTA dataset, compared with other advanced methods, the present invention has the best effect. Among them, the backbone networks using R50-HFFPN and R101-HFFPN are the results of the present invention. The evaluation index mAP represents the average detection performance of the algorithm for various targets, mAP sThe average detection performance of the presented algorithm for small targets. The detection performance of this method on the DOTA dataset reaches 76.64% / 76.89% (ResNet50 / ResNet101), which is higher than that of other advanced methods; on the COCO dataset, three representative two-stage object detection algorithms (Faster RCNN), single-stage object detection algorithms (RetinaNet), and anchor-free object detection algorithms (FCOS) have all achieved consistent performance improvements in small target detection metrics, demonstrating the application potential of the present invention in detectors and having better effects in small target detection.
[0083] Table 1 Comparison with the latest technical methods on the DOTA dataset
[0084]
[0085] Table 2 Performance improvement of three representative algorithms on the COCO dataset
[0086]
[0087] Extensive experiments conducted on the DOTA and COCO datasets show that the proposed HFFPN achieves significant and consistent performance improvements compared to other competing methods.
[0088] References:
[0089] [1] Chen, Z., Chen, K., Lin, W., See, J., Yu, H., Ke, Y., Yang, C.: Piou loss: Towards accurate oriented object detection in complex environments. In: ECCV. pp. 195 - 211. Springer (2020).
[0090] [2] Ding, J., Xue, N., Long, Y., Xia, G.S., Lu, Q.: Learning roi transformer for oriented object detection in aerial images. In: CVPR. pp. 2849 - 2858 (2019).
[0091] [3]Guo, Z., Liu, C., Zhang, X., Jiao, J., Ji, X., Ye, Q.: Beyond bounding-box: Convex-hull feature adaptation for oriented and densely packed object detection. In: CVPR. pp. 8792-8801 (2021).
[0092] [4]Han, J., Ding, J., Li, J., Xia, G.S.: Align deep features for oriented object detection. IEEE Transactions on Geoscience and Remote Sensing 60, 1-11 (2021).
[0093] [5]Han, J., Ding, J., Xue, N., Xia, G.S.: Redet: A rotation-equivariant detector for aerial object detection. In: CVPR. pp. 2786-2795 (2021).
[0094] [6]Li, W., Chen, Y., Hu, K., Zhu, J.: Oriented reppoints for aerial object detection. In: CVPR. pp. 1829-1838 (2022).
[0095] [7]Lin, T.Y., Goyal, P., Girshick, R., He, K., Dollár, P.: Focal loss for dense object detection. In: ICCV. pp. 2980-2988 (2017).
[0096] [8]Pan, X., Ren, Y., Sheng, K., Dong, W., Yuan, H., Guo, X., Ma, C., Xu, C.: Dynamic refinement network for oriented and densely packed object detection. In: CVPR. pp. 11207-11216 (2020).
[0097] [9]Qian, W., Yang, X., Peng, S., Yan, J., Guo, Y.: Learning modulated loss for rotated object detection. In: AAAI. vol. 35, pp. 2458-2466 (2021).
[0098]
[10] Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28 (2015).
[0099]
[11] Tian, Z., Shen, C., Chen, H., He, T.: Fcos: Fully convolutional one-stage object detection. In: ICCV. pp. 9627-9636 (2019).
[0100]
[12] Wei, H., Zhang, Y., Chang, Z., Li, H., Wang, H., Sun, X.: Oriented objects as pairs of middle lines. ISPRS Journal of Photogrammetry and Remote Sensing 169, 268-279 (2020).
[0101]
[13] Xie, X., Cheng, G., Wang, J., Yao, X., Han, J.: Oriented r-cnn for object detection. In: ICCV. pp. 3520-3529 (2021).
[0102]
[14] Xu, Y., Fu, M., Wang, Q., Wang, Y., Chen, K., Xia, G.S., Bai, X.: Gliding vertex on the horizontal bounding box for multi-oriented object detection. TPAMI 43(4), 1452-1459 (2020).
[0103]
[15] Yang, X., Hou, L., Zhou, Y., Wang, W., Yan, J.: Dense label encoding for boundary discontinuity free rotation detection. In: CVPR. pp. 15819-15829 (2021).
[0104]
[16] Yang, X., Yan, J., Feng, Z., He, T.: R3det: Refined single-stage detector with feature refinement for rotating object. In: AAAI. vol. 35, pp. 3163-3171 (2021).
[0105]
[17] Yang, X., Yang, J., Yan, J., Zhang, Y., Zhang, T., Guo, Z., Sun, X., Fu, K.: Scrdet: Towards more robust detection for small, cluttered and rotated objects. In: ICCV. pp. 8232-8241 (2019).
[0106] The above embodiments are only for illustrating the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the present invention.
Claims
1. A small target detection method based on a hierarchical focused feature pyramid, characterized in that it includes the following steps: Step 1, preprocess the image to be detected, and send the preprocessed image to be detected and its corresponding image-level label into the neural network; Step 2, the neural network extracts features from the image and sends the features into the hierarchical focused feature pyramid for fusion, including the following steps: Step a1, given a dataset set with image-level labels, divide the set into a training image sample set and a test image sample set; Step a2, arbitrarily select an image I from the training image sample set, and input the image I and its corresponding image-level label y into the feature extraction backbone network of the neural network to obtain the image features C of each layer extracted by the network. The specific operation is as follows: Among them, is the convolutional module in the feature extraction backbone network, and t is the number of feature layers; Step a3, the image feature C passes through the hierarchical feature subtraction module HFSM. During the top-down fusion process of C, a hierarchical feature subtraction operation is performed to obtain the intermediate feature M. The specific operation is as follows: Among them, σ(·) is a convolutional layer of size 1×1, and UP(·) is an upsampling operation with a ratio of 2; and represent element-wise addition and element-wise subtraction respectively; |·| represents the operation of taking the absolute value, and l is a hyperparameter that implements the hierarchical strategy and determines between which feature layers the hierarchical operation is performed; Step a4, the intermediate feature M passes through a convolutional layer with a size of 3×3 for further feature fusion to obtain the fused feature P. The specific operation is as follows: P i = conv 3×3 (M i ) Step a5, the fused feature P is fed into the feature fusion-guided attention, and for a certain level of feature among them Take its previous-level feature After upsampling it, connect it with P in the channel dimension i to obtain the fusion-guided feature The specific operation is as follows: F g = concat(P i , UP(P i+1 )) Among them, concat(·) represents the operation of concatenating two feature maps in the channel dimension; Step a6, fuse the guiding feature F g Successively pass through the channel attention module and the spatial attention module to obtain the attention feature The specific operation is as follows: Among them, F ca is the channel attention feature obtained by the channel attention module, which is then immediately fed into the spatial attention module to obtain the attention feature F a ; avgpool1(·) and avgpool2(·) respectively represent the average pooling operations in the spatial dimension and in the channel dimension; CR(·) and CS(·) are respectively 1×1 convolutions with ReLU activation layers and 1×1 convolutions with sigmoid activation layers; δ(·) represents the operation of extending in the dimension; a Step a7, F a Obtain an attention-weighted graph with a dimension of 1 through a 1×1 convolution W a = conv 1×1 (F a ) Step a8, attention weighted graph W a is multiplied by the features at this level to obtain the final output feature O. The specific operation is as follows: Step 3, in the training stage, the model uses the features obtained after fusion to output the position and category of the target in the image to be detected; Step 4, in the test stage, the image to be detected also enters the hierarchical focused feature pyramid after feature extraction, and uses the fused features to output the coordinates, category, and score of the predicted box of the image to be detected.
2. The small target detection method based on a hierarchical focused feature pyramid according to claim 1, characterized in that in Step 1, for the preprocessing, first perform standardization processing on the image, then scale the image to a size of 256×256, and finally randomly crop it to a size of 224×224.
3. The small target detection method based on a hierarchical focused feature pyramid according to claim 1, characterized in that in Step 2, the hierarchical focused feature pyramid includes: a top-down feature fusion branch, a hierarchical feature subtraction module HFSM, and a feature fusion guided attention FFGA; the hierarchical feature subtraction module HFSM introduces a feature subtraction operation to obtain detailed information and strengthens this part of the information to improve the detection ability for small targets; at the same time, to avoid the loss of main body information caused by the feature subtraction operation in the high-level semantic information, the hierarchical feature subtraction module HFSM introduces a hierarchical strategy to control the feature subtraction operation to be performed between effective feature layers; the feature fusion guided attention FFGA is a generalized self-attention mechanism, which uses the global information during feature fusion to guide the features of this layer to focus on effective information, suppress noise information, and guide the discovery of small target features.
4. The small target detection method based on a hierarchical focused feature pyramid according to claim 1, characterized in that In step 3, during the training phase, the model uses the features obtained after fusion to output the position and category of the target in the image to be detected. That is, the model uses the output feature O to guide the regression of the detection box, predicts the target category, and compares the results of regression and classification with the label y, thereby continuously updating the model parameters for learning.
5. A small target detection method based on a hierarchical focus feature pyramid as described in claim 1, wherein In step 4, during the testing phase, the image to be detected is input into the network. After feature extraction, it also enters the hierarchical focus feature pyramid. After the hierarchical focus feature pyramid obtains features containing more small target information, the model uses these features to output the coordinates, category, and score of the predicted box.