A saliency-guided high-resolution small target detection method and system

CN122821076APending Publication Date: 2026-09-25NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511273789.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,目前的高分辨率特征图上的小目标检测面临计算成本过高,计算精度受限的问题

Benefits of technology

[0014]本申请提供了一种显著性引导高分辨率小目标检测方法及系统,通过训练好的目标检测模型对目标可见光图像进行小目标检测,得到目标可见光图像对应的分类检测预测结果,其中,训练好的目标检测模型为改进的RetinaNet模型,所述改进的RetinaNet模型包括改进的特征金字塔,所述改进的特征金字塔提取多个不同尺度的特征图,分别记为P2特征图、P3特征图、P4特征图、P5特征图、P6特征图、P7特征图,改进的RetinaNet模型,在P2特征图、P3特征图后分别连接显著性引导目标检测模块,在P4特征图、P5特征图、P6特征图、P7特征图后分别连接检测头;所述显著性引导目标检测模块,用于对P2特征图或P3特征图进行目标检测,得到P2特征图或P3特征图对应的检测结果,解决了现有高分辨率特征图上的小目标检测计算成本过高,计算精度受限的问题,降低了小目标检测的计算成本,提高了小目标检测的计算精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821076A_ABST
    Figure CN122821076A_ABST
Patent Text Reader

Abstract

The application discloses a saliency-guided high-resolution small target detection method and system, and relates to the field of computer vision.The method comprises the following steps: inputting a target visible light image into a trained target detection model to obtain a corresponding classification detection prediction result; the trained target detection model is an improved RetinaNet model, which comprises an improved feature pyramid; the improved feature pyramid extracts multiple feature maps of different scales, which are respectively denoted as a P2 feature map, a P3 feature map, a P4 feature map, a P5 feature map, a P6 feature map and a P7 feature map; the improved RetinaNet model is connected with a saliency-guided target detection module after the P2 feature map and the P3 feature map, and is connected with a detection head after the P4 feature map, the P5 feature map, the P6 feature map and the P7 feature map; and the application can reduce the calculation cost of small target detection and improve the calculation precision of small target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision, and in particular to a saliency-guided high-resolution small target detection method and system. Background Technology

[0002] Small object detection has always been a highly challenging problem in computer vision. Previous studies have attempted to optimize its performance through data augmentation, multi-scale feature fusion, and contextual information learning. Increasing the number of small objects in an image is an effective data augmentation strategy. KISANTAL et al. balanced samples of different scales by randomly pasting small objects of different sizes at different locations within the same image. BOSQUET et al. combined generative adversarial networks with object segmentation, image inpainting, and fusion techniques to generate high-quality synthetic data. Multi-scale feature fusion enables cross-level information complementarity. FPN employs a top-down and bottom-up feature pyramid network structure to integrate deep semantic information and shallow representation features of small objects. Building on this, PANet uses a bidirectional path architecture, allowing deep features to retain accurate localization information. Contextual information can provide clues about the surrounding environment of small objects. CoupleNet uses a dual-branch coupling module, where the local branch handles object-level details and the global branch integrates the scene background, significantly improving accuracy in complex scenes. While the above methods improve small object detection performance by introducing auxiliary features, they inevitably introduce additional noise into the system. In contrast, utilizing high-resolution images or feature maps can directly improve the effective resolution of small objects, thereby enhancing their feature representation and providing more reliable detection evidence. However, current small object detection on high-resolution feature maps faces the problems of excessively high computational costs and limited computational accuracy. Summary of the Invention

[0003] The purpose of this application is to provide a saliency-guided high-resolution small target detection method and system, which can reduce the computational cost of small target detection and improve the computational accuracy of small target detection.

[0004] To achieve the above objectives, this application provides the following solution:

[0005] Firstly, this application provides a saliency-guided high-resolution small target detection method, including:

[0006] Acquire a visible light image of the target;

[0007] The target visible light image is input into a trained target detection model to obtain the classification and detection prediction result corresponding to the target visible light image. The trained target detection model is an improved RetinaNet model, which includes an improved feature pyramid. The improved feature pyramid extracts multiple feature maps of different scales, denoted as P2 feature map, P3 feature map, P4 feature map, P5 feature map, P6 feature map, and P7 feature map, respectively. The resolution of the P2 feature map, P3 feature map, P4 feature map, P5 feature map, P6 feature map, and P7 feature map decreases sequentially.

[0008] The improved RetinaNet model connects a saliency-guided target detection module after the P2 and P3 feature maps, respectively, and a detection head after the P4, P5, P6, and P7 feature maps, respectively. The saliency-guided target detection module is used to perform target detection on the P2 or P3 feature maps to obtain the detection results corresponding to the P2 or P3 feature maps.

[0009] Secondly, this application provides a saliency-guided high-resolution small target detection system, comprising:

[0010] The target visible light image acquisition module is used to acquire the target visible light image;

[0011] The target detection module is used to input the visible light image of the target into a trained target detection model to obtain the classification and detection prediction result corresponding to the visible light image of the target. The trained target detection model is an improved RetinaNet model, which includes an improved feature pyramid. The improved feature pyramid extracts multiple feature maps of different scales, denoted as P2 feature map, P3 feature map, P4 feature map, P5 feature map, P6 feature map, and P7 feature map, respectively. The resolution of the P2 feature map, P3 feature map, P4 feature map, P5 feature map, P6 feature map, and P7 feature map decreases sequentially.

[0012] The improved RetinaNet model connects a saliency-guided target detection module after the P2 and P3 feature maps, respectively, and a detection head after the P4, P5, P6, and P7 feature maps, respectively. The saliency-guided target detection module is used to perform target detection on the P2 or P3 feature maps to obtain the detection results corresponding to the P2 or P3 feature maps.

[0013] According to the specific embodiments provided in this application, the following technical effects are disclosed:

[0014] This application provides a saliency-guided high-resolution small target detection method and system. A trained target detection model is used to detect small targets in a visible light image, yielding a classification and prediction result. The trained target detection model is an improved RetinaNet model, which includes an improved feature pyramid. The improved feature pyramid extracts multiple feature maps of different scales, denoted as P2, P3, P4, P5, P6, and P7 feature maps. The improved RetinaNet model connects a saliency-guided target detection module after the P2 and P3 feature maps, and a detection head after the P4, P5, P6, and P7 feature maps. The saliency-guided target detection module performs target detection on either the P2 or P3 feature map, obtaining the corresponding detection result. This method solves the problems of high computational cost and limited accuracy in small target detection on existing high-resolution feature maps, reducing the computational cost and improving the accuracy of small target detection. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is an application environment diagram of a saliency-guided high-resolution small target detection method in Embodiment 1 of this application.

[0017] Figure 2 This is a schematic flowchart of a saliency-guided high-resolution small target detection method provided in Embodiment 1 of this application.

[0018] Figure 3 This is a flowchart illustrating the saliency-guided target detection method provided in Embodiment 1 of this application.

[0019] Figure 4 This is a schematic diagram of the target detection model provided in Embodiment 1 of this application.

[0020] Figure 5 This is a schematic diagram of the existing RetinaNet model structure.

[0021] Figure 6 This is a schematic diagram of sparse convolution provided in Embodiment 1 of this application.

[0022] Figure 7This is a schematic diagram of a target visible light image provided in Embodiment 1 of this application.

[0023] Figure 8 A saliency diagram provided for Embodiment 1 of this application.

[0024] Figure 9 This is a schematic diagram showing the classification and detection prediction results of the existing target detection model and the target detection model of this application, as provided in Embodiment 1 of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] Incorporating image saliency into object detection helps improve detection accuracy. Aidan et al. integrated model saliency and human saliency into the loss function, guiding the model to focus on task-relevant regions perceived by human observers. Current deep learning-based architectures can also simulate human visual attention mechanisms to generate saliency maps. Song et al. used a saliency detection model at a coarse level to locate suspicious regions in remote sensing images, and then implemented an efficient network at a fine level to predict the object category and location within these regions. Besides serving as a priori condition for object selection, saliency maps can also enhance object features by fusing them with the original image or features. An et al. utilized a progressive attention mechanism combining voxels and saliency maps to suppress redundant background features and alleviate severe foreground-background imbalance. Shi et al. developed a unified detection framework that combines saliency maps generated from saliency detection with shallow features extracted by an object detection network. They also used a weight-sharing mechanism between the two tasks to achieve mutual reinforcement and performance improvement. These methods demonstrate the effectiveness of saliency-assisted object detection. Therefore, we believe that saliency can be used to guide small object detection in high-resolution feature maps. First, small targets are located through saliency detection, and then accurate detection is performed in the identified regions to optimize speed and accuracy.

[0027] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0028] Example 1

[0029] The saliency-guided high-resolution small target detection method provided in this application can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on another server. Terminal 102 can send a target visible light image to server 104. After receiving the target visible light image, server 104 inputs the target visible light image into a trained target detection model to obtain a classification detection prediction result corresponding to the target visible light image. Server 104 can then feed back the obtained classification detection prediction result corresponding to the target visible light image to terminal 102. Furthermore, in some embodiments, the saliency-guided high-resolution small target detection method can also be implemented independently by server 104 or terminal 102. For example, terminal 102 can directly perform image target detection on the target visible light image, or server 104 can obtain the target visible light image from the data storage system and perform image target detection on the target visible light image.

[0030] The terminal 102 can be, but is not limited to, various desktop computers, laptops, and tablets. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers, or it can be a cloud server.

[0031] In one exemplary embodiment, such as Figure 2 As shown, a saliency-guided high-resolution small target detection method is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps 201 to 202.

[0032] Step 201: Obtain a visible light image of the target.

[0033] Step 202: Input the target visible light image into the trained target detection model to obtain the classification detection prediction result corresponding to the target visible light image; the trained target detection model is an improved RetinaNet model, which includes an improved feature pyramid. The improved feature pyramid extracts multiple feature maps of different scales, denoted as P2 feature map, P3 feature map, P4 feature map, P5 feature map, P6 feature map, and P7 feature map, respectively; the resolution of the P2 feature map, P3 feature map, P4 feature map, P5 feature map, P6 feature map, and P7 feature map decreases sequentially.

[0034] The improved RetinaNet model connects a saliency-guided target detection module after the P2 and P3 feature maps, respectively, and a detection head after the P4, P5, P6, and P7 feature maps, respectively. The saliency-guided target detection module is used to perform target detection on the P2 or P3 feature maps to obtain the detection results corresponding to the P2 or P3 feature maps.

[0035] Implementing steps 201 to 202 above solves the problems of high computational cost and limited accuracy in detecting small targets using high-resolution feature maps. By accurately locating small targets, it fully utilizes the advantages of high-resolution feature maps while avoiding redundant computation in background regions, thereby improving the speed and accuracy of small target detection. This application designs a saliency-guided target detection method. First, a cascaded saliency detection module is designed, fusing multi-layer features and progressively optimizing the generation of saliency maps to achieve accurate localization of small targets. Subsequently, a dedicated sparse convolution-based detection head performs detection at the small target location to achieve accurate recognition. The sparse convolution mask is generated based on the saliency map, which trains the localization and detection modules together, allowing them to optimize each other. The saliency prior information generated by the localization module can highlight the features of small objects, enabling the detector to extract unique features from cluttered backgrounds, thereby improving detection accuracy. At the same time, the saliency-guided target detection module provides rich semantic information, which helps to more accurately identify salient regions. This method not only reduces computational costs but also improves the localization accuracy of small targets, enhancing the performance of small target detection.

[0036] This method adopts a localization-then-detection paradigm. First, a cascaded saliency detection module is used to fuse multi-layer features and gradually optimize the generation of the saliency map to achieve accurate localization of small targets. Then, a sparse convolutional mask is generated based on the saliency map, so that the detector operates only at the location where the small target exists. Finally, a dedicated sparse detection head uses sparse convolution to extract features of the small target and performs recognition under limited feature conditions to generate the target's bounding box and confidence score.

[0037] This application uses the RetinaNet model as the baseline network, and assumes the input target visible light image is I∈R. H×W×3 Where H and W are the width and height of the target's visible light image, the flowchart of the saliency-guided target detection method is shown below. Figure 3 Specifically, it includes:

[0038] Step 1: Input the image and extract multi-scale features.

[0039] like Figure 5As shown, the existing RetinaNet network consists of two parts: a backbone network with a feature pyramid (FPN), which extracts multi-scale feature maps, namely P2, P3, P4, P5, P6, and P7 feature maps. Each feature map is followed by a head (detection head), which includes two detection heads for classification and regression, respectively.

[0040] The improved RetinaNet model provided in this application includes a backbone network and a detection head network. First, a visible light image of the target is input into the backbone network to extract multi-scale features f. l ∈R H'×W'×C Here, 'l' represents the feature pyramid level. Because small targets have fewer pixels, lower-resolution high-level features tend to represent the semantic information of large and medium-sized objects, limiting the effectiveness of small target features. Conversely, lower-level feature maps have higher resolution and stronger small target information, making them more suitable for small target detection. Therefore, unlike RetinaNet, which only performs detection on feature maps from layers P3 to P7, this application adds a high-resolution feature map at layer P2 to improve small target detection performance. Considering the high computational cost of high-resolution feature maps (high-resolution feature f3 accounts for nearly half of the FLOPs when using RetinaNet) and the sparse distribution of small targets in the image, this application maintains a dense detection method on low-resolution features f4 to f7, while performing sparse detection on high-resolution features f2 and f3. This significantly reduces computational cost and improves the speed of small target detection.

[0041] Step 2: Establish a cascaded significance detection module.

[0042] The salience-guided target detection module ( Figure 4 The SGODM in the model includes a cascaded significance detection module. Figure 4 CSDM and sparse detection head (in the middle) Figure 4 The cascaded saliency detection module is used to perform saliency detection on the P2 feature map or the P3 feature map to obtain the saliency map corresponding to the P2 feature map or the P3 feature map; the sparse detection head includes a classification detection head and a localization detection head; the classification detection head is used to classify based on the target feature map and the saliency map corresponding to the target feature map to obtain the classification prediction result corresponding to the target feature map; the localization detection head is used to locate based on the target feature map and the saliency map corresponding to the target feature map to obtain the bounding box coordinate prediction result corresponding to the target feature map; the target feature map is the P2 feature map or the P3 feature map; the classification prediction result and the bounding box coordinate prediction result corresponding to the target feature map constitute the detection result corresponding to the target feature map.

[0043] For the P3 feature map, the cascaded saliency detection module includes a convolutional layer and a sigmoid activation function layer; the convolutional layer is used to extract features from the P3 feature map to obtain intermediate features; the sigmoid activation function layer is used to normalize the intermediate features to obtain the saliency map corresponding to the P3 feature map.

[0044] For the P2 feature map, the cascaded saliency detection module is used to: upsample the saliency map corresponding to the P3 feature map to obtain upsampled features; perform element-wise multiplication on the upsampled features and the P2 feature map to obtain multiplicative features; concatenate the multiplicative features and the P2 feature map to obtain concatenated features; and perform saliency detection on the concatenated features to obtain the saliency map corresponding to the P2 feature map.

[0045] Saliency detection models can simulate the human visual system's perception of salient regions in images, facilitating rapid target localization. This application directly performs saliency detection on high-resolution feature maps, utilizing their fine features to generate accurate saliency maps, thus improving the reliability of localization. In this application, a standard 3×3 convolutional layer is first applied to the P3 feature map to obtain intermediate features. The number of input channels matches the number of channels in the P3 feature map, and the output channel is set to 1. Then, the intermediate features... Sigmoid normalization is performed to generate the final single-channel saliency map, i.e., the saliency map S3 corresponding to the P3 feature map. This process can be described as follows:

[0046]

[0047] Here, w1 represents the convolution weights, and b1 is the bias term. This compact operation does not introduce redundant network architecture, effectively avoiding additional computational overhead.

[0048] Since the saliency map S3 corresponding to the P3 feature map contains the location information of small targets, this location information can be combined with the P2 feature layer to suppress background interference. Therefore, firstly, the saliency map S3 corresponding to the P3 feature map is upsampled to obtain the upsampled feature UP(S3). Then, a fusion method similar to residuals is used to merge the upsampled feature UP(S3) with the P2 feature map. Finally, saliency detection is performed on the fused feature map to obtain the saliency map S2 corresponding to the P2 feature map. This process can be described as follows:

[0049]

[0050] Where f2 is the P2 feature map; The first part represents the feature map after saliency enhancement, i.e., the concatenated feature; b2 and w2 represent the convolution weights and biases, respectively; the operator ⊙ represents the element-wise multiplication operation; UP represents the upsampling operation; f2⊙UP(S3) is the multiplicative feature; For the intermediate features extracted after concatenating features, After normalization, the saliency map S2 corresponding to the P2 feature map is obtained. This cascaded structure can progressively optimize the saliency map generation process, which helps to improve the accuracy and robustness of saliency detection and enhance the ability to locate small objects.

[0051] Step 3: Saliency-guided sparse mask generation.

[0052] When using sparse convolution for detection, improper mask settings may lead to unnecessary background calculations or omission of crucial foreground information, thereby reducing the speed and accuracy of small target detection. The saliency map generated by saliency detection can highlight the possible locations of small targets; a higher intensity value indicates a greater likelihood of target presence. Therefore, this application utilizes the saliency map as a sparse mask to guide convolution operations. To determine the optimal intensity threshold, the threshold is first set to 0 (i.e., features at all locations are processed), and then an intensity value that maintains detection accuracy while maximizing speed is found through testing experiments. After obtaining the optimal threshold, this application considers all regions with saliency values ​​higher than the optimal threshold σ (σ > 0) as potential target regions, thus generating a sparse convolution mask.

[0053]

[0054] Among them, H l S represents the sparse convolution mask of the l-th layer. l This is the saliency map of layer l; l = 2, 3. This guides the object detection module to more accurately locate and process small object features.

[0055] Step 4: Establish a sparse detection head.

[0056] This application designs a sparse detection head composed of sparse convolutions for high-resolution feature maps. It includes a feature enhancement module and a task-specific prediction layer, as detailed below. Figure 2 As shown, the feature enhancement module consists of four alternating 3×3 sparse convolutional layers (each with 256 channels) and employs the ReLU activation function to refine the feature representation. Subsequent task-specific layers use 3×3 sparse convolutions to generate class probabilities or bounding box coordinates.

[0057] Both the classification detection head and the localization detection head include a feature enhancement module and a prediction layer. The feature enhancement module contains four alternating sparse convolutional layers, each followed by a ReLU activation function layer. The classification detection head calculates the classification score for each category and uses the category with the highest classification score as the classification prediction result for the target feature map.

[0058] Sparse convolutional layers can effectively avoid unnecessary background computation by restricting computation to the foreground region through sparse masks (sparse convolutional masks). For example... Figure 6 As shown, this sparse mask contains a large number of zero-value pixels and a small number of active points. Element-wise multiplication allows for the selective extraction and processing of features from small objects, significantly improving detection speed. This operation can be formally described as:

[0059] f sp =w·f sp H i ⊙f+b(5);

[0060] Among them, H i Denotes the sparse convolution mask of the i-th layer; f and f sp represents the original features of the sparse convolution input and the output of the sparse convolution layer, respectively; w and b represent the weights and biases of the sparse convolution layer, respectively.

[0061] To achieve parameter efficiency, the RetinaNet model employs a shared detector head structure across all feature levels. However, this sharing strategy has limitations when processing feature maps of varying resolutions. For low-resolution feature maps, traditional dense convolutions capture global feature representations by computing all spatial locations. In contrast, sparse convolutions selectively process only non-zero activation regions on high-resolution feature maps, resulting in more discriminative sparse feature representations. Due to the differences in semantic information and spatial distribution between these two approaches, parameter sharing between the sparse head and the standard detector head degrades the performance of small object detection. Therefore, the sparse detector head independently optimizes its parameters during training to further improve small object detection performance. This architecture facilitates precise optimization of the sparse head's parameters during training, leading to more efficient sparse feature processing.

[0062] Step 5: Training and Testing.

[0063] This application first inputs the sample visible light image into the backbone network to extract multi-layer features. For low-layer features (including P4, P5, P6 and P7 feature maps), dense detection is performed directly. For high-layer features (including P2 and P3 feature maps), cascaded saliency detection is first used for localization to generate a sparse convolution mask. Then, a sparse detection head is used to classify and locate the location of small targets.

[0064] Before inputting the visible light image into the trained target detection model, the saliency-guided high-resolution small target detection method further includes: acquiring a sample set; the sample set includes several sample visible light images and the true classification label, true bounding box coordinates, and true saliency map corresponding to each sample visible light image; using a loss function, the target detection model is trained using the sample set to obtain a trained target detection model.

[0065] During training, both the cascaded saliency detection module and the sparse convolution-based object detection head (sparse detection head) operate on the same feature map, enabling them to optimize synergistically and further improve detection performance. This embodiment uses a hybrid loss function to optimize each feature layer:

[0066]

[0067] Among them, L l (C l B l ,S l () represents the loss function, which includes classification loss. Bounding box loss and saliency plot loss C l and B represents the predicted classification label and the true classification label, respectively. l and S represents the predicted bounding box coordinates and the ground truth bounding box coordinates, respectively. l and These represent the predicted saliency map and the true saliency map, respectively (the true saliency map can be directly generated using a binary mask, where pixels within the bounding box of a small target are labeled as 255, and other pixels as 0). The classification loss is calculated using the focus loss, the bounding box loss using the smoothing L1 loss, and the saliency map loss using the binary cross-entropy loss.

[0068] To ensure multi-scale feature learning, this embodiment introduces β. l This parameter is used to achieve layer-by-layer loss rebalancing. Therefore, the model's total loss function can be defined as:

[0069]

[0070] This application uses a stochastic gradient descent (SGD) optimizer to train all detectors for 50,000 iterations with an initial learning rate of 0.001, which is reduced by a factor of 10 at iterations 30,000 and 40,000, respectively. The rebalancing weights β... lIt is set to increase linearly from 1 to 2.6. In each iteration:

[0071] (1) Input the sample visible light image I and extract multi-scale features f l ;

[0072] (2) Obtain the saliency map S3 of the P3 feature layer through saliency detection, fuse S3 and the P2 feature map f2 to enhance the small target features of the P2 feature layer, and then obtain the saliency map S2 through saliency detection.

[0073] (3) Generating a sparse convolutional mask H based on the saliency map l ;

[0074] (4) Perform dense detection on low-resolution feature maps and perform targeted detection using a sparse detection head on high-resolution feature maps to obtain the target classification and bounding box location coordinates.

[0075] (5) Calculate the loss function L l =L cls +L bbox +L sal And then propagate in reverse;

[0076] After training, the best weight information is saved, and then the weights are used directly in the testing phase to detect the input target visible light image to obtain the classification detection prediction result of the target visible light image. The classification detection prediction result includes the classification prediction label and the bounding box position coordinates.

[0077] This application focuses on establishing a small target detection model on high-resolution feature maps. First, the saliency map is optimized layer by layer through a cascaded saliency detection module to accurately locate small targets. Then, a sparse convolutional mask is generated based on the saliency map for subsequent prediction. Finally, a sparse detection head is used to detect small targets at the locations where they exist in the high-resolution feature map.

[0078] In this embodiment, let the target visible light image be as follows: Figure 7 As shown in the saliency plot... Figure 8 As shown, the corresponding target visible light image and saliency map have the same size. The target visible light image has 3 channels, while the saliency map has 1 channel.

[0079] Using the above steps to guide high-resolution small object detection with saliency maps, this embodiment compares the object detection model with baseline networks RetinaNet and RetinaNet* (RetinaNet model with added high-resolution feature map P2 for detection). The classification and detection prediction results of existing object detection models and the object detection model of this application are as follows: Figure 9As shown, the leftmost column of images represents the classification and detection prediction results obtained using the RetinaNet model, the middle column represents the classification and detection prediction results obtained using the RetinaNet* model, and the rightmost column represents the classification and detection prediction results obtained using the object detection model of this application. To quantitatively evaluate the performance of the above methods, this application selects the mean accuracy mAP and the mean accuracy AP when the IoU is 0.5. 0.5 Average precision (AP) for small-scale targets (area less than 32×32 pixels) S The five metrics—average recall (AR), frames per second (FPS), and image processing frequency—are used to evaluate the accuracy of object detection. Higher values ​​for these metrics indicate more accurate object detection. The results are shown in Table 1. The results demonstrate that the saliency-guided high-resolution small object detection method proposed in this application outperforms the comparative methods in all metrics.

[0080] Table 1 Quantitative Results of Target Detection

[0081] RetinaNet 21.81 39.38 13.63 31.04 18.66 RetinaNet* 24.09 44.68 16.37 32.62 14.41 The target detection model of this application 25.91 47.90 19.00 34.52 21.14

[0082] This application designs a saliency-guided high-resolution small target detection method. First, a cascaded saliency detection module fuses multi-layer features to generate a saliency map for small target localization. This module operates directly on the high-resolution feature map and employs lightweight convolutions to replace the traditional heavy saliency detection head, improving localization accuracy while maintaining detection speed. Then, a sparse detection head uses sparse convolutions to operate only in small target regions, and its parameters are not shared with other detection heads during training. This allows the sparse detection head to be specifically optimized to better handle representative sparse small target features, further improving detection accuracy. Furthermore, we generate sparse convolutional masks based on the saliency map, thereby achieving joint training and mutual enhancement of the localization and detection modules in an end-to-end manner. The saliency map guides the detector to focus on potential target regions to improve small target detection accuracy, while the detection results optimize saliency map generation for more accurate small target localization. Compared to baseline networks, this application introduces high-resolution feature maps to enhance small target detection performance while avoiding unnecessary background region computation, thus improving both detection speed and accuracy.

[0083] This application also provides an application scenario in which the above-described saliency-guided high-resolution small target detection method is applied. Specifically, the saliency-guided high-resolution small target detection method provided in this embodiment can be applied to an image classification and detection scenario. The image classification and detection scenario includes an image acquisition stage and an image classification and detection stage; the target visible light image enters the image classification and detection stage from the content production stage, and the corresponding classification and detection prediction results are obtained through human-machine collaboration. The saliency-guided high-resolution small target detection method provided in this embodiment belongs to the machine labeling stage in the image classification and detection stage. Specifically, in the image classification and detection stage for the target visible light image, the target visible light image can be labeled based on a collaborative method of machine labeling and manual labeling, that is, corresponding classification prediction labels and bounding box position coordinates are added to the target visible light image.

[0084] Example 2

[0085] Based on the same inventive concept, this application also provides a salience-guided high-resolution small target detection system for implementing the salience-guided high-resolution small target detection method described above. The solution provided by this system is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the salience-guided high-resolution small target detection system provided below can be found in the limitations of the salience-guided high-resolution small target detection method described above, and will not be repeated here.

[0086] This embodiment provides a saliency-guided high-resolution small target detection system, which includes the following modules.

[0087] The target visible light image acquisition module is used to acquire the target visible light image.

[0088] The target detection module is used to input the visible light image of the target into a trained target detection model to obtain the classification detection prediction result corresponding to the visible light image of the target. The trained target detection model is an improved RetinaNet model, which includes an improved feature pyramid. The improved feature pyramid extracts multiple feature maps of different scales, denoted as P2 feature map, P3 feature map, P4 feature map, P5 feature map, P6 feature map, and P7 feature map, respectively. The resolution of the P2 feature map, P3 feature map, P4 feature map, P5 feature map, P6 feature map, and P7 feature map decreases sequentially.

[0089] The improved RetinaNet model connects a saliency-guided target detection module after the P2 and P3 feature maps, respectively, and a detection head after the P4, P5, P6, and P7 feature maps, respectively. The saliency-guided target detection module is used to perform target detection on the P2 or P3 feature maps to obtain the detection results corresponding to the P2 or P3 feature maps.

[0090] As an optional implementation, the saliency-guided target detection module includes a cascaded saliency detection module and a sparse detection head;

[0091] The cascaded saliency detection module is used to perform saliency detection on the P2 feature map or the P3 feature map to obtain the saliency map corresponding to the P2 feature map or the P3 feature map.

[0092] The sparse detection head includes a classification detection head and a localization detection head. The classification detection head is used to classify based on the target feature map and the corresponding saliency map to obtain the classification prediction result corresponding to the target feature map. The localization detection head is used to locate based on the target feature map and the corresponding saliency map to obtain the bounding box coordinate prediction result corresponding to the target feature map. The target feature map is a P2 feature map or a P3 feature map. The classification prediction result and the bounding box coordinate prediction result corresponding to the target feature map constitute the detection result corresponding to the target feature map.

[0093] Specifically, for the P3 feature map, the cascaded saliency detection module includes a convolutional layer and a sigmoid activation function layer; the convolutional layer is used to extract features from the P3 feature map to obtain intermediate features; the sigmoid activation function layer is used to normalize the intermediate features to obtain the saliency map corresponding to the P3 feature map.

[0094] For the P2 feature map, the cascaded saliency detection module is used to: upsample the saliency map corresponding to the P3 feature map to obtain upsampled features; perform element-wise multiplication on the upsampled features and the P2 feature map to obtain multiplicative features; concatenate the multiplicative features and the P2 feature map to obtain concatenated features; and perform saliency detection on the concatenated features to obtain the saliency map corresponding to the P2 feature map.

[0095] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0096] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A saliency-guided high-resolution small target detection method, characterized in that, The saliency-guided high-resolution small target detection method includes: Acquire a visible light image of the target; The target visible light image is input into a trained target detection model to obtain the classification and detection prediction result corresponding to the target visible light image. The trained target detection model is an improved RetinaNet model, which includes an improved feature pyramid. The improved feature pyramid extracts multiple feature maps of different scales, denoted as P2 feature map, P3 feature map, P4 feature map, P5 feature map, P6 feature map, and P7 feature map, respectively. The resolution of the P2 feature map, P3 feature map, P4 feature map, P5 feature map, P6 feature map, and P7 feature map decreases sequentially. The improved RetinaNet model connects a saliency-guided target detection module after the P2 and P3 feature maps, respectively, and a detection head after the P4, P5, P6, and P7 feature maps, respectively. The saliency-guided target detection module is used to perform target detection on the P2 or P3 feature maps to obtain the detection results corresponding to the P2 or P3 feature maps.

2. The saliency-guided high-resolution small target detection method according to claim 1, characterized in that, The saliency-guided target detection module includes a cascaded saliency detection module and a sparse detection head; The cascaded saliency detection module is used to perform saliency detection on the P2 feature map or the P3 feature map to obtain the saliency map corresponding to the P2 feature map or the P3 feature map. The sparse detection head includes a classification detection head and a localization detection head; the classification detection head is used to classify based on the target feature map and the saliency map corresponding to the target feature map to obtain the classification prediction result corresponding to the target feature map; the localization detection head is used to locate based on the target feature map and the saliency map corresponding to the target feature map to obtain the bounding box coordinate prediction result corresponding to the target feature map. The target feature map is either a P2 feature map or a P3 feature map; the classification prediction result and the bounding box coordinate prediction result corresponding to the target feature map constitute the detection result corresponding to the target feature map.

3. The saliency-guided high-resolution small target detection method according to claim 2, characterized in that, For the P3 feature map, the cascaded saliency detection module includes a convolutional layer and a sigmoid activation function layer; the convolutional layer is used to extract features from the P3 feature map to obtain intermediate features; The Sigmoid activation function layer is used to normalize the intermediate features to obtain the saliency map corresponding to the P3 feature map; For the P2 feature map, the cascaded saliency detection module is used to: upsample the saliency map corresponding to the P3 feature map to obtain upsampled features; perform element-wise multiplication on the upsampled features and the P2 feature map to obtain multiplicative features; and concatenate the multiplicative features and the P2 feature map to obtain concatenated features. The saliency of the spliced ​​features is detected to obtain the saliency map corresponding to the P2 feature map.

4. The saliency-guided high-resolution small target detection method according to claim 2, characterized in that, Both the classification detection head and the localization detection head include a feature enhancement module and a prediction layer. The feature enhancement module contains four alternately arranged sparse convolutional layers, each followed by a ReLU activation function layer.

5. The saliency-guided high-resolution small target detection method according to claim 1, characterized in that, Before inputting the visible light image into the trained target detection model, the saliency-guided high-resolution small target detection method further includes: Obtain a sample set; the sample set includes several sample visible light images and the true classification label, true bounding box coordinates and true saliency map corresponding to each sample visible light image; Using the loss function and the sample set, the target detection model is trained to obtain a trained target detection model.

6. The saliency-guided high-resolution small target detection method according to claim 5, characterized in that, Loss functions include classification loss, bounding box loss, and saliency map loss.

7. The saliency-guided high-resolution small target detection method according to claim 6, characterized in that, The classification loss is calculated using focus loss, the bounding box loss is calculated using smoothed L1 loss, and the saliency map loss is calculated using binary cross-entropy loss.

8. A saliency-guided high-resolution small target detection system, characterized in that, The saliency-guided high-resolution small target detection system includes: The target visible light image acquisition module is used to acquire the target visible light image; The target detection module is used to input the visible light image of the target into a trained target detection model to obtain the classification and detection prediction result corresponding to the visible light image of the target. The trained target detection model is an improved RetinaNet model, which includes an improved feature pyramid. The improved feature pyramid extracts multiple feature maps of different scales, denoted as P2 feature map, P3 feature map, P4 feature map, P5 feature map, P6 feature map, and P7 feature map, respectively. The resolution of the P2 feature map, P3 feature map, P4 feature map, P5 feature map, P6 feature map, and P7 feature map decreases sequentially. The improved RetinaNet model connects a saliency-guided target detection module after the P2 and P3 feature maps, respectively, and a detection head after the P4, P5, P6, and P7 feature maps, respectively. The saliency-guided target detection module is used to perform target detection on the P2 or P3 feature maps to obtain the detection results corresponding to the P2 or P3 feature maps.

9. The saliency-guided high-resolution small target detection system according to claim 8, characterized in that, The saliency-guided target detection module includes a cascaded saliency detection module and a sparse detection head; The cascaded saliency detection module is used to perform saliency detection on the P2 feature map or the P3 feature map to obtain the saliency map corresponding to the P2 feature map or the P3 feature map. The sparse detection head includes a classification detection head and a localization detection head; the classification detection head is used to classify based on the target feature map and the saliency map corresponding to the target feature map to obtain the classification prediction result corresponding to the target feature map; the localization detection head is used to locate based on the target feature map and the saliency map corresponding to the target feature map to obtain the bounding box coordinate prediction result corresponding to the target feature map. The target feature map is either a P2 feature map or a P3 feature map; the classification prediction result and the bounding box coordinate prediction result corresponding to the target feature map constitute the detection result corresponding to the target feature map.

10. The saliency-guided high-resolution small target detection system according to claim 9, characterized in that, For the P3 feature map, the cascaded saliency detection module includes a convolutional layer and a sigmoid activation function layer; the convolutional layer is used to extract features from the P3 feature map to obtain intermediate features; The Sigmoid activation function layer is used to normalize the intermediate features to obtain the saliency map corresponding to the P3 feature map; For the P2 feature map, the cascaded saliency detection module is used to: upsample the saliency map corresponding to the P3 feature map to obtain upsampled features; perform element-wise multiplication on the upsampled features and the P2 feature map to obtain multiplicative features; concatenate the multiplicative features and the P2 feature map to obtain concatenated features; and perform saliency detection on the concatenated features to obtain the saliency map corresponding to the P2 feature map. Both the classification detection head and the localization detection head include a feature enhancement module and a prediction layer. The feature enhancement module contains four alternately arranged sparse convolutional layers, each followed by a ReLU activation function layer.