Multi-scale feature fusion detection method and system for salient target of remote sensing image
By constructing a multi-scale uncertainty confusing feature digestion encoder-decoder deep supervision framework and probability distribution mapping module, the uncertainty problem of significant object detection in remote sensing images is solved, and the detection accuracy and adaptability are improved, especially in complex remote sensing scenarios, which can more accurately identify and locate significant objects.
Patent Information
- Application Number
- CN202510108974.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-07-18
AI Technical Summary
The existing remote sensing significant object detection methods have high uncertainty when facing remote sensing scenes with multi-scale, complex backgrounds and blurred edges, resulting in a degradation of detection performance. It is difficult to effectively evaluate and eliminate uncertainty under the influence of environmental factors such as lighting and shadows in remote sensing images.
A deep supervision framework for multi-scale uncertainty obfuscated feature digestion is constructed, a multi-scale feature map is generated by using the encoder, an uncertainty mapping module based on probability distribution is designed for nonlinear weighting, and pixel uncertainty is evaluated in combination with Gaussian distribution, and uncertainty is suppressed through the multi-scale focus fusion module, and layer-by-layer supervision is used to utilize binary cross entropy and IOU losses.
It improves the discrimination ability and detection accuracy of remote sensing significant object detection, can more accurately identify significant objects in complex scenarios, reduce sample prediction uncertainty, stronger adaptability and stability, and can accurately capture key details.
Smart Images

Figure CN120339666A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of remote sensing image processing methods, and particularly relates to a multi-scale feature fusion detection method and system for prominent targets in remote sensing images. Background Art
[0002] Prominent target detection is an important task in the field of computer vision. The aim is to automatically locate and segment the targets or regions in an image that can most attract human visual attention. Prominent target detection in remote sensing images has important application value in assisting in key tasks such as agricultural resource monitoring, ecological environment protection, urban planning, and disaster assessment, providing key support for accurate decision-making.
[0003] In the field of prominent target detection, traditional methods mainly rely on manually crafted feature extraction and machine learning-based classification means. However, when facing remote sensing scenes with multi-scale, complex backgrounds, and blurred edges, such methods usually perform unsatisfactorily. With the rapid development of deep learning technology, significant progress has been made in remote sensing prominent target detection. Extracting local features of images through multi-layer convolution to achieve prominent target recognition and localization has become the mainstream. When the Transformer technology was introduced, methods based on Transformer began to gradually attract the attention of many researchers. These methods greatly enhance the long-range dependence modeling ability of the model by constructing global dependence relationships and context information. Remote sensing prominent target detection methods based on deep learning usually generate a saliency map through non-linear mapping and use the Softmax or Sigmoid function to calculate the probability that each pixel belongs to the foreground or background. Although the probability value usually ranges between 0 and 1, where 0 represents the background and 1 represents the foreground, once the probability of a pixel approaches 0.5, the credibility of determining whether the pixel belongs to the foreground or background becomes extremely low, which will lead to greater uncertainty and thus affect the overall detection performance.
[0004] In summary, whether it is possible to effectively evaluate and eliminate the uncertainty of sample prediction has become an important means to improve the performance of prominent target detection. In terms of uncertainty evaluation, most researchers focus on the background of natural images and do not fully consider the unique challenges of remote sensing images. Remote sensing images usually have complex and heterogeneous backgrounds, large target scale variations, and are affected by environmental factors such as light and shadow, which will introduce additional uncertainties. Therefore, it is impossible to explore the essential features of prominent target detection in complex remote sensing scenes. Summary of the Invention
[0005] In view of the above deficiencies or improvement requirements of the prior art, the present invention provides a multi-scale feature fusion detection method and system for prominent targets in remote sensing images. First, an encoder-decoder deep supervision framework for resolving multi-scale uncertainty confusion features is constructed. An encoder is used to generate four feature maps of different scales, and the changing edges in these feature maps have relatively high uncertainty. Secondly, an uncertainty mapping module based on probability distribution is designed to non-linearly weight the pixels in the fuzzy area by combining Gaussian distribution to evaluate the uncertainty of the target pixels. Thirdly, the evaluated uncertainty feature maps are passed through a multi-scale focusing fusion module to fuse the global and local multi-scale features, suppress the uncertainty of the confusion features, and further improve the discriminative ability and detection accuracy of the model.
[0006] To achieve the above object, according to the first aspect of the present invention, there is provided a multi-scale feature fusion detection method for prominent targets in remote sensing images, including:
[0007] S100: Initialize the calculation parameters of the remote sensing prominent target detection model based on probability uncertainty evaluation, and use a feature extractor to extract features from the remote sensing image to generate feature maps with multi-scales and relatively high uncertainty at the changing edges;
[0008] S200: Use the designed uncertainty evaluation mapping module to take the high-level semantic information of the remote sensing image feature map obtained in step S100 as prior knowledge, and perform non-linear weighting processing on the uncertain pixels;
[0009] S300: Use the uncertainty evaluation mapping module in step S200 to reconstruct the probability distribution maps of the foreground and background respectively, and then use this uncertainty information to weight and fuse the high-level and low-level features to obtain clearer foreground and background feature representations;
[0010] S400: Use the designed multi-scale focusing fusion module to focus and enhance the deep features obtained in step S300, retain the details of the shallow features, and save them to achieve high highlighting of important regions and information supplementation of texture details;
[0011] S500: Perform global and local multi-scale joint processing on the deep features and shallow features processed in step S400 to suppress the uncertainty of the confusion features and further improve the discriminative ability and detection accuracy of the model;
[0012] S600: Use a combination of binary cross-entropy loss and IOU loss to perform layer-by-layer supervision on the prediction results of multiple stages output by the model to ensure effective cooperation between features at different levels.
[0013] Further, step S200 includes:
[0014] S201: Using the high-level semantic features as prior knowledge, convert the prediction output into a probability form as follows:
[0015] prob_map = σ(P i ), (2)
[0016] where P i is the prior feature, and σ(·) represents converting P i into a probability map through Sigmoid;
[0017] S202: According to the generated probability map, perform non-linear weighting using the Gaussian distribution, specifically:
[0018]
[0019] where prob_map is the probability-form prediction output in Equation (2), U fore and U back represent the foreground and background probability weights respectively, μ and σ represent the mean and standard deviation of the Gaussian distribution respectively, exp(·) is the exponential function with the natural constant e as the base, represents the foreground uncertainty of the high-level feature map, represents the background uncertainty of the high-level feature map, (·) i represents the number of the stage, i ∈ {1, 2, 3, 4}, and the foreground and background uncertainty maps of the low-level feature map and are also generated in the same way.
[0020] Furthermore, in step S300, multiply the foreground and background uncertainty maps after Gaussian weighting with the foreground and background of the high- and low-level feature maps for weighted fusion, so that the model can identify the foreground and background regions, enhance the foreground significant features and suppress the background interference, and achieve more accurate classification decisions in regions with high uncertainty, specifically:
[0021]
[0022] In Equation (4), represents the high-level feature initially extracted at the i-th stage, represents the low-level feature initially extracted at the (i - 1)-th stage, Concat(·) represents the fusion operation of the two feature maps, CBR(·) represents the Conv convolution, BatchNorm normalization, and ReLu non-linear activation operations on the feature map, and Feature-H and Feature-L represent the high-level and low-level features after weighted fusion respectively.
[0023] Furthermore, the local multi-scale joint processing described in step S400 includes:
[0024] Adjust the channel dimension through a point convolution layer, fuse the high-level feature Feature-H and perform batch normalization, then implement non-linear transformation through ReLu, and finally further optimize the feature representation through another point convolution layer and batch normalization. The calculation formula is as shown in (5):
[0025]
[0026] In formula (5), F top represents the feature map after the high-level feature map Feature-H first passes through the point convolution layer, batch normalization, and ReLu. PWConv represents the point convolution layer, BN represents the batch normalization operation, and F local represents the feature map F top The local branch feature map after passing through another point convolution layer and batch normalization.
[0027] Furthermore, the global multi-scale joint processing described in step S400 includes:[[]]
[0028] Process the high-level feature Feature-H through a global average pooling layer to capture the global context, and then perform point convolution processing. The calculation formula is as shown in (6):
[0029]
[0030] In formula (6), F bottom represents the feature map obtained after the high-level feature map Feature-H passes through the global average pooling process. GAP represents the global average pooling, and F global represents F bottom The global branch feature map obtained after passing through the point convolution processing.
[0031] Furthermore, step S500 includes:[[]]
[0032] Add the outputs of the global branch and the local branch, and obtain the probability distribution of the important features through Sigmoid processing. Then, perform focus enhancement on the high-level features. The calculation formulas are as shown in (7) and (8):
[0033] F prob = σ(F local + F global ), (7)
[0034] F focus = Feature-H + F prob · Feature-H, (8)
[0035] In formula (7), F probIt means adding the global branch feature map and the local branch feature map, and converting the obtained feature map into a probability distribution map through Sigmoid. σ(·) represents the Sigmoid processing of the feature map.
[0036] Furthermore, step S500 includes:
[0037] Performing feature fusion on the result after focus enhancement and the low-level feature map to achieve the supplementation and enhancement of texture details. The calculation formula is shown in Equation (9):
[0038] F output = Conv 3×3 (Concat(Upsample(F focus ), Feature-L)), (9)
[0039] In Equation (9), F output represents performing feature fusion on F focus and the low-level feature map Feature-L to supplement texture detail information. Upsample represents upsampling the feature map, and Conv 3×3 represents performing a 3×3 convolution on the feature map.
[0040] Furthermore, step S600 includes: Adopting a deep supervision mechanism to perform layer-by-layer supervision on predictions at different levels, enabling this model to perform salient object detection in feature spaces of different scales and ensuring the effective integration of shallow features and deep features. The calculation formula is shown in (10).
[0041]
[0042] In Equation (10), GT represents the true saliency label map, and λ1 and λ2 are the weight coefficients of the BCE loss and the IoU loss, controlling their relative contributions, with the default value being 1.
[0043] Furthermore, step S100 includes:
[0044] S101: Performing preliminary processing on the input remote sensing image, initializing a series of basic configuration parameters, and loading its pre-trained weights;
[0045] S102: Constructing an encoder-decoder deep supervision framework for multi-scale uncertainty confusion feature resolution. For the input remote sensing image I ∈ R C×H×W , using the selected feature extractor PVTv2 to perform multi-stage feature extraction on the image. These feature maps will serve as the basic data for subsequent operations such as uncertainty evaluation and feature fusion. The calculation formula is shown in (1):
[0046] F n = Block n(I), n ∈ {1, 2, 3, 4}, (1)
[0047] In formula (1), Block n (·) represents a feature extractor; F n represents the feature maps extracted at different stages.
[0048] According to the second aspect of the present invention, there is provided a remote sensing image salient object detection system based on probability uncertainty assessment, including:
[0049] A feature extraction module, configured to initialize the calculation parameters of a remote sensing salient object detection model based on probability uncertainty assessment, use the feature extractor to extract features from the remote sensing image, and generate feature maps with multiple scales and relatively high uncertainty at the changing edges;
[0050] A weighted processing module, configured to use the designed uncertainty assessment mapping module to use the high-level semantic information of the remote sensing image feature map obtained in step S100 as prior knowledge to perform non-linear weighted processing on the pixels with uncertainty;
[0051] An assessment mapping module, configured to separately reconstruct the probability distribution maps of the foreground and the background, and then use this uncertainty information to weightedly fuse high-level and low-level features to obtain clearer foreground and background feature representations;
[0052] A multi-scale focusing fusion module, configured to enhance the focus of the obtained deep features, retain the details of the shallow features, and save them to achieve a high degree of highlighting of important regions and information supplementation of texture details;
[0053] A multi-scale joint processing module, configured to perform global and local multi-scale joint processing on the deep features and the shallow features, suppress the uncertainty of the confusing features, and further improve the discrimination ability and detection accuracy of the model;
[0054] A layer-by-layer supervision module, configured to use a combination of binary cross-entropy loss and Intersection over Union loss to perform layer-by-layer supervision on the prediction results of multiple stages output by the model, and ensure the effective cooperation between features at different levels.
[0055] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0056] 1. The method of the present invention is divided into three parts: remote sensing image feature extraction, uncertainty assessment mapping, and multi-scale focused fusion. First, an encoder-decoder deep supervision framework for multi-scale uncertainty confusion feature resolution is constructed. The encoder is used to generate four feature maps of different scales, and the changing edges in these feature maps have relatively high uncertainty. Secondly, an uncertainty mapping module based on probability distribution is designed, which combines the Gaussian distribution to perform non-linear weighted processing on the pixels in the blurred area to evaluate the uncertainty of the target pixels. Thirdly, the evaluated uncertainty feature map is passed through the multi-scale focused fusion module to fuse the global and local multi-scale features, reduce the uncertainty of the confusion features, and further improve the discrimination ability and detection accuracy of the model.
[0057] 2. The method of the present invention can accurately identify and locate significant targets in remote sensing images, has stronger adaptability and stability in complex scenarios compared with existing detection methods, and can effectively reduce the uncertainty of sample prediction.
[0058] 3. The method of the present invention uses the trained remote sensing significant target detection model with probability uncertainty assessment to process a series of representative complex scenarios. Compared with other existing methods, the present invention can capture and display the key details in the image more accurately, and solves the problems in the prior art such as detection being interfered by shadows, the background of significant targets being complex, the scene containing multiple small targets, and the significant targets being narrow and slender. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a schematic flow chart of the multi-scale feature fusion detection method for significant targets in remote sensing images according to the embodiment of the present invention;
[0060] Figure 2 It is a schematic diagram of the uncertainty assessment mapping module in the embodiment of the present invention;
[0061] Figure 3 It is a schematic diagram of the multi-scale focused fusion module in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0063] Please refer to Figure 1 The flow chart. A multi-scale feature fusion detection method for significant targets in remote sensing images provided by the present invention includes the following steps:
[0064] Step 1: Initialize the calculation parameters of the remote sensing image processing model based on probability uncertainty assessment. Use the feature extractor to extract features from the remote sensing image, generating a feature map with multiple scales and relatively high uncertainty at the changing edges. Specifically, in Step 1, the input remote sensing image is preliminarily processed, a series of basic configuration parameters are initialized, the optimizer type, learning rate, and batch size are set, its pre-trained weights are loaded, and the size of the input image is uniformly adjusted to a unified size. Secondly, the present invention constructs an encoder-decoder deep supervision framework for multi-scale uncertainty confusion feature resolution. For the input remote sensing image I ∈ R C×H×W , use the selected feature extractor PVTv2 to perform multi-stage feature extraction on the image. These feature maps will serve as the basic data for subsequent operations such as uncertainty assessment and feature fusion. The calculation formula is shown in (1):
[0065] F n =Block n (I), n ∈ {1, 2, 3, 4}, (1)
[0066] In formula (1), Block n (·) represents the feature extractor; F n represents the feature maps extracted at different stages.
[0067] Step 2: The present invention designs an uncertainty assessment mapping module, which is specifically as Figure 2 shown. Through this module, the high-level semantic information of the remote sensing image feature map obtained in Step 1 is used as prior knowledge to perform non-linear weighting on the pixels with uncertainty.
[0068] First, taking the high-level semantic features as prior knowledge, convert the prediction output into a probability form. Its calculation expression is shown in formula (2):
[0069] prob_map=σ(P i ), (2)
[0070] In formula (2), P i is the prior feature, and σ(·) represents converting P i into a probability map through Sigmoid.
[0071] Secondly, based on the generated probability map, calculate the uncertainties of the foreground and background. To address the uncertainty problem when the probability is close to the threshold and enhance the model's attention to fuzzy regions, the present invention uses a Gaussian distribution for non-linear weighting. The reason for choosing the Gaussian distribution is that it has a symmetric bell-shaped curve property, which can enhance the weight when the probability is close to the mean. Through this weighting method, we can make the model pay more attention to those fuzzy regions that are difficult to classify due to the probability being close to the threshold, thereby improving the feature expression and detection accuracy of these regions. Its expression is shown in Equation (3):
[0072]
[0073] In Equation (3), μ and σ represent the mean and standard deviation of the Gaussian distribution respectively. represents the foreground uncertainty of the high-level feature map, represents the background uncertainty of the high-level feature map. The foreground and background uncertainty maps of the low-level feature map ( and ) are also generated in the same way.
[0074] Step 3: Use the uncertainty evaluation mapping module in Step 2 to reconstruct the probability distribution maps of the foreground and background respectively, and then use this uncertainty information to weighted fuse the high- and low-level features to obtain clearer foreground and background feature representations. The specific process is as follows: Use the foreground and background uncertainty maps after Gaussian weighting to weighted fuse the foreground and background of the high- and low-level feature maps respectively. By multiplying these weighted uncertainty maps with the original feature maps, the model can identify which regions in the feature map belong to the foreground and which belong to the background, so that the model can effectively enhance the significant features of the foreground region and suppress the interference of the background region. Specifically, Gaussian weighting makes the regions that are difficult to distinguish, especially the parts close to the foreground and background boundaries, receive higher attention. Due to the characteristics of the Gaussian distribution, the regions close to the mean will be given higher weights, which helps the model to more precisely focus on the transition region between the foreground and the background, so as to make more accurate classification decisions in these regions with higher uncertainties. Its calculation formula is shown in Equation (4):
[0075]
[0076] In Equation (4), Concat(·) represents the operation of fusing two feature maps, and CBR(·) represents the operations of Conv convolution, BatchNorm normalization, and ReLu non-linear activation on the feature map. Feature-H and Feature-L represent the high-level and low-level features after weighted fusion respectively.
[0077] Step 4: Use the designed multi-scale focusing fusion module to enhance the deep features obtained in Step 3, retain the details of the shallow features, and achieve a high degree of highlighting of important regions and information supplementation of texture details. Ensure that the model can more accurately extract and understand key deep semantic information during the feature fusion process, while avoiding misleading caused by the lack of semantic understanding of shallow features during the fusion process, such as Figure 3 as shown
[0078] On the local branch, first adjust the channel dimension through a point convolution layer, fuse the high-level feature Feature-H and perform batch normalization, then implement a non-linear transformation through ReLu, and finally further optimize the feature representation through another point convolution layer and batch normalization. The shallow features use a small receptive field to focus on the enhancement of detailed features, and accurately identify the edges and internal detailed information of the target area. The calculation formula is as shown in (5):
[0079]
[0080] In formula (5), F top represents the feature map of the high-level feature map Feature-H passing through the point convolution layer, batch normalization, and ReLu for the first time. PWConv represents the point convolution layer, BN represents the batch normalization operation, and F local represents the local branch feature map of the feature map F top after passing through another point convolution layer and batch normalization.
[0081] On the global branch, process the high-level feature Feature-H through a global average pooling layer to capture the global context, and then perform point convolution processing. The deep features have a larger receptive field to capture important semantic information in the image and enhance the model's understanding of the context information. The calculation formula is as shown in (6).
[0082]
[0083] In formula (6), F bottom represents the feature map obtained by processing the high-level feature map Feature-H through global average pooling. GAP represents global average pooling, and F global represents the global branch feature map obtained by performing point convolution processing on F bottom .
[0084] Step 5: Adopt a global and local multi-scale denoising fusion processing strategy for the deep features and shallow features obtained in Step 4. In this way, the uncertainty caused by confusing features can be effectively suppressed, thereby further enhancing the performance of the model in terms of discriminative ability and detection accuracy, enabling it to more accurately complete object detection and related tasks, and improving the overall discriminative ability and detection accuracy of the model.
[0085] Add the outputs of the global branch and the local branch, and obtain the probability distribution of important features through Sigmoid processing. Then, focus on enhancing the high-level features to further highlight and optimize the key feature information, and improve the accuracy and effectiveness of the overall feature expression. The calculation formulas are shown in (7) and (8):
[0086] F prob =σ(F local +F global ), (7)
[0087] F focus =Feature-H+F prob ·Feature-H, (8)
[0088] In formula (7), F prob represents the addition of the global branch feature map and the local branch feature map, and the obtained feature map is converted into a probability distribution map through Sigmoid. σ(·) represents the Sigmoid processing of the feature map.
[0089] In formula (8), F focus represents the feature map obtained by focusing on enhancing the high-level feature map Feature-H.
[0090] Secondly, fuse the results after focusing enhancement with the low-level feature map to achieve the supplement and enhancement of texture details, making the overall features more abundant and accurate. The calculation formula is shown in formula (9).
[0091] F output =Conv 3×3 (Concat(Upsample(F focus ),Feature-L)), (9)
[0092] In formula (9), F output represents the feature fusion of F focus and the low-level feature map Feature-L to supplement texture detail information. Upsample represents the upsampling of the feature map, and Conv 3×3 represents the 3×3 convolution of the feature map.
[0093] Step 6: Use the combination of binary cross-entropy loss and Intersection over Union (IOU) loss to perform layer-by-layer supervision on the prediction results of multiple stages of the model output, ensuring the effective cooperation between different-level features.
[0094] A deep supervision mechanism is adopted to perform layer-by-layer supervision on predictions at different levels, enabling the model to detect significant targets in feature spaces of different scales, ensuring the effective integration of shallow features and deep features, thereby improving the overall detection performance and ensuring that the model can also function stably in various complex scenarios of remote sensing images. The calculation formula is shown in (10).
[0095]
[0096] In formula (10), GT represents the true saliency label map, and λ1 and λ2 are the weight coefficients of the BCE loss and the IoU loss, controlling their relative contributions, with the default value being 1.
[0097] Step 7: Use the trained remote sensing saliency target detection model with probability uncertainty evaluation to process a series of representative complex scenarios, including detecting problems such as being interfered by shadows, having a complex background of significant targets, containing multiple small targets in the scene, and the significant targets being narrow and slender. Compared with other existing methods, the present invention can capture and display the key details in the image more accurately. The comparison result data with other detection methods is shown in Table 1.
[0098] Table 1 Quantitative comparison with existing mainstream saliency target detection methods on the EORSSD dataset
[0099]
[0100]
[0101] Table 2 Quantitative comparison with existing mainstream saliency target detection methods on the ORSSD dataset
[0102]
[0103]
[0104] As can be seen from Tables 1 and 2, the present invention has achieved optimal or near-optimal performance on multiple evaluation metrics of salient object detection on the EORSSD and ORSSD datasets. The quantitative experimental results on the two datasets show that the overall detection performance of the present invention in the salient object detection task is better than other comparison methods. The uncertainty evaluation mechanism proposed by the present invention significantly improves the classification accuracy of the model for fuzzy pixels. Secondly, the probability uncertainty mapping module performs non-linear weighting on the uncertainty pixels close to the threshold, adjusts the weights of the uncertain regions, enabling the model to obtain more accurate discrimination ability when processing fuzzy regions and reducing the misjudgment risk brought by the fixed threshold. The multi-scale focused fusion module enhances the focus on deep semantic features during the fusion process through global and local joint reinforcement, while ensuring that the detailed information of shallow features is retained, making the distinction between foreground and background of the model clearer. The present invention shows significant robustness and accuracy especially in the remote sensing image scenario with multi-scale and multi-background interference.
[0105] In another embodiment of the present invention, there is provided a multi-scale feature fusion detection system for salient objects in remote sensing images, including
[0106] A feature extraction module, which is used to initialize the calculation parameters of the remote sensing salient object detection model based on probability uncertainty evaluation, extract features from the remote sensing image using a feature extractor, and generate a feature map with multi-scale and relatively high uncertainty of changing edges;
[0107] A weighting processing module, which is used to use the designed uncertainty evaluation mapping module to take the high-level semantic information of the remote sensing image feature map obtained in step S100 as prior knowledge and perform non-linear weighting processing on the uncertain pixels;
[0108] An evaluation mapping module, which is used to respectively reconstruct the probability distribution maps of the foreground and background, and then use this uncertainty information to weight and fuse high-level and low-level features to obtain a clearer foreground and background feature representation;
[0109] A multi-scale focused fusion module, which is used to enhance the focus of the obtained deep features, retain the details of the shallow features, and save them to achieve a highly prominent important area and information supplement of texture details;
[0110] A multi-scale joint processing module, which is used to perform global and local multi-scale joint processing on the deep features and shallow features, suppress the uncertainty of the confusing features, and further improve the discrimination ability and detection accuracy of the model;
[0111] A layer-by-layer supervision module, which is used to use the combination of binary cross-entropy loss and Intersection over Union loss to perform layer-by-layer supervision on the prediction results of multiple stages output by the model, and ensure the effective cooperation between features at different levels.
[0112] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. A multi-scale feature fusion detection method for prominent targets in remote sensing images, characterized in that, Including: S100: Initialize the calculation parameters of the remote sensing salient object detection model, extract features from the remote sensing image to generate a feature map with multiple scales and relatively high uncertainty in the changing edges; S200: Use the high-level semantic information of the remote sensing image feature map obtained in step S100 as prior knowledge to perform non-linear weighting on the pixels with uncertainty; S300: Use the probability distribution maps of the foreground and background reconstructed in step S200 respectively, and weighted-fuse the high-level and low-level features based on the uncertainty information to obtain clearer foreground and background feature representations; S400: Enhance the focus of the deep features obtained in step S300, retain the details of the shallow features, and save the information supplement for highly prominent important regions and texture details; S500: Perform global and local multi-scale joint processing on the deep features and shallow features processed in step S400, suppress the uncertainty of the confusing features, and further improve the discriminative ability and detection accuracy of the model; S600: Use the combination of binary cross-entropy loss and IOU loss to perform layer-by-layer supervision on the prediction results of multiple stages output by the model, and ensure the effective cooperation between features at different levels.
2. The multi-scale feature fusion detection method for prominent targets in remote sensing images according to claim 1, characterized in that Step S200 includes: S201: Use the high-level semantic features as prior knowledge, and convert the prediction output into a probability form as: prob_map = σ(P i ), (2) Among them, P i is the prior feature, and σ(·) represents converting P i into a probability map through Sigmoid; S202: According to the generated probability map, perform non-linear weighting using a Gaussian distribution, specifically: where prob_map is the probability form prediction output in Equation (2), U fore and U back represent the foreground and background probability weights respectively, μ and σ represent the mean and standard deviation of the Gaussian distribution respectively, exp(·) is the exponential function with the base of the natural constant e, represents the foreground uncertainty of the high-level feature map, represents the background uncertainty of the high-level feature map, (·) i represents which stage it is in, i ∈ {1, 2, 3, 4}, the foreground and background uncertainty maps of the low-level feature map and are also generated in the same way.
3. A multi-scale feature fusion detection method for prominent targets in remote sensing images according to claim 2, characterized in that In step S300, multiply the foreground and background uncertainty maps after Gaussian weighting with the foreground and background of the high-level and low-level feature maps for weighted fusion, so that the model can identify the foreground and background regions, enhance the foreground salient features and suppress the background interference, and achieve more accurate classification decisions in regions with high uncertainty, specifically: In formula (4), represents the high-level features initially extracted in the i-th stage, represents the low-level features initially extracted in the (i - 1)-th stage. Concat(·) represents the operation of fusing the two feature maps, and CBR(·) represents the operations of Conv convolution, BatchNorm normalization, and ReLu non-linear activation on the feature map. Feature-H and Feature-L respectively represent the high-level features and low-level features after weighted fusion.
4. A multi-scale feature fusion detection method for prominent targets in remote sensing images according to claim 3, characterized in that, The local multi-scale joint processing described in step S400 includes: Adjust the channel dimension through a point convolution layer, fuse the high-level feature Feature-H and perform batch normalization, then perform non-linear transformation through ReLu, and finally further optimize the feature representation through another point convolution layer and batch normalization. The calculation formula is as shown in (5): In formula (5), F top represents the feature map after the advanced feature map Feature-H first passes through the point convolution layer, batch normalization, and ReLu. PWConv represents the point convolution layer, BN represents the batch normalization operation, and F local represents the feature map F top is the local branch feature map after passing through another point convolution layer and batch normalization.
5. A method for multi-scale feature fusion detection of significant targets in remote sensing images according to claim 4, characterized in that, The global multi-scale joint processing described in step S400 includes: Process the high-level feature Feature-H through a global average pooling layer to capture the global context, and then perform point convolution processing. The calculation formula is as shown in (6): In formula (6), F bottom represents the feature map obtained by globally average pooling the high-level feature map Feature-H. GAP represents global average pooling, and F global represents F bottom the global branch feature map obtained by point convolution processing. PWConv represents the point convolution layer, and BN represents the batch normalization operation.
6. A multi-scale feature fusion detection method for prominent targets in remote sensing images according to claim 5, characterized in that Step S500 includes: Add the outputs of the global branch and the local branch, and obtain the probability distribution of the important features through Sigmoid processing. Then perform focus strengthening on the high-level features. The calculation formulas are as shown in (7) and (8): F prob = σ(F local + F global ), (7) F focus = Feature-H + F prob ·Feature-H, (8) In formula (7), F prob represents the global branch feature map F global added to the local branch feature map F local and the resulting feature map is converted into a probability distribution map through Sigmoid. σ(·) represents the Sigmoid processing of the feature map, and F focus is the output after focusing and strengthening, and Feature-H is the high-level feature map.
7. A multi-scale feature fusion detection method for significant targets in remote sensing images according to claim 5, characterized in that Step S500 includes: Fuse the results after focus strengthening with the low-level feature map to achieve the supplement and enhancement of texture details. The calculation formula is as shown in formula (9): F output = Conv 3×3 (Concat(Upsample(F focus ), Feature - L)), (9) In Equation (9), the final output F output represents the fused feature map F of the focused enhancement feature map focus and the low-level feature map Feature-L to supplement texture detail information. Upsample represents upsampling the feature map, Concat(·) represents the fusion operation of the two feature maps, and Conv 3×3 represents a 3×3 convolution of the feature map.
8. A multi-scale feature fusion detection method for prominent targets in remote sensing images according to claim 7, characterized in that, Step S600 includes: Adopt a deep supervision mechanism to perform layer-by-layer supervision on the predictions at different levels, so that this model can perform salient object detection in the feature space at different scales, and ensure the effective integration of shallow features and deep features. The calculation formula is as shown in (10). In Equation (10), L total represents the overall loss, GT represents the true saliency label map, and P i is the prior feature. L IOU represents the IOU loss, and L BCE represents the BCE loss. λ1 and λ2 are the weight coefficients of the BCE loss and the IOU loss, controlling their relative contributions.
9. A multi-scale feature fusion detection method for prominent targets in remote sensing images according to claim 8, characterized in that Step S100 includes: S101: Conduct preliminary processing on the input remote sensing image, initialize a series of basic configuration parameters, and load its pre-trained weights S102: An encoder-decoder deep supervision framework for constructing multi-scale uncertainty-confused feature resolution is built. For the input remote sensing image I ∈ R C×H×W , the selected feature extractor PVTv2 is used to perform multi-stage feature extraction on the image. These feature maps will serve as the basic data for subsequent operations such as uncertainty evaluation and feature fusion. The calculation formula is shown in (1): F n = Block n (I), n ∈ {1, 2, 3, 4}, (1) In Equation (1), I is the input remote sensing image, and Block n (·) represents the feature extractor; F n represents the feature maps extracted in different stages.
10. A multi-scale feature fusion detection system for prominent targets in remote sensing images, characterized in that, including a feature extraction module, which is used to initialize the calculation parameters of the remote sensing salient object detection model based on probability uncertainty assessment, extract features from the remote sensing image using a feature extractor, and generate a feature map with multiple scales and relatively high uncertainty in the changing edges; a weighted processing module, which is used to use the designed uncertainty assessment mapping module to take the high-level semantic information of the remote sensing image feature map obtained in step S100 as prior knowledge and perform non-linear weighted processing on the pixels with uncertainty; an evaluation mapping module, which is used to reconstruct the probability distribution maps of the foreground and background respectively, and then use this uncertainty information to weighted fuse high-level and low-level features to obtain clearer foreground and background feature representations; a multi-scale focus fusion module, which is used to enhance the focus of the obtained deep features, retain the details of the shallow features, and save them to achieve a high degree of highlighting of important regions and information supplementation of texture details; a multi-scale joint processing module, which is used to perform global and local multi-scale joint processing on deep features and shallow features, suppress the uncertainty of confusing features, and further improve the discrimination ability and detection accuracy of the model; a layer-by-layer supervision module, which is used to layer-by-layer supervise the prediction results of multiple stages of the model output using a combination of binary cross-entropy loss and IOU loss to ensure effective cooperation between features at different levels.